Unsupervised learning helps understand genome-wide patterns and miRNA regulation.
problem Understanding genome-wide biological insights and miRNA regulation.
method Review of unsupervised learning algorithms for genome informatics and miRNA regulation.
result Reviewed several unsupervised learning methods for genome and miRNA analysis.
AI improves healthcare diagnostics and predictions.
problem Data heterogeneity and model limitations in AI for health.
method Review of AI applications in health informatics.
result AI enhances disease diagnosis and prediction.
Geotechnics adopts data-driven methods from materials informatics.
problem Soil complexity and lack of comprehensive data.
method Leveraging deep learning and transfer learning for feature extraction.
result Revolutionary potential of advanced computational tools in geotechnics.
In this paper are presented methods of impact analysis on informatics system security accidents, qualitative and quantitative methods, starting with risk and informational system security definitions. It is presented the relationship between the risks of exploiting vulnerabilities of security system, security level of …
PiNet improves graph classification efficiency and accuracy.
problem Graph level classification challenges.
method Attention-based pooling mechanism for graph convolution operations.
result Superior performance and high sample efficiency.
Big data analytics improves healthcare through early detection and quality life.
problem Limited access to healthcare data hinders evidence-based decision-making.
method Analysis of healthcare data using various tools and techniques.
result Big data analytics enhances healthcare quality and patient outcomes.
Paper optimizes portfolios with non-identical asset return variances using statistical mechanics.
problem Optimizing portfolios with assets having different return variance.
method Replica analysis of statistical mechanical informatics.
result Asymptotic behaviors of minimal investment risk and concentrated investment level determined analytically.
Modeling territorial control in civil wars using HMMs.
problem Lack of fine-grained data on territorial control in civil wars.
method Theoretical model of territorial control and Hidden Markov Models (HMMs).
result HMMs can estimate levels of territorial control in civil wars.
This paper surveys deep learning applications in EHR analysis.
problem Leveraging EHR data for clinical informatics tasks.
method Reviews deep learning architectures and techniques applied to EHR data.
result Identifies current limitations and future research directions.
Paper proposes active learning for structured output design, improving Gaussian process model predictions.
problem Finding optimal input parameters for achieving desired structured outputs.
method Developed new acquisition functions to minimize prediction error of Gaussian process model, incorporating output correlations.
result Effectiveness demonstrated in synthetic and real data experiments, including materials informatics.
Discovering genomic structure through learned neural architectures.
problem Decoding the complex, unknown structure of human genomes using deep learning.
method Developed a novel search algorithm to learn optimal architectures for genomic data.
result Architectures learned from RNA expression data predict gene regulatory structure and identify key sequence motifs.
Deep learning method extracts medical knowledge from YouTube videos.
problem Improving healthcare information dissemination through machine learning.
method Developed a deep learning method to classify YouTube videos by medical knowledge level.
result Preliminary results show satisfactory performance in extracting medical knowledge from videos.
Paper uses genome Markov structure for outlier detection and read classification.
problem Identifying outliers and classifying reads in genome databases.
method Applying second-order Markov models to triplet base distributions.
result Improved accuracy in outlier identification and read classification.
SVM and N-best algorithm classify microbial marker clades from genome sequences.
problem Classifying microbial clades from genome sequences, especially new species.
method Support vector machine (SVM) with N-best algorithm, time series feature extraction, random fragment generation, k-mer size selection.
result Recognition accuracy rates above 28% in top-1 candidate, above 91% in top-10 candidate.
Elastic co-clustering improves clustering of single-cell genomic data.
problem Improving clustering performance of single-cell genomic datasets.
method Elastic coupled co-clustering in an unsupervised transfer learning framework.
result Our algorithm significantly improves clustering performance over traditional methods.
Genomic models learn DNA sequences to predict functions.
problem Understanding complex genetic interactions.
method Training LLMs on DNA sequences to predict functions.
result gLMs can predict functions of DNA elements.
The paper predicts diseases using both clinical and genomics data.
problem Clinical predictions using genomics data are not common.
method Integrated clinical and genomics datasets, machine learning, Principal Component Analysis for feature selection.
result 73% accuracy in predicting 75 disease classes.
Prototype Matching Network (PMN) improves genomic TFBS prediction.
problem Predicting Transcription Factor Binding Sites (TFBSs) with hundreds of TFs as labels.
method Prototype Matching Network (PMN) that learns motif-like features and TF-TF interactions.
result PMN significantly outperforms baselines on a large TFBS dataset.
New method combines ensembling and regularization for genomic disease prediction.
problem Genomic diseases require accurate prediction and biomarker identification.
method Integrates regularization with ensembling techniques for high-dimensional binary classification.
result Identifies critical biomarkers overlooked by competing methods.
Kernel method embeds noisy datasets, capturing shared structures.
problem Limited power in capturing nonlinear structures, noisiness, high-dimensionality, and interpretability issues.
method Kernel spectral joint embeddings using duo-landmark integral operators.
result Consistent recovery of low-dimensional noiseless signals and convergence to eigenfunctions of integral operators.
The increased affordability of whole genome sequencing has motivated its use for phenotypic studies. We address the problem of learning interpretable models for discrete phenotypes from whole genomes. We propose a general approach that relies on the Set Covering Machine and a k-mer representation of the genomes. We sho…
Dilated convolutions model long-distance genomic dependencies effectively.
problem Detecting regulatory elements from raw DNA with long-distance dependencies.
method Developed and used a novel dataset for dilated convolutional neural networks.
result Dilated convolutions are effective at modeling regulatory elements in the human genome.
Dr.S recommends cancer drugs based on genomic data.
problem Personalizing cancer treatments using genomic information.
method Machine learning to identify optimal drug-gene associations.
result Developed a Drug Recommendation System (Dr.S) for cancer cell lines.
Develops robust learning framework under distributional perturbations.
problem Learning robust to data distributional changes.
method Distributionally Robust Optimization (DRO) under Wasserstein metric.
result Establishes performance guarantees and tractable formulations.
Paper finds previous work on submanifolds incorrect.
problem Incorrect definition of semi-invariant submanifolds.
method Examined previous work's definition and found it flawed.
result Previous results on semi-invariant submanifolds are invalid.
New method recovers structured missing data in genomic studies.
problem Structured missingness in genomic data integration.
method Structured Matrix Completion (SMC) for efficient matrix recovery.
result Establishes optimal recovery rate and performs well in simulations and real data.
PKB method uses pathway information for cancer sample classification.
problem Cancer genomic data's high dimensionality and limited sample sizes.
method Pathway-based Kernel Boosting (PKB) method integrating gene pathway information for sample classification.
result PKB method outperforms other methods and identifies relevant pathways.
Generates new human genomic sequences for LAI training.
problem Lack of accessible reference data sets for LAI.
method Class-conditional VAE-GAN to generate realistic sequences.
result Generated sequences improve LAI method performance.
Paper introduces Hedged Bandits for ensemble learning with performance guarantees.
problem Lack of performance guarantees in existing ensemble learning methods.
method Hedged Bandits method providing asymptotic and short run performance guarantees.
result Hedged Bandits method outperforms existing ensemble learning methods.
Proposes FDR-corrected sparse CCA for neuroimaging and genomics.
problem High-dimensional datasets in neuroimaging and genomics make false discoveries a concern.
method FDR-corrected sparse canonical correlation analysis (CCA) for high-dimensional settings.
result The proposed method controls the FDR of canonical vectors in high-dimensional settings.
Develops a faster soybean genome clustering method combining spectral and vector quantization.
problem Clustering soybean whole genome sequences efficiently.
method Combines Spectral Clustering and Vector Quantization for computational efficiency.
result Significantly outperforms existing methods in cluster quality and time complexity.
Copula-based fusion improves breast cancer risk stratification.
problem Combining clinical and genomic risk scores using simple rules fails to capture their joint relationship.
method Used copulas to model the joint relationship between clinical and genomic risk scores.
result Copula-based fusion improves risk stratification, identifying subgroups with the worst prognosis.
Method extracts cancer signatures from genome data, reducing noise and variability.
problem Identifying stable cancer signatures from noisy genomic data.
method Applied statistical risk models from finance to cancer genome data, using NMF.
result Extracted signatures have lower variability and improved stability.
The paper tackles sample complexity in high-dimensional data, focusing on correlation mining.
problem Understanding reliable inference in variable-rich, sample-starved data.
method Develops a unified statistical framework to quantify sample complexity for various inferential tasks.
result Illustrates high-dimensional learning rates and sample complexity for correlation mining.
Private cancer prediction model trained on federated genomic data.
problem Train a private cancer prediction model on federated genomic data.
method Differentially private federated learning (FL) for genomic cancer prediction.
result Ranked 3rd in a competition for private cancer prediction.
SEISM tests neural network features for regulatory genomics.
problem Testing neural network features for regulatory genomics.
method Selective inference procedure for sequence motifs.
result Sampling under specific parameters characterizes composite null hypothesis.
TF-MoDISco finds transcription factor motifs from genomic data.
problem Identifying transcription factor motifs from genomic sequence data.
method Algorithm for motif discovery from basepair-level importance scores.
result Improved version v0.5.6.5 of TF-MoDISco.
Proposes a deep learning framework for evaluating patient similarities from EHRs.
problem Evaluating clinical similarities between patients for various healthcare applications.
method A deep learning framework with medical concept embedding, preserving temporal information.
result Significant improvement in patient similarity evaluation over baselines.
Paper proposes scalable method for analyzing multi-omic data.
problem Integrating high-dimensional multi-omic data for cancer subtyping.
method Mixed graphical model approach using Birth-Death MCMC algorithm.
result Our method outperforms LASSO and standard BDMCMC in computational efficiency and model selection accuracy.
Neural network classifies liver cancer patients based on genomic data.
problem Classifying liver cancer patients into high-risk and low-risk groups.
method Data expansion using wavelet analysis, compression of wavelet coefficients, training a neural network model.
result The neural network model accurately classifies patients without survival time information.
Paper proposes a framework to predict therapeutic properties of compounds.
problem Predict therapeutic properties of compounds with heterogeneous data.
method Domain-adversarial multi-task framework using adversarial learning.
result Framework improves performance over competitive baselines.
Measures DNA quality degradation effects.
problem Identifying degraded DNA sequence data.
method Novel quality quantification based on intentional degradation effects.
result Quantified measures of degradation can be used for multiple purposes.
iRF detects stable high-order interactions in genomics data.
problem Understanding high-order interactions in genomics data.
method Iterative Random Forest algorithm (iRF) for stable high-order interaction detection.
result iRF identifies stable high-order interactions with computational cost similar to Random Forest.
As the amount and complexity of genetic information increases it is necessary that we explore some efficient ways of handling these data. This study takes the "divide and conquer" approach for analyzing high dimensional genomic data. Our aims include reducing the dimensionality of the problem that has to be dealt one a…
Transfer learning improves model accuracy on sparse materials datasets.
problem Irreducible errors in analyses due to differing measurements across datasets.
method Three transfer learning techniques: multi-task, difference, and explicit latent variable architectures.
result Explicit latent variable method is most accurate for activation energies of NO reduction steps.
Method computes embeddings for RNA-seq data without genome alignment.
problem No need for genome alignment for RNA-seq data analysis.
method RNN transforms kmers into 2D latent space for transcriptomic analysis.
result Captures DNA sequence similarity and abundance in latent space.
Understanding functional organization of genetic information is a major challenge in modern biology. Following the initial publication of the human genome sequence in 2001, advances in high-throughput measurement technologies and efficient sharing of research material through community databases have opened up new view…
Review of automatic de-identification systems for EHR, highlighting challenges beyond accuracy.
problem Challenges in surrogate generation and patient privacy in de-identification of EHR.
method Comprehensive review of 18 recently published systems, focusing on accuracy and challenges.
result Despite accuracy improvements, challenges remain in surrogate generation and patient privacy.