Novel method identifies proteomic risk markers for Alzheimer disease.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Background: High-throughput proteomics techniques, such as mass spectrometry (MS)-based approaches, produce very high-dimensional data-sets. In a clinical setting one is often interested in how mass spectra differ between patients of different classes, for example spectra from healthy patients vs. spectra from patients…
Deep Learning can significantly benefit cancer proteomics and genomics. In this study, we attempt to determine a set of critical proteins that are associated with the FLT3-ITD mutation in newly-diagnosed acute myeloid leukemia patients. A Deep Learning network consisting of autoencoders forming a hierarchical model fro…
Liquid chromatography coupled with tandem mass spectrometry, also known as shotgun proteomics, is a widely-used high-throughput technology for identifying proteins in complex biological samples. Analysis of the tens of thousands of fragmentation spectra produced by a typical shotgun proteomics experiment begins by assi…
We introduce Graph-Sparse Logistic Regression, a new algorithm for classification for the case in which the support should be sparse but connected on a graph. We val- idate this algorithm against synthetic data and benchmark it against L1-regularized Logistic Regression. We then explore our technique in the bioinformat…
A novel GPDA method for high-dimensional functional data.
Study repurposes open data to find potential COVID-19 drugs.
New method improves robustness of double robust estimators under complete misspecification.
In the Pioneer 100 (P100) Wellness Project (Price and others, 2017), multiple types of data are collected on a single set of healthy participants at multiple timepoints in order to characterize and optimize wellness. One way to do this is to identify clusters, or subgroups, among the participants, and then to tailor pe…
GPU optimization speeds up large-scale classification tasks.
BSFP method reveals latent patterns in multi-omic data for predicting lung function in HIV-associated OLD.
Valid inference from data and predictions.
High-dimensional data common in genomics, proteomics, and chemometrics often contains complicated correlation structures. Recently, partial least squares (PLS) and Sparse PLS methods have gained attention in these areas as dimension reduction techniques in the context of supervised data analysis. We introduce a framewo…
Variable selection in high dimensional space has challenged many contemporary statistical problems from many frontiers of scientific disciplines. Recent technology advance has made it possible to collect a huge amount of covariate information such as microarray, proteomic and SNP data via bioimaging technology while ob…
Sparse neural networks visualize paired transcriptomic and electrophysiological data.
In systems biology, it is common to measure biochemical entities at different levels of the same biological system. One of the central problems for the data fusion of such data sets is the heterogeneity of the data. This thesis discusses two types of heterogeneity. The first one is the type of data, such as metabolomic…
A new hybrid federated learning algorithm for combining clinical and omics data.
Federated learning improves bioinformatics by sharing data legally.
Undirected graphical models, or Markov networks, are a popular class of statistical models, used in a wide variety of applications. Popular instances of this class include Gaussian graphical models and Ising models. In many settings, however, it might not be clear which subclass of graphical models to use, particularly…
Unified framework for robust causal directionality in quantum systems under MNAR observation.
Ensemble models provide more accurate feature importance estimates than single models.
Accurately predicting drug responses to cancer is an important problem hindering oncologists' efforts to find the most effective drugs to treat cancer, which is a core goal in precision medicine. The scientific community has focused on improving this prediction based on genomic, epigenomic, and proteomic datasets measu…
The worldwide surge of multiresistant microbial strains has propelled the search for alternative treatment options. The study of Protein-Protein Interactions (PPIs) has been a cornerstone in the clarification of complex physiological and pathogenic processes, thus being a priority for the identification of vital compon…
Biological data are extremely diverse, complex but also quite sparse. The recent developments in deep learning methods are offering new possibilities for the analysis of complex data. However, it is easy to be get a deep learning model that seems to have good results but is in fact either overfitting the training data …
Active inference uses machine learning to prioritize data labeling for more efficient statistical inference.
We report a scalable hybrid quantum-classical machine learning framework to build Bayesian networks (BN) that captures the conditional dependence and causal relationships of random variables. The generation of a BN consists of finding a directed acyclic graph (DAG) and the associated joint probability distribution of t…
Study of Bayes optimal learning in high-dimensional linear regression with network side information.
New method estimates sparse canonical vectors efficiently.
The rapid development of high-throughput technologies has enabled the generation of data from biological or disease processes that span multiple layers, like genomic, proteomic or metabolomic data, and further pertain to multiple sources, like disease subtypes or experimental conditions. In this work, we propose a gene…
"Mixed Data" comprising a large number of heterogeneous variables (e.g. count, binary, continuous, skewed continuous, among other data types) are prevalent in varied areas such as genomics and proteomics, imaging genetics, national security, social networking, and Internet advertising. There have been limited efforts a…
Synthetic data can be used to ask more questions and accelerate discovery with provable validity guarantees.
Scalable methods integrate multiview data for clinical outcomes.
Over the past decades, statisticians and machine-learning researchers have developed literally thousands of new tools for the reduction of high-dimensional data in order to identify the variables most responsible for a particular trait. These tools have applications in a plethora of settings, including data analysis in…
In this paper, we study the challenge of feature selection based on a relatively small collection of sample pairs . The observations are thereby supposed to follow a noisy single-index model, depending on a certain set of signal variables. A major difficulty is tha…
OPAL optimizes labeling strategy for precise inference from uncertain models.
A transfer learning method builds high-dimensional models using disparate datasets.
Proposes improved classification via transfer learning with regularized linear discriminant analysis.
The problem of learning the structure of a high dimensional graphical model from data has received considerable attention in recent years. In many applications such as sensor networks and proteomics it is often expensive to obtain samples from all the variables involved simultaneously. For instance, this might involve …
CARVE validates clustering results using resampling and stability analysis.
sJIVE combines structure and prediction in multi-source data.
New method stabilizes model selection with theoretical guarantees.
iDeepViewLearn combines deep learning and feature selection for multiview learning.
engGNN combines external and generated graphs to improve disease classification and biomarker discovery.
A new method for combining multiple data views in supervised learning.
Bayesian model clusters diverse 'omics data for disease subtyping.
Bayesian models link multiview data to outcomes.
New algorithm improves on EM for streaming data, outperforming existing methods.
Unified framework for SGMoE resolves estimation and selection issues.