Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

1122 · Jul 201919922001200920172026
23 results for metabolomics

Bayesian GAMs improve predictive performance for high-dimensional data.

problem Sparse regularization in GAMs leads to excess shrinkage and difficulty in selecting nonlinear effects.
method Developed a novel spike-and-slab LASSO prior and scalable EM-Coordinate Descent algorithm.
result Improved predictive and computational performance compared to existing models.

GEMSS discovers multiple sparse solutions in high-dimensional data.

problem Identifying multiple sparse feature combinations in high-dimensional, underdetermined systems.
method GEMSS (Gaussian Ensemble for Multiple Sparse Solutions) uses a structured spike-and-slab prior, mixture of Gaussians, and Jaccard-based penalty to optimize a single objective function via stochastic gradient descent.
result GEMSS consistently outperforms five feature selection methods on 128 experiments and real-world datasets.

BSFP method reveals latent patterns in multi-omic data for predicting lung function in HIV-associated OLD.

problem Limited understanding of multi-omic molecular phenomena and clinical outcomes in obstructive lung disease.
method Bayesian Simultaneous Factorization and Prediction (BSFP) method for multi-omic data, accommodating imputation and full posterior inference.
result BSFP reveals distinct clusters of patients with OLD and multi-omic patterns related to lung function decline.

In systems biology, it is common to measure biochemical entities at different levels of the same biological system. One of the central problems for the data fusion of such data sets is the heterogeneity of the data. This thesis discusses two types of heterogeneity. The first one is the type of data, such as metabolomic…

2019-08-23abs ↗pdf ↗

This paper extends compositional data analysis using graph signal processing.

problem Traditional log-ratios between all variables are not suitable for specific variable relationships.
method Linking compositional data analysis with graph signal processing, it considers only selected log-ratios.
result The approach retains desirable properties of scale invariance and compositional coherence.

iDeepViewLearn combines deep learning and feature selection for multiview learning.

problem Learning nonlinear relationships in data from multiple complementary views.
method Combines deep learning flexibility with statistical feature selection using deep neural networks and graph Laplacian regularization.
result Identifies genes and CpG sites that differentiate between breast cancer survivors and non-survivors.

In the Pioneer 100 (P100) Wellness Project (Price and others, 2017), multiple types of data are collected on a single set of healthy participants at multiple timepoints in order to characterize and optimize wellness. One way to do this is to identify clusters, or subgroups, among the participants, and then to tailor pe…

2019-01-12abs ↗pdf ↗

Pipeline integrates cross-sectional and longitudinal multi-omics data for IBD research.

problem Integrating diverse data types from the same individuals for disease understanding.
method Statistical and deep learning methods for variable selection, feature extraction, and joint integration.
result Identified microbial pathways, metabolites, and genes discriminating IBD status.

engGNN combines external and generated graphs to improve disease classification and biomarker discovery.

problem Challenges in integrating omics data due to high dimensionality and small sample sizes.
method Dual-graph framework that integrates external biological networks with data-driven generated graphs.
result engGNN outperforms state-of-the-art methods in disease classification and biomarker discovery.

Selective prediction framework reduces errors in molecular structure identification from MS/MS.

problem High-stakes applications require reliable molecular structure identification from MS/MS data.
method Selective prediction framework using risk-coverage tradeoff and uncertainty quantification.
result First-order confidence measures and retrieval-level aleatoric uncertainty achieve strong risk-coverage tradeoffs.

With the maturation of metabolomics science and proliferation of biobanks, clinical metabolic profiling is an increasingly opportunistic frontier for advancing translational clinical research. Automated Machine Learning (AutoML) approaches provide exciting opportunity to guide feature selection in agnostic metabolic pr…

2017-10-09abs ↗pdf ↗

BOOOM optimizes orthonormal matrices without needing gradients.

problem Optimizing over the Stiefel manifold in non-convex, non-smooth settings.
method Global Givens rotation-based parametrization and Recursive Modified Pattern Search.
result BOOOM achieves strong performance across various optimization problems.

A new approach to the sparse Canonical Correlation Analysis (sCCA)is proposed with the aim of discovering interpretable associations in very high-dimensional multi-view, i.e.observations of multiple sets of variables on the same subjects, problems. Inspired by the sparse PCA approach of Journee et al. (2010), we also s…

2019-09-17abs ↗pdf ↗