Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

115231346461 · Jun 202019922001200920182026
48 results for genomic analysis

Bayesian analysis uncovers flux couplings in metabolic networks.

problem Uncertainty and unrealistic assumptions in traditional flux analysis methods.
method Introduces Bayesian metabolic flux analysis to model reactions probabilistically and infer flux distributions.
result Reveals informative flux couplings and more unobserved fluxes in metabolic networks.

Graphs represent gene segment organization, revealing complex interrelationships in a scrambled genome.

problem Understanding gene segment organization and interrelationships in a scrambled genome.
method Directed graphs representing gene segments and their relationships, with graph properties mapped to higher-dimensional space for analysis.
result Emerging star-like structures indicate complex interrelationships, including segments from multiple genes interleaving or overlapping.

fiBAG integrates multiplatform genomic data to identify disease markers.

problem Understanding complex mechanisms underlying human diseases from multiplatform genomic data.
method fiBAG uses Gaussian process models and Bayes factors to identify functional evidence and guide variable selection.
result fiBAG improves detection of disease-related markers compared to non-integrative methods.

PKB method uses pathway information for cancer sample classification.

problem Cancer genomic data's high dimensionality and limited sample sizes.
method Pathway-based Kernel Boosting (PKB) method integrating gene pathway information for sample classification.
result PKB method outperforms other methods and identifies relevant pathways.

PKB framework boosts genomic data analysis by integrating pathway knowledge.

problem Boosting discovery power and connecting new findings with biological mechanisms in genomic data.
method Pathway-based Kernel Boosting (PKB) framework integrating clinical and pathway information for prediction of various outcomes.
result PKB substantially outperforms other methods in predicting drug response and cancer survival.

Understanding functional organization of genetic information is a major challenge in modern biology. Following the initial publication of the human genome sequence in 2001, advances in high-throughput measurement technologies and efficient sharing of research material through community databases have opened up new view…

2011-02-27abs ↗pdf ↗

Spectral simplicial theory improves feature selection for complex data.

problem Complex data sets and high-dimensional feature spaces require efficient feature selection methods.
method Extends spectral techniques to abstract simplicial complexes, incorporating topological data analysis.
result Spectral simplicial methods provide a unified approach for feature selection in multi-modal genomic data.

The paper develops methods for causal inference from single-cell RNA sequencing data with multiple outcomes.

problem Causal inference from single-cell RNA sequencing data with multiple heterogeneous outcomes.
method Generic semiparametric inference framework for doubly robust estimation with multiple derived outcomes.
result Demonstrates the use of semiparametric inferential results for estimating causal effects in genomics.

SNeCT integrates multi-platform genomic data using Tucker decomposition with network constraints.

problem Integrative analysis of large-scale, high-dimensional, sparse genomic data with prior knowledge incorporation.
method Parallel stochastic gradient descent on a network-constrained optimization function.
result Decomposed factor matrices stratify cancers, find similar patients, and personalize interpretation.

BioBO optimizes gene perturbation design using Bayesian optimization with biological priors.

problem Efficient design of genomic perturbation experiments in drug discovery.
method Integrates Bayesian optimization with multimodal gene embeddings and enrichment analysis.
result Improves labeling efficiency by 25-40% and identifies top-performing perturbations more effectively.

The paper solves a genome assembly problem by recovering hidden Hamiltonian cycles from noisy measurements.

problem Inferring an unknown Hamiltonian cycle in a genome assembly problem from noisy edge measurements.
method Introduced a linear programming relaxation (F2F LP) to recover the hidden Hamiltonian cycle with high probability.
result A simple linear programming relaxation recovers the hidden Hamiltonian cycle with high probability as non o \infty.

Quantile normalisation is a popular normalisation method for data subject to unwanted variations such as images, speech, or genomic data. It applies a monotonic transformation to the feature values of each sample to ensure that after normalisation, they follow the same target distribution for each sample. Choosing a "g…

2017-06-01abs ↗pdf ↗

SVM and N-best algorithm classify microbial marker clades from genome sequences.

problem Classifying microbial clades from genome sequences, especially new species.
method Support vector machine (SVM) with N-best algorithm, time series feature extraction, random fragment generation, k-mer size selection.
result Recognition accuracy rates above 28% in top-1 candidate, above 91% in top-10 candidate.

Elastic co-clustering improves clustering of single-cell genomic data.

problem Improving clustering performance of single-cell genomic datasets.
method Elastic coupled co-clustering in an unsupervised transfer learning framework.
result Our algorithm significantly improves clustering performance over traditional methods.

Canonical Correlation Analysis (CCA) is a classical tool for finding correlations among the components of two random vectors. In recent years, CCA has been widely applied to the analysis of genomic data, where it is common for researchers to perform multiple assays on a single set of patient samples. Recent work has pr…

2012-06-18abs ↗pdf ↗

With the wealth of high-throughput sequencing data generated by recent large-scale consortia, predictive gene expression modelling has become an important tool for integrative analysis of transcriptomic and epigenetic data. However, sequencing data-sets are characteristically large, and previously modelling frameworks …

2015-07-21abs ↗pdf ↗

Prototype Matching Network (PMN) improves genomic TFBS prediction.

problem Predicting Transcription Factor Binding Sites (TFBSs) with hundreds of TFs as labels.
method Prototype Matching Network (PMN) that learns motif-like features and TF-TF interactions.
result PMN significantly outperforms baselines on a large TFBS dataset.

Machine learning accurately diagnoses cancer from whole genome sequencing data.

problem Accurate cancer diagnosis at all stages.
method Novel MLAC (Machine Learning Against Cancer) method using next-gen RNA sequencing.
result Perfect precision, sensitivity, and specificity achieved for most tumor types.

GSAE autoencoder models gene sets for better cancer subtype and prognosis analysis.

problem Inter-gene set associations not considered in gene set-based analyses.
method Gene superset autoencoder model incorporating prior gene sets.
result Gene supersets retain biological features and are reproducible for cancer subtype and prognosis.

Dilated convolutions model long-distance genomic dependencies effectively.

problem Detecting regulatory elements from raw DNA with long-distance dependencies.
method Developed and used a novel dataset for dilated convolutional neural networks.
result Dilated convolutions are effective at modeling regulatory elements in the human genome.

With different genomes available, unsupervised learning algorithms are essential in learning genome-wide biological insights. Especially, the functional characterization of different genomes is essential for us to understand lives. In this book chapter, we review the state-of-the-art unsupervised learning algorithms fo…

2015-08-03abs ↗pdf ↗

Nucleosome positioning is an important process required for proper genome packing and its accessibility to execute the genetic program in a cell-specific, timely manner. In the recent years hundreds of papers have been devoted to the bioinformatics, physics and biology of nucleosome positioning. The purpose of this rev…

2015-08-27abs ↗pdf ↗

Advances of modern sensing and sequencing technologies generate a deluge of high dimensional space-temporal physiological and next-generation sequencing (NGS) data. Physiological traits are observed either as continuous random functions, or on a dense grid and referred to as function-valued traits. Both physiological a…

2014-10-27abs ↗pdf ↗