A simple algorithm for GWAS estimating SNP effects.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Secure linear regression at speed of plaintext methods.
ParKCa combines multiple causal inference methods to infer new causes from known and unknown factors.
Deep learning detects genetic interactions in type 2 diabetes.
Common complex diseases are likely influenced by the interplay of hundreds, or even thousands, of genetic variants. Converging evidence shows that genetic variants with low marginal effects (LME) play an important role in disease development. Despite their potential significance, discovering LME genetic variants and as…
Method clusters SNPs to improve detection of disease-associated variants.
New methods improve genetic studies of complex diseases.
Method distinguishes genetic correlations from causation in GWAS.
Deep models improve GWAS by identifying genetic interactions.
A genome-wide association study (GWAS) correlates marker variation with trait variation in a sample of individuals. Each study subject is genotyped at a multitude of SNPs (single nucleotide polymorphisms) spanning the genome. Here we assume that subjects are unrelated and collected at random and that trait values are n…
As an increasing number of genome-wide association studies reveal the limitations of attempting to explain phenotypic heritability by single genetic loci, there is growing interest for associating complex phenotypes with sets of genetic loci. While several methods for multi-locus mapping have been proposed, it is often…
SEISM tests neural network features for regulatory genomics.
A new model identifies genetic risk factors using gene-level priors.
In statistical genetics an important task involves building predictive models for the genotype-phenotype relationships and thus attribute a proportion of the total phenotypic variance to the variation in genotypes. Numerous models have been proposed to incorporate additive genetic effects into models for prediction or …
Genome-wide association studies have proven to be essential for understanding the genetic basis of disease. However, many complex traits---personality traits, facial features, disease subtyping---are inherently high-dimensional, impeding simple approaches to association mapping. We developed a nonparametric Bayesian re…
Bayesian method for robust causal inference using many-dimensional instrumental variables.
Proposes a two-stage method for testing variable interactions with FDR control.
Following the publication of an attack on genome-wide association studies (GWAS) data proposed by Homer et al., considerable attention has been given to developing methods for releasing GWAS data in a privacy-preserving way. Here, we develop an end-to-end differentially private method for solving regression problems wi…
LEARNER improves low-rank matrix estimation using source population data.
We introduce the C++ application and R package ranger. The software is a fast implementation of random forests for high dimensional data. Ensembles of classification, regression and survival trees are supported. We describe the implementation, provide examples, validate the package with a reference implementation, and …
An approach for learning ancestral causal relationships in high dimensions, validated on human genome-wide data.
regularized logistic regression has now become a workhorse of data mining and bioinformatics: it is widely used for many classification problems, particularly ones with many features. However, regularization typically selects too many features and that so-called false positives are unavoidable. In this pape…
Proposes spBART for risk prediction using epigenetic signatures and covariates.
Improves normalizing flows by incorporating data dependencies.
With the wealth of high-throughput sequencing data generated by recent large-scale consortia, predictive gene expression modelling has become an important tool for integrative analysis of transcriptomic and epigenetic data. However, sequencing data-sets are characteristically large, and previously modelling frameworks …
A new algorithm estimates causal effects in studies with multiple causes.
Federated learning improves bioinformatics by sharing data legally.
New method accurately identifies causal genes from GWAS data.
With different genomes available, unsupervised learning algorithms are essential in learning genome-wide biological insights. Especially, the functional characterization of different genomes is essential for us to understand lives. In this book chapter, we review the state-of-the-art unsupervised learning algorithms fo…
SAERMA combines deep learning and rule mining to identify SNP interactions.
Enhances FDR control in variable selection using neural networks.
Efficiently solves Elastic Net in high dimensions with Newton method.
Genome-wide association studies (GWAS) offer new opportunities to identify genetic risk factors for Alzheimer's disease (AD). Recently, collaborative efforts across different institutions emerged that enhance the power of many existing techniques on individual institution data. However, a major barrier to collaborative…
IEN speeds up T-Rex+GVS for fast, efficient GWAS.
New method learns parameter groups and structures in multi-response models.
Constrained least squares regression is an essential tool for high-dimensional data analysis. Given a partition of input variables, this paper considers a particular class of nonconvex constraint functions that encourage the linear model to select a small number of variables from a small number of groups …
Genome-wide association studies (GWAS) have achieved great success in the genetic study of Alzheimer's disease (AD). Collaborative imaging genetics studies across different research institutions show the effectiveness of detecting genetic risk factors. However, the high dimensionality of GWAS data poses significant cha…
New models capture complex genetic causes of diseases.
When performing regression on a dataset with variables, it is often of interest to go beyond using main linear effects and include interactions as products between individual variables. For small-scale problems, these interactions can be computed explicitly but this leads to a computational complexity of at least $…
Develops methods for GWAS of high dimensional phenotypes using summary statistics.
Generates new human genomic sequences for LAI training.
Machine learning has been gaining traction in recent years to meet the demand for tools that can efficiently analyze and make sense of the ever-growing databases of biomedical data in health care systems around the world. However, effectively using machine learning methods requires considerable domain expertise, which …
In genome-wide interaction studies, to detect gene-gene interactions, most methods are divided into two folds: single nucleotide polymorphisms (SNP) based and gene-based methods. Basically, the methods based on the gene are more effective than the methods based on a single SNP. Recent years, while the kernel canonical …
Motivation: Analysis of relationships of drug structure to biological response is key to understanding off-target and unexpected drug effects, and for developing hypotheses on how to tailor drug thera-pies. New methods are required for integrated analyses of a large number of chemical features of drugs against the corr…
In high-dimensional linear models, the sparsity assumption is typically made, stating that most of the parameters are equal to zero. Under the sparsity assumption, estimation and, recently, inference have been well studied. However, in practice, sparsity assumption is not checkable and more importantly is often violate…
T-Rex selector selects variables fast and controls FDR in high-dimensional data.
The paper analyzes the power of MX CI tests and finds likelihood-based statistics most powerful.
New method reduces memory usage for high-dimensional variable selection.