Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

4.9%9.7%14.6%19.5% · Mar 202619922001200920182026
48 results for Genetic risk factors

A distributed feature selection framework identifies genetic risk factors for Alzheimer's disease.

problem High dimensionality of GWAS data makes it hard to detect genetic risk factors for Alzheimer's disease.
method Distributed Feature Selection Framework (DFSF) with distributed group Lasso screening rules and stability selection.
result The method efficiently identifies relevant genetic risk factors for Alzheimer's disease across multiple institutions.

Proposes a hybrid deep learning network for better heart failure survival prediction.

problem Improving survival prediction in heart failure patients.
method Joint analysis of cardiac motion features and clinical risk factors using a hybrid deep learning network.
result Optimal integration of clinical risk factors into deep prediction networks.

Novel method identifies proteomic risk markers for Alzheimer disease.

problem Lack of comprehensive proteomic risk markers for Alzheimer disease diagnosis.
method Deep belief network-based feature selection method using proteomic and clinical data.
result Identified an optimal subset of proteins achieving 90% accuracy in Alzheimer disease diagnosis.

ENN method uses expectile regression for genetic data analysis of complex diseases.

problem Discover additional genetic variants contributing to complex diseases.
method Developed an expectile neural network (ENN) method integrating expectile regression and neural networks.
result ENN method outperforms existing expectile regression in discovering genetic variants predisposing to sub-populations.

BayesMR estimates causal effects and directionality from genetic data.

problem Challenges in finding good genetic instruments and estimating causal effects.
method Bayesian Mendelian randomization approach that accounts for pleiotropy and reverse causation.
result BayesMR provides a posterior distribution over causal effects and uncertainty.

Sparse GFA identifies disease factors in FTD subgroups.

problem Heterogeneity in neurological disorders hinders understanding and treatment.
method Sparse Group Factor Analysis (GFA) with regularised horseshoe priors.
result Identified latent disease factors differentially expressed in FTD subgroups.

GA-MSSR optimizes forex trading rules for higher returns and reduced risk.

problem Noisy market data affects the consistency and profitability of trading algorithms.
method Optimized trading rules derived from technical indicators using a Genetic Algorithm.
result GA-MSSR achieved superior performance with significant positive returns and reduced risk factors.

This paper optimizes Iran's stock portfolio using neural networks and genetic algorithms.

problem Optimizing capital allocation in Iran's stock market with low risk and high return.
method Markowitz Mean-Variance-Skewness model with neural network prediction of stock returns and risks.
result Designing 8 different portfolios for various risk tolerance levels.

This study introduces a new GAS blending ensemble model for Bitcoin price prediction.

problem Predicting Bitcoin price fluctuations in the cryptocurrency market.
method Integrates advanced ensemble learning methods, feature selection algorithms, and sentiment analysis.
result The GAS model demonstrates excellent performance in daily Bitcoin trend prediction.

Enhances genetic programming for stock alpha discovery with warm start and structural constraints.

problem Overwhelming search space and computational burden in traditional genetic programming for alpha factor discovery.
method Proposes a new GP framework with warm start and structural constraints to enhance search performance and interpretability.
result Superior out-of-sample prediction results and higher portfolio returns compared to benchmarks.

New methods improve genetic studies of complex diseases.

problem Improving genetic studies of complex diseases using high-dimensional clinical data.
method Evaluation of unsupervised disentangled representation learning methods (autoencoders, VAE, beta-VAE, FactorVAE) for genetic association studies.
result FactorVAEs and beta-VAEs outperform standard VAEs and non-variational autoencoders in genetic studies of asthma and COPD.

Gradient boosting enhances existing Mendelian models for genetic disease risk prediction.

problem Improving existing Mendelian models for genetic disease risk prediction.
method Combining gradient boosting with existing Mendelian models.
result Improved model outperforms both original and gradient boosting-only models.

This paper optimizes portfolio rebalancing under uncertain security returns using meta-heuristic algorithms.

problem Optimizing portfolio rebalancing under uncertain security returns with transaction costs.
method Meta-heuristic algorithms (genetic algorithm) for solving the portfolio rebalancing problem.
result Meta-heuristic algorithms provide better results than global optimization solvers for portfolio rebalancing under uncertainty.

Lapse-supported life insurance exacerbates adverse selection risks.

problem Lapse-supported life insurance increases adverse selection costs.
method Modeling 'Term to 100' contracts and analyzing three methods of managing lapse surplus.
result Adverse selection losses can be almost unlimited under certain conditions.

A new model selects low-carbon mutual funds considering ESG criteria, risk, and investor preferences.

problem Aligning financial investments with a low-carbon economy.
method Tri-criterion portfolio selection model using a preference-based multi-objective genetic algorithm (ev-MOGA).
result The model successfully incorporates carbon risk exposure and loss-adverse attitudes into portfolio construction.

DL/FBF improves GPSR solutions by selecting compact, generalising expressions.

problem Overfitting and structural bloat in symbolic regression with genetic programming.
method Description length (DL) and fractional Bayes factor (FBF) criteria for selecting compact, generalising expressions.
result DL/FBF post-selection improves test performance compared to AIC/BIC baseline.

Metaheuristics optimize portfolios with pre-assignment and margin trading for better risk-adjusted returns.

problem Maximizing returns while minimizing risk in portfolio optimization.
method Incorporates pre-assignment constraints and margin trading strategies using Genetic Algorithms and Particle Swarm Optimization.
result Metaheuristic-based portfolio optimization yields superior risk-adjusted returns compared to traditional methods.

Develops a SAS approach for high-dimensional risk prediction using unlabeled data.

problem Challenges in risk modeling with EHR data due to lack of direct disease outcomes and high dimensionality.
method Surrogate Assisted Semi-supervised Learning (SAS) approach leveraging unlabeled and labeled data.
result Valid inference for predicted risk even when underlying model is dense and mis-specified.

Method controls extrapolation in prediction profiles for statistical and machine learning models.

problem Avoiding invalid predictions due to extrapolation in prediction profiles.
method Genetic algorithm optimization over constrained factor regions.
result Optimal factor settings without constraint are often invalid and extrapolated.

Optimizes stock portfolios with profit, risk, and sustainability.

problem Balancing profit, risk, and sustainability in stock portfolio management.
method Developed a novel utility function combining Sharpe ratio and ESG scores; used genetic algorithm for optimization.
result System outperforms traditional reinforcement learning methods and improves on risk and sustainability metrics.

C-mix model for censored durations in genetic data, improving prediction accuracy.

problem Predicting adverse events in genetic datasets with high-dimensional covariates.
method High-dimensional mixture model with Elastic-Net regularization and QNEM algorithm.
result Outperforms state-of-the-art models in terms of C-index and AUC(t).

Linear Mixed Models (LMMs) are important tools in statistical genetics. When used for feature selection, they allow to find a sparse set of genetic traits that best predict a continuous phenotype of interest, while simultaneously correcting for various confounding factors such as age, ethnicity and population structure…

2015-07-16abs ↗pdf ↗

Genetic Programming constructs features for physics experiments, improving classification accuracy.

problem Lack of interpretable feature construction for experimental physics.
method Combining Genetic Programming with dimensional consistency constraints.
result Constructed features improve classification accuracy by a significant margin.

Binacox detects multiple cut-points in high-dimensional Cox models for genetic cancer data.

problem Detecting multiple cut-points in high-dimensional Cox models with many continuous features.
method Combines one-hot encoding with binarsity penalty for feature selection and regularization.
result Significantly outperforms state-of-the-art survival models in terms of C-index and computational speed.

Neural networks improve cancer risk prediction from family history data.

problem Improving cancer risk prediction from family history data using machine learning.
method Developed and trained neural network models on large pedigrees to predict hereditary cancers.
result Neural networks can achieve nearly optimal prediction performance and outperform traditional models in misreported data.

Kernel method detects higher order interactions in multi-view data for schizophrenia.

problem Detecting higher order interactions in multi-view biological data.
method Kernel method on reproducing kernel Hilbert space (RKHS) with mixed-effects linear model.
result Identified 13 triplets with significant correlations to hippocampal volume in schizophrenia.

Proposes a two-stage method for estimating heterogeneous treatment effects using gradient boosting trees.

problem Estimating heterogeneous treatment effects in randomized clinical trials with high-dimensional predictive markers.
method Two-stage statistical learning procedure using gradient boosting trees (XGBoost) to estimate main effects and HTE.
result Improves efficiency in estimating heterogeneous treatment effects through nonparametric function estimation.

Advances of modern sensing and sequencing technologies generate a deluge of high dimensional space-temporal physiological and next-generation sequencing (NGS) data. Physiological traits are observed either as continuous random functions, or on a dense grid and referred to as function-valued traits. Both physiological a…

2014-10-27abs ↗pdf ↗

BoGA combines evolutionary search with Bayesian optimization for efficient protein design.

problem Designing novel proteins with specific characteristics is challenging due to sequence space complexity.
method BoGA integrates a genetic algorithm with Bayesian optimization to efficiently explore sequence space.
result BoGA accelerates discovery of high-confidence binders for diverse protein design objectives.

Efficiently infers graph edges from genetic similarity data in landscape genetics.

problem Inferring unknown graph edges from genetic similarity data in a heterogeneous landscape.
method Developed an efficient first-order optimization method to solve the inverse landscape genetics problem.
result Our method provides fast and reliable convergence, significantly outperforming existing heuristics.