A distributed feature selection framework identifies genetic risk factors for Alzheimer's disease.
problem High dimensionality of GWAS data makes it hard to detect genetic risk factors for Alzheimer's disease.
method Distributed Feature Selection Framework (DFSF) with distributed group Lasso screening rules and stability selection.
result The method efficiently identifies relevant genetic risk factors for Alzheimer's disease across multiple institutions.
Deep models improve GWAS by identifying genetic interactions.
problem Missing non-linear interaction effects in GWAS.
method Gradient-based DeepLIFT technique to interpret deep models.
result Known and novel genetic risk factors identified.
New method groups genetic data into coherent topics for disease insights.
problem Analyzing large, multi-dimensional genetic data sets.
method Conditional Hierarchical Bayesian Tucker Decomposition for genetic data analysis.
result Our models are more coherent than baseline models.
Genome-wide association studies (GWAS) offer new opportunities to identify genetic risk factors for Alzheimer's disease (AD). Recently, collaborative efforts across different institutions emerged that enhance the power of many existing techniques on individual institution data. However, a major barrier to collaborative…
System identifies health risks using semantic and machine learning.
problem Identifying risk factors associated with health conditions in subpopulations.
method Developed a combined semantic and machine learning system using a health risk ontology and knowledge graph.
result Dynamic discovery of risk factors and their subpopulations.
A new model identifies genetic risk factors using gene-level priors.
problem Identifying genetic risk factors from nucleotide-level genetic variants.
method Sparse Group Lasso with Group-level Graph structure (SGLGG) model.
result SGLGG effectively identifies phenotype-associated risk SNPs.
Proposes a hybrid deep learning network for better heart failure survival prediction.
problem Improving survival prediction in heart failure patients.
method Joint analysis of cardiac motion features and clinical risk factors using a hybrid deep learning network.
result Optimal integration of clinical risk factors into deep prediction networks.
Novel method identifies proteomic risk markers for Alzheimer disease.
problem Lack of comprehensive proteomic risk markers for Alzheimer disease diagnosis.
method Deep belief network-based feature selection method using proteomic and clinical data.
result Identified an optimal subset of proteins achieving 90% accuracy in Alzheimer disease diagnosis.
ENN method uses expectile regression for genetic data analysis of complex diseases.
problem Discover additional genetic variants contributing to complex diseases.
method Developed an expectile neural network (ENN) method integrating expectile regression and neural networks.
result ENN method outperforms existing expectile regression in discovering genetic variants predisposing to sub-populations.
BayesMR estimates causal effects and directionality from genetic data.
problem Challenges in finding good genetic instruments and estimating causal effects.
method Bayesian Mendelian randomization approach that accounts for pleiotropy and reverse causation.
result BayesMR provides a posterior distribution over causal effects and uncertainty.
Sparse GFA identifies disease factors in FTD subgroups.
problem Heterogeneity in neurological disorders hinders understanding and treatment.
method Sparse Group Factor Analysis (GFA) with regularised horseshoe priors.
result Identified latent disease factors differentially expressed in FTD subgroups.
GA-MSSR optimizes forex trading rules for higher returns and reduced risk.
problem Noisy market data affects the consistency and profitability of trading algorithms.
method Optimized trading rules derived from technical indicators using a Genetic Algorithm.
result GA-MSSR achieved superior performance with significant positive returns and reduced risk factors.
This paper optimizes Iran's stock portfolio using neural networks and genetic algorithms.
problem Optimizing capital allocation in Iran's stock market with low risk and high return.
method Markowitz Mean-Variance-Skewness model with neural network prediction of stock returns and risks.
result Designing 8 different portfolios for various risk tolerance levels.
This study introduces a new GAS blending ensemble model for Bitcoin price prediction.
problem Predicting Bitcoin price fluctuations in the cryptocurrency market.
method Integrates advanced ensemble learning methods, feature selection algorithms, and sentiment analysis.
result The GAS model demonstrates excellent performance in daily Bitcoin trend prediction.
Enhances genetic programming for stock alpha discovery with warm start and structural constraints.
problem Overwhelming search space and computational burden in traditional genetic programming for alpha factor discovery.
method Proposes a new GP framework with warm start and structural constraints to enhance search performance and interpretability.
result Superior out-of-sample prediction results and higher portfolio returns compared to benchmarks.
New models capture complex genetic causes of diseases.
problem Capturing causal relationships between genetic factors and diseases.
method Implicit causal models combining neural architectures and Bayesian inference.
result Significantly outperformed existing genetics methods.
Study develops a predictive model to reduce hospital readmissions.
problem Inaccurate readmission risk prediction models in clinical settings.
method Used Genetic Algorithm and Greedy Ensemble to optimize a readmission risk prediction model.
result Developed a useful risk prediction model for reducing unplanned readmissions.
New methods improve genetic studies of complex diseases.
problem Improving genetic studies of complex diseases using high-dimensional clinical data.
method Evaluation of unsupervised disentangled representation learning methods (autoencoders, VAE, beta-VAE, FactorVAE) for genetic association studies.
result FactorVAEs and beta-VAEs outperform standard VAEs and non-variational autoencoders in genetic studies of asthma and COPD.
Gradient boosting enhances existing Mendelian models for genetic disease risk prediction.
problem Improving existing Mendelian models for genetic disease risk prediction.
method Combining gradient boosting with existing Mendelian models.
result Improved model outperforms both original and gradient boosting-only models.
This paper optimizes portfolio rebalancing under uncertain security returns using meta-heuristic algorithms.
problem Optimizing portfolio rebalancing under uncertain security returns with transaction costs.
method Meta-heuristic algorithms (genetic algorithm) for solving the portfolio rebalancing problem.
result Meta-heuristic algorithms provide better results than global optimization solvers for portfolio rebalancing under uncertainty.
Lapse-supported life insurance exacerbates adverse selection risks.
problem Lapse-supported life insurance increases adverse selection costs.
method Modeling 'Term to 100' contracts and analyzing three methods of managing lapse surplus.
result Adverse selection losses can be almost unlimited under certain conditions.
Bayesian method identifies multi-way interactions among predictors.
problem Identifying meaningful interactions among multiple variables.
method Factorization mechanism and Gibbs sampling for posterior inference.
result Posterior consistency of the regression model.
A new model selects low-carbon mutual funds considering ESG criteria, risk, and investor preferences.
problem Aligning financial investments with a low-carbon economy.
method Tri-criterion portfolio selection model using a preference-based multi-objective genetic algorithm (ev-MOGA).
result The model successfully incorporates carbon risk exposure and loss-adverse attitudes into portfolio construction.
New algorithm predicts lung cancer progression and mortality.
problem Predicting semi-competing risk outcomes in lung cancer.
method Neural Expectation-Maximization algorithm for multi-state outcomes.
result Estimates non-parametric baseline hazards and risk functions.
DL/FBF improves GPSR solutions by selecting compact, generalising expressions.
problem Overfitting and structural bloat in symbolic regression with genetic programming.
method Description length (DL) and fractional Bayes factor (FBF) criteria for selecting compact, generalising expressions.
result DL/FBF post-selection improves test performance compared to AIC/BIC baseline.
Metaheuristics optimize portfolios with pre-assignment and margin trading for better risk-adjusted returns.
problem Maximizing returns while minimizing risk in portfolio optimization.
method Incorporates pre-assignment constraints and margin trading strategies using Genetic Algorithms and Particle Swarm Optimization.
result Metaheuristic-based portfolio optimization yields superior risk-adjusted returns compared to traditional methods.
Deep learning detects genetic interactions in type 2 diabetes.
problem Detecting genetic interactions in complex diseases like type 2 diabetes.
method Stacked Autoencoder for non-linear epistatic interactions.
result Deep learning can uncover missing heritability in complex diseases.
Develops a SAS approach for high-dimensional risk prediction using unlabeled data.
problem Challenges in risk modeling with EHR data due to lack of direct disease outcomes and high dimensionality.
method Surrogate Assisted Semi-supervised Learning (SAS) approach leveraging unlabeled and labeled data.
result Valid inference for predicted risk even when underlying model is dense and mis-specified.
ADNN uses prior knowledge to construct financial features.
problem Feature construction in financial trading.
method Tailored neural network structure with domain knowledge.
result ADNN constructs more informative features than genetic programming.
Method controls extrapolation in prediction profiles for statistical and machine learning models.
problem Avoiding invalid predictions due to extrapolation in prediction profiles.
method Genetic algorithm optimization over constrained factor regions.
result Optimal factor settings without constraint are often invalid and extrapolated.
Optimizes stock portfolios with profit, risk, and sustainability.
problem Balancing profit, risk, and sustainability in stock portfolio management.
method Developed a novel utility function combining Sharpe ratio and ESG scores; used genetic algorithm for optimization.
result System outperforms traditional reinforcement learning methods and improves on risk and sustainability metrics.
Method screens weakly associated predictors in high-dimensional data.
problem Identifying weakly associated predictors in ultrahigh-dimensional data.
method Covariance-insured screening methodology.
result Validates the method through simulations and real data studies.
We study the performance of various agent strategies in an artificial investment scenario. Agents are equipped with a budget, x(t), and at each time step invest a particular fraction, q(t), of their budget. The return on investment (RoI), r(t), is characterized by a periodic function with different types and leve…
C-mix model for censored durations in genetic data, improving prediction accuracy.
problem Predicting adverse events in genetic datasets with high-dimensional covariates.
method High-dimensional mixture model with Elastic-Net regularization and QNEM algorithm.
result Outperforms state-of-the-art models in terms of C-index and AUC(t).
Linear Mixed Models (LMMs) are important tools in statistical genetics. When used for feature selection, they allow to find a sparse set of genetic traits that best predict a continuous phenotype of interest, while simultaneously correcting for various confounding factors such as age, ethnicity and population structure…
Genetic Programming constructs features for physics experiments, improving classification accuracy.
problem Lack of interpretable feature construction for experimental physics.
method Combining Genetic Programming with dimensional consistency constraints.
result Constructed features improve classification accuracy by a significant margin.
Binacox detects multiple cut-points in high-dimensional Cox models for genetic cancer data.
problem Detecting multiple cut-points in high-dimensional Cox models with many continuous features.
method Combines one-hot encoding with binarsity penalty for feature selection and regularization.
result Significantly outperforms state-of-the-art survival models in terms of C-index and computational speed.
Neural networks improve cancer risk prediction from family history data.
problem Improving cancer risk prediction from family history data using machine learning.
method Developed and trained neural network models on large pedigrees to predict hereditary cancers.
result Neural networks can achieve nearly optimal prediction performance and outperform traditional models in misreported data.
Kernel method detects higher order interactions in multi-view data for schizophrenia.
problem Detecting higher order interactions in multi-view biological data.
method Kernel method on reproducing kernel Hilbert space (RKHS) with mixed-effects linear model.
result Identified 13 triplets with significant correlations to hippocampal volume in schizophrenia.
Proposes a two-stage method for estimating heterogeneous treatment effects using gradient boosting trees.
problem Estimating heterogeneous treatment effects in randomized clinical trials with high-dimensional predictive markers.
method Two-stage statistical learning procedure using gradient boosting trees (XGBoost) to estimate main effects and HTE.
result Improves efficiency in estimating heterogeneous treatment effects through nonparametric function estimation.
Kernel method optimizes personalized dose rules for patients.
problem Finding optimal individualized dose rules for patients.
method Kernel assisted learning method for estimating optimal dose rules.
result The method identifies the optimal individualized dose rule and produces favorable outcomes.
Advances of modern sensing and sequencing technologies generate a deluge of high dimensional space-temporal physiological and next-generation sequencing (NGS) data. Physiological traits are observed either as continuous random functions, or on a dense grid and referred to as function-valued traits. Both physiological a…
New method models construction safety risks using injury reports.
problem Improving construction safety through empirical and quantitative analysis.
method Genetic-inspired framework, data-driven approach, Kernel Density Estimators, Copulas.
result Safety risk distribution similar to natural phenomena.
Introduces factor risk measures to assess risk relative to multiple factors.
problem Measuring risk relative to multiple factors.
method Introduces a double-argument mapping as a risk measure to assess risk relative to a vector of factors.
result Characterizes various types of factor risk measures including distortion, quantile, linear, and coherent measures.
In this paper we present an evolutionary optimization approach to solve the risk parity portfolio selection problem. While there exist convex optimization approaches to solve this problem when long-only portfolios are considered, the optimization problem becomes non-trivial in the long-short case. To solve this problem…
BoGA combines evolutionary search with Bayesian optimization for efficient protein design.
problem Designing novel proteins with specific characteristics is challenging due to sequence space complexity.
method BoGA integrates a genetic algorithm with Bayesian optimization to efficiently explore sequence space.
result BoGA accelerates discovery of high-confidence binders for diverse protein design objectives.
Semi-supervised GAN creates synthetic genetic data for disease prediction.
problem Expensive and time-consuming to build large labeled genetic databases.
method Semi-supervised Genetic Generative Adversarial Network (gGAN).
result Model achieved satisfactory results with real genetic data.
Efficiently infers graph edges from genetic similarity data in landscape genetics.
problem Inferring unknown graph edges from genetic similarity data in a heterogeneous landscape.
method Developed an efficient first-order optimization method to solve the inverse landscape genetics problem.
result Our method provides fast and reliable convergence, significantly outperforming existing heuristics.