Method selects features via kernel-based independence measures.
problem Feature selection in high-dimensional data.
method Optimization of conditional covariance trace.
result Method outperforms other feature selection algorithms.
We introduce extensions of stability selection, a method to stabilise variable selection methods introduced by Meinshausen and Bühlmann (J R Stat Soc 72:417-473, 2010). We propose to apply a base selection method repeatedly to random observation subsamples and covariate subsets under scrutiny, and to select covariates …
Improves personalized treatment selection using covariates.
problem Ranking and selecting the best alternative based on covariates.
method Linear model for covariate effects, two-stage procedures for error types, generalized slippage configuration.
result Procedures provide statistical guarantees for correct selection.
Proposes FarmHazard model for hazard regression with correlated covariates.
problem Model selection challenges in high-dimensional data with correlated covariates.
method Factor-Augmented Regularized Model for Hazard Regression (FarmHazard) that learns latent factors and idiosyncratic components.
result Proves model selection and estimation consistency under mild conditions.
Local learning method selects covariates for causal effect estimation in the presence of latent variables.
problem Estimating causal effects from nonexperimental data with latent variables.
method Local learning approach that identifies valid adjustment sets for causal relationships.
result Ensures soundness and completeness of causal effect estimation under standard assumptions.
DARTS optimizes covariate selection in trials with limited data.
problem Limited budget for high-dimensional pretreatment data.
method Dynamic Adaptive Rerandomization via Thompson Sampling (DARTS).
result DARTS efficiently concentrates budget on informative features.
Evaluating prediction models under covariate shift and selective labels
problem Model performance evaluation under distribution shift and selection bias
method Double machine learning
result Accurate estimation of target risk
Selective inference for group lasso estimators across various distributions and covariates.
problem Developing selective inference methods for group lasso estimators.
method Randomized group-regularized optimization problem with post-selection likelihood.
result Selective point estimator and Wald-type confidence regions for regression parameters.
DRCS selects a subset of data to minimize worst-case test error under covariate shift.
problem Selecting a subset of data that performs well across different deployment scenarios when data distributions differ.
method DRCS derives an upper bound for the worst-case test error assuming covariate shift and selects instances to minimize this bound.
result DRCS achieves distributionally robust training instance selection.
Study tackles variable selection with missing covariates and outcomes using machine learning and imputation.
problem Missing data in both covariates and outcomes complicates variable selection in health studies.
method Exploits machine learning flexibility and bootstrap imputation for variable selection, comparing multiple methods.
result XGBoost and BART perform best in variable selection with bootstrap imputation, achieving high F1 scores and low Type I errors. A new method selects covariates for causal effect estimation without strong assumptions.
problem Estimating causal effects without global causal structure learning and strong assumptions.
method Local covariate selection method that avoids pretreatment and causal sufficiency assumptions.
result The method achieves accurate causal effect estimation with improved computational efficiency.
A scalable algorithm for GP regression selects relevant covariates efficiently.
problem Scalable variable selection in large GP regression models.
method VGPR algorithm using Vecchia approximation for sparse precision matrix, mini-batch subsampling.
result Improved scalability and accuracy in selecting relevant covariates.
MEBoost selects variables in regression with measured error, improving accuracy over naive methods.
problem Variable selection in regression models with covariates measured with error.
method Iterative algorithm that corrects for measurement error using estimating equations.
result MEBoost outperforms naive methods in selecting accurate covariates, especially under high measurement error.
PS framework selects best policy from library for CSO problems.
problem Policy selection in CSO with heterogeneous performance across covariate space.
method PS framework constructs library of candidate policies and learns a meta-policy to select the best one.
result PS consistently outperforms best single policy in heterogeneous CSO problems.
Flexible Cox model for time-dependent covariates with complex sparsity patterns.
problem Lack of flexibility in enforcing specific sparsity patterns in time-dependent Cox models.
method Proposes a flexible framework for variable selection in time-dependent Cox models, accommodating complex selection rules.
result Achieves accurate estimation with low false alarm rates for complex covariate structures.
CSL selects best model from library based on covariates.
problem Selecting the best model from a library based on covariates.
method Uses meta learning and cross-validation to find a local minimum.
result Converges at a rate faster than Op(n−1/4) and offers extensive empirical evidence. Bayesian method selects important covariates in modal regression.
problem Bayesian modal regression with heavy-tailed responses.
method Expectation-maximization algorithm for parameter estimation; test statistic for variable selection.
result Efficacy of the proposed method in identifying important covariates.
A new method clusters covariates considering class labels for better classification.
problem Clustering covariates independently of class labels can lead to poor results.
method Formulates as convex optimization, uses ADMM for solving, and selects model via marginal likelihood.
result Proposed method offers a unique global minimum and improves classification.
Novel method constructs covariance functions for Bayesian optimisation.
problem Bayesian optimisation with existing knowledge.
method Uses m-kernels to convert existing covariance functions to problem-specific ones.
result Constructs covariance functions matching the problem at hand.
New method splits unknown covariance Gaussians into independent parts.
problem Splitting multivariate Gaussian data with unknown covariance.
method Developed a general algorithm for decomposing unknown covariance Gaussians.
result Demonstrated decomposition for single multivariate Gaussian with unknown covariance.
Paper improves prediction accuracy in sparse linear models with missing data.
problem Enhancing prediction accuracy in sparse linear models with missing information.
method Introduces an approach combining sparse regression and covariance matrix estimation.
result Improves matrix completion accuracy and feature selection precision, reducing prediction MSE.
Paper compares feature selection methods using GCM and LOCO, showing GCM methods generally outperform LOCO.
problem Feature selection and importance estimation in model-agnostic settings.
method Comparison of feature selection methods related to GCM and LOCO under three model settings.
result GCM-related methods generally outperform LOCO under suitable regularity conditions, as shown by theoretical and empirical results.
Proposes an L1-regularized functional SVM for binary classification with functional covariates.
problem Binary classification with multivariate functional covariates.
method L1-regularized functional support vector machine (SVM) with an accompanying algorithm.
result The proposed classifier performs well in prediction and feature selection.
New covariance estimator for financial portfolios.
problem Estimating large financial covariances in non-stationary environments.
method Exponentially weighted averages and cross-validation for nonlinearly shrinking sample eigenvalues.
result Our estimator performs well in large dimensions compared to existing estimators.
Optimal selective classification using likelihood ratios improves model reliability.
problem Enhancing predictive model reliability by allowing uncertain predictions.
method Neyman--Pearson lemma applied to likelihood ratios for optimal selection.
result Neyman--Pearson-informed methods outperform existing baselines under covariate shifts.
Training logs can improve model comparison precision, but careful covariate selection is key.
problem Improving precision in comparing stochastically trained models.
method Use arm-specific covariate adjustment, where each model is adjusted with statistics from its own runs.
result Simple adjustments based on early training logs often reduce uncertainty in model comparisons.
The study evaluates different parameter selection methods for Gaussian process interpolation.
problem Choosing optimal parameters for Gaussian process interpolation.
method Empirical study using scoring rules and leave-one-out selection criteria.
result The choice of model family is often more important than the selection criterion.
The paper addresses the selection of synthetic data for improving classifier performance, focusing on the role of covariance shift.
problem The effectiveness of synthetic data in improving classifier performance is questioned, and the specific properties affecting this performance are unclear.
method The paper uses high-dimensional regression to analyze synthetic data selection, focusing on the covariance shift between synthetic and target distributions.
result The covariance shift between synthetic and target distributions affects the generalization error of classifiers, but the mean shift does not.
Geodesic curves improve flexibility in covariance estimation.
problem Inflexible covariance families limit spatiotemporal modeling.
method Use geodesic curves to build more flexible covariance families.
result Natural projection minimizes geodesic distance to sample covariance.
Proposes spBART for risk prediction using epigenetic signatures and covariates.
problem Complex high-dimensional epigenetic data and low-dimensional covariates for risk prediction.
method Semi-parametric Bayesian Additive Regression Trees (spBART) with cross-validation for variable selection.
result Achieves strong out-of-sample discrimination (AUC = 0.96) in held-out validation set.
Paper solves portfolio selection under uncertain covariance matrix using robust optimization.
problem Optimizing portfolio selection under model uncertainty in covariance matrix.
method Formulates as a min-max mean-variance problem, solves using McKean-Vlasov dynamic programming.
result Provides explicit solutions for optimal robust portfolio strategies and robust efficient frontier.
Extends deep learning for nonlinear Cox regression variable selection.
problem Variable selection for nonlinear Cox regression model.
method Extends LassoNet to survival data for nonlinear Cox model.
result Valid and effective method demonstrated through simulations.
Robust Conformalized Selection controls FDR under noisy responses.
problem Existing conformal selection methods fail to control FDR under contaminated calibration data.
method RCS framework for selective classification with valid FDR control under label contamination.
result RCS framework controls FDR and maintains power under contaminated calibration data.
A method to simplify Gaussian graphical models for high-dimensional data.
problem Difficulty in inferring networks of dependencies between variables when sample size is small.
method Approximate covariance matrix as block-diagonal, select threshold using slope heuristic, infer network in blocks.
result The method reduces the number of parameters to estimate and improves network inference.
Safe-DRFS selects features robust to covariate shifts for reliable performance.
problem Feature selection fails in diverse deployment environments.
method Safe-DRFS extends safe screening to distributionally robust settings under covariate shift.
result Safe-DRFS identifies a feature subset encompassing optimal subsets across distribution shifts.
Kernel ridge regression for causal inference with missing data.
problem Estimating treatment effects with missing data in selected samples.
method Kernel ridge regression estimators for nonparametric dose response curves and semiparametric treatment effects.
result Uniform consistency and finite sample rates for continuous treatment, root-n consistency for discrete treatment.
This work proposes a model averaging method for SVM that avoids redundant covariates and achieves asymptotic optimality.
problem Redundant covariates impair SVM performance in high-dimensional settings.
method Frequentist model averaging procedure for SVM using cross-validation to select optimal weights.
result The proposed method achieves asymptotic optimality in SVM model averaging.
Sharp rates for prediction error in high-dimensional sparse models.
problem High-dimensional sparse linear models with limited predictive power.
method Forward regression for model selection and least squares estimation.
result Sharp convergence rates without beta-min or irrepresentability conditions.
Develops a method for estimating networks and covariate associations in compositional data.
problem Estimating network interactions and covariate associations for compositional data.
method Hierarchical Bayesian model with spike-and-slab priors for edge and covariate selection, variational EM for inference.
result The proposed method outperforms existing methods in network recovery accuracy.
Method selects significant spatial covariates in noisy data.
problem Identifying true spatial covariates in noisy data.
method Combines sparsity-promoting estimation with noise-robust model selection.
result Method reliably recovers true covariates under diverse noise scenarios.
Efficient knockoffs for large-scale feature selection.
problem Large-scale feature selection problems.
method Gaussian model-X knockoffs with efficient methods for solving semidefinite programs.
result Efficient knockoffs can be generated with linear complexity in the dimension.
A new algorithm reduces regret in high-dimensional online learning problems.
problem High-dimensional covariates with unknown reward function.
method BV-LASSO algorithm incorporating binning and voting for nonparametric variable selection.
result Achieves optimal regret ildeO(T(dx∗+dy+1)/(dx∗+dy+2)). This paper analyzes the quality of covariance selection in graphical models using AUC bounds.
problem Quality assessment of covariance selection in graphical models.
method Formulated as a detection problem, the paper uses AUC bounds to measure model quality.
result Quality of tree approximation models decays exponentially with increasing dimension.
A framework for online learning of regularization parameters in streaming data.
problem Efficiently learning regularization parameters in online learning settings.
method Stochastic gradient descent to iteratively estimate a time-varying sparsity parameter.
result Convergence results in non-stochastic settings for linear and graphical models.
New method selects variables for GP regression using sparse projection.
problem Identifying environmental factors affecting metal corrosion.
method Sparse projection of input variables, gradient descent optimization, non-convex marginal likelihood.
result Proposed method outperforms benchmarks in variable selection accuracy.
In a Gaussian graphical model, the conditional independence between two variables are characterized by the corresponding zero entries in the inverse covariance matrix. Maximum likelihood method using the smoothly clipped absolute deviation (SCAD) penalty (Fan and Li, 2001) and the adaptive LASSO penalty (Zou, 2006) hav…
Efficient Bayesian variable selection for binomial and negative binomial data.
problem Computational challenges in Bayesian variable selection for complex models.
method Tempered Gibbs Sampling and MCMC scheme.
result Demonstrated effectiveness on cancer data with thousands of covariates.
This paper explores how effective sample size, dimensionality, and model performance are related in covariate shift adaptation.
problem Understanding the relationship between effective sample size, dimensionality, and generalization in covariate shift adaptation.
method Building a unified theory connecting effective sample size, data dimensionality, and generalization in the context of covariate shift adaptation.
result Dimensionality reduction or feature selection can increase effective sample size, supporting the practice of reducing dimensionality before covariate shift adaptation.