A new method selects variables for random survival forests using maximally selected rank statistics.
problem Random survival forests can be biased in selecting variables, especially for non-linear effects.
method Use maximally selected rank statistics for variable selection in random survival forests, comparing on p-value scale.
result The new method outperforms other approaches in prediction performance and computational speed.
Private variable selection method controls FDR with simulations showing reasonable power.
problem Performing variable selection with privacy constraints.
method Private knockoff filter using Gaussian and Laplace mechanisms.
result Achieves controlled false discovery rate (FDR) in variable selection.
Proposes a new method using Copula Entropy for variable selection.
problem Variable selection in machine learning and statistics.
method Copula Entropy (CE) based ranks for variable selection, model-free and tuning-free.
result CE based method selects variables more effectively and derives better interpretable results.
Bayesian neural network improves feature selection and prediction.
problem Improving feature selection and prediction accuracy in neural networks.
method BNN-ARD with l2-norm feature importance measure.
result Improves variable selection and predictive performance on real-world data.
Two statistical tasks are shown to have equivalent sample complexity.
problem Determining if a function depends on only a few variables and identifying those variables.
method Proved statistical equivalence of feature selection and junta testing through sample complexity analysis.
result Brute-force algorithm is sample-optimal for both tasks with optimal sample size.
Boosting combines ML with stats for flexible modeling.
problem Flexible modeling in biomedicine.
method Combines machine learning and classical stats.
result Advances in variable selection, functional regression, and time-to-event modeling.
A new knockoff statistic using conditional prediction function improves variable selection in complex models.
problem Controlling false discovery rate in complex models with nonlinear relationships.
method Introducing a knockoff statistic based on the conditional prediction function for use with machine learning models.
result The CPF statistics provide superior power in detecting prognostic variables over existing knockoff statistics.
Study suggests variable selection may not significantly reduce power in multivariate tests.
problem The feasibility of parsimonious variable selection in Hotelling's T2-test.
method Investigation of power loss when selecting small subsets of variables from multivariate data.
result Some evidence suggests no significant power loss over a wide range of alternatives.
Proposes a method to select variables for kernel two-sample tests.
problem Determining whether two samples have the same distribution using informative variables.
method A framework based on kernel maximum mean discrepancy (MMD) for selecting a subset of variables.
result The sample size requirements for the three kernels depend on the number of selected variables, not the data dimension.
Proposes a two-stage method for selecting correlated predictors in high-dimensional data.
problem Selecting correlated predictors in high-dimensional data with unknown group structures.
method Two-stage approach: variable clustering followed by group selection.
result The two-stage method improves prediction accuracy and active predictor selection.
The study improves the perceptron's storage capacity by optimizing variable selection.
problem Distinguishing genuine structure from random correlations in high-dimensional data.
method Replica method from statistical mechanics for optimal variable selection.
result Optimal variable selection can surpass the Cover--Gardner bound for pattern classification.
Extends model-x framework to handle missing data.
problem Inability to control false selections in missing data settings.
method Posterior sampled imputation, univariate imputation, joint imputation and sampling knockoffs.
result Preserves theoretical guarantees of model-x framework in missing data setting.
Bayesian method selects important covariates in modal regression.
problem Bayesian modal regression with heavy-tailed responses.
method Expectation-maximization algorithm for parameter estimation; test statistic for variable selection.
result Efficacy of the proposed method in identifying important covariates.
VarPro selects features without model dependence, achieving balanced performance.
problem Finding a small set of features with high explanatory power.
method Rule-based variable priority approach, avoiding model-specific methods and artificial data.
result VarPro has a consistent filtering property for noise variables and achieves balanced performance.
New deep learning model interprets tabular data with variable selection and explainability.
problem Deep learning models lack interpretability and variable selection.
method Proposes a new network architecture that combines deep learning with generalized linear models.
result The model provides superior predictive power and interpretable results.
A neural network approach unifies Lasso for variable selection.
problem Combining statistical and machine learning techniques for variable selection.
method Representing Lasso through a neural network and developing a new optimization algorithm.
result The new optimization algorithm achieves better performance than previous methods.
Proposes a method for stable variable selection in high-dimensional data.
problem Challenges of variable selection in high-dimensional, correlated data.
method Resample-aggregate framework using diffusion models.
result Stable subset of predictors with calibrated stability scores.
Bayesian approach controls FDR in high-dimensional models.
problem High-dimensional variable selection and inference.
method Adapted Mirror Statistic to Bayesian framework for FDR control.
result Effective FDR control without data splitting.
A fast, approximate method for variable selection in GLMs tackles correlated data.
problem Variable selection in generalized linear models with correlated data.
method Replica method of statistical mechanics and vector approximate message passing.
result The proposed algorithm provides fast convergence and high approximation accuracy.
A new variable selection method using model-based boosting and random permutations.
problem Sparse and fast variable selection in high-dimensional data.
method Model-based gradient boosting with randomly permuted variables to stop early.
result Competes with state-of-the-art methods in high-dimensional classification.
The paper introduces a valuation framework for variable selection in econometric models.
problem Optimizing variable selection in econometric models to balance gains and losses.
method Derives a valuation framework based on expected marginal gains and losses, introduces three unbiased solutions.
result New approaches significantly outperform existing methods in variable selection.
Bayesian method tackles variable selection in high-dimensional data.
problem Challenges in Bayesian variable selection with large P.
method Efficient MCMC scheme with sublinear cost per iteration, extended to generalized linear models.
result Demonstrated effectiveness on cancer and maize genomic data.
MIBoost boosts variable selection with multiple imputation.
problem Missing data complicates variable selection in statistical models.
method Gradient boosting with multiple imputation, MIBoost.
result MIBoost yields comparable predictive performance to other methods.
New cross-validation method reduces overfitting uncertainty.
problem Overfitting in cross-validation.
method Statistically principled inference tool based on cross-validation.
result Guaranteed probability of selecting the best model.
We develop the necessary theory in computational algebraic geometry to place Bayesian networks into the realm of algebraic statistics. We present an algebra{statistics dictionary focused on statistical modeling. In particular, we link the notion of effiective dimension of a Bayesian network with the notion of algebraic…
This paper explores the following question: what kind of statistical guarantees can be given when doing variable selection in high-dimensional models? In particular, we look at the error rates and power of some multi-stage regression methods. In the first stage we fit a set of candidate models. In the second stage we s…
A fast algorithm selects best subsets in high-dimensional models.
problem Identifying sparse models in high-dimensional generalized linear models.
method Splicing technique for fast and consistent best subset selection.
result Our algorithm achieves high certainty in selecting best subsets with polynomial computational complexity.
New methods for selecting variables in complex biomedical data.
problem Selecting important variables in multivariate, functional, and complex biomedical data.
method Optimization-based variable selection methods for various regression models.
result Outperforms state-of-the-art methods in accuracy and speed.
Proposes statistical inference for L2-Boosting.
problem Statistical inference for L2-Boosting. method Post-selection inference for iterative variable selection in L2-Boosting. result Developed tests and confidence intervals for L2-Boosting. CPI overcomes limitations of permutation importance by providing accurate variable selection.
problem Misidentification of unimportant variables in complex models due to covariate correlations.
method Developed a model agnostic and computationally lean Conditional Permutation Importance (CPI) approach.
result CPI provides accurate type-I error control and more parsimonious variable selection.
Proposes a modified Morgan-Pitman test for evaluating variances in machine learning models.
problem Limited ability to account for sampling variability in model selection.
method Enhances the classic Morgan-Pitman test for robustness in non-linear models with heavy-tailed distributions or outliers.
result Demonstrates the test's effectiveness and practical utility in model evaluation and selection.
New insights into variable selection with different model assumptions.
problem Sparse recovery with ℓ∞ error guarantees in variable selection. method Separation between oblivious and adaptive models of ℓ∞ sparse recovery. result Proves a surprising contrast between oblivious and adaptive models in ℓ∞ sparse recovery. We study the computational complexity of Markov chain Monte Carlo (MCMC) methods for high-dimensional Bayesian linear regression under sparsity constraints. We first show that a Bayesian approach can achieve variable-selection consistency under relatively mild conditions on the design matrix. We then demonstrate that t…
Proposes a criterion for selecting relevant auxiliary variables in incomplete data analysis.
problem Selecting useful auxiliary variables for incomplete data analysis.
method Formulates model selection problem, proposes an information criterion based on Kullback-Leibler divergence.
result Proposed information criterion is an asymptotically unbiased estimator of Kullback-Leibler divergence.
Extends knockoff filter for composite null hypotheses in variable selection.
problem Handling composite null hypotheses in variable selection.
method Developed two methods for composite inference with knockoffs: S-OLS and FRPP.
result Proposed heuristic variants of S-OLS outperforming BH procedure for composite nulls.
New algorithm solves complex variable selection problems in high dimensions.
problem Grouped variable selection in high-dimensional data.
method Optimal solutions for the ℓ0-regularized formulation using discrete optimization.
result Exact algorithms solve problems with 5 million features and 1000 observations in minutes to hours.
Study examines inference methods after variable selection in Cox models.
problem Bias and misleading inference after variable selection in Cox models.
method Simulation study of inference procedures for Lasso and adaptive Lasso in Cox models.
result Performance of inference procedures varies, with debiased Lasso showing promise.
Modern biotechnologies often result in high-dimensional data sets with much more variables than observations (n ≪ p). These data sets pose new challenges to statistical analysis: Variable selection becomes one of the most important tasks in this setting. We assess the recently proposed flexible framework for variab…
Transformers can learn optimal variable selection in group-sparse classification.
problem Understanding how transformers leverage attention to select relevant variables in group-sparse classification.
method Training a one-layer transformer using gradient descent to select variables from one group of input variables.
result A one-layer transformer can correctly leverage the attention mechanism to select variables, disregarding irrelevant ones.
We study high-dimensional Gaussian mixture classification using statistical physics methods.
problem Classifying high-dimensional Gaussian mixture with general covariance matrices.
method Replica method from statistical physics for asymptotic analysis of convex classifiers.
result Construction and validation of a de-biased estimator for variable selection.
Enhances FDR control in variable selection using neural networks.
problem Balancing rigorous error control with statistical power in high-dimensional variable selection.
method Learning-augmented T-Rex Selector framework with a neural network trained on synthetic datasets.
result Achieves superior detection of true variables compared to existing approaches.
Penalized regression is an attractive framework for variable selection problems. Often, variables possess a grouping structure, and the relevant selection problem is that of selecting groups, not individual variables. The group lasso has been proposed as a way of extending the ideas of the lasso to the problem of group…
Method identifies causal drivers from background features.
problem Distinguishing causal influence from hidden confounding.
method Stability of regression coefficients measured by statistic V.
result V converges to zero if and only if no causal drivers exist.
We study the problem of variable selection in convex nonparametric regression. Under the assumption that the true regression function is convex and sparse, we develop a screening procedure to select a subset of variables that contains the relevant variables. Our approach is a two-stage quadratic programming method that…
New algorithm improves latent variable model estimation.
problem Estimating parameters in latent variable models.
method Jarzynski-adjusted Langevin algorithm (JALA) for SMC methods.
result JALA-EM provides maximum marginal likelihood estimate.
Proposes two-stage robust and sparse distributed inference for large-scale data.
problem Statistical inference in large-scale, high-dimensional, and outlier-contaminated data.
method Two-stage approach: model selection with robust Lasso, fusion of local selections, and bootstrap methods for inference.
result Robust and computationally efficient inference procedures for variable selection, confidence intervals, and standard deviation approximations.
A/B testing improves marketing decisions by selecting effective stratification variables.
problem Improving the sensitivity of A/B testing through stratified sampling.
method Designing an algorithm to select a subset of stratification variables for variance reduction.
result The subset selection method outperforms other variance reduction techniques in A/B testing.
Bayesian approach selects subsets of variables for interpretable prediction and identifies key factors in educational outcomes.
problem Challenges in subset selection for stability, regularization, and inference.
method Bayesian perspective on subset selection, deriving optimal subsets and variable importance metrics.
result Better prediction, interval estimation, and variable selection compared to competing methods.