Proposes a two-stage method for selecting correlated predictors in high-dimensional data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
sgboost reduces variable selection bias in boosting with balanced group selection.
In this paper, we introduce Adaptive Cluster Lasso(ACL) method for variable selection in high dimensional sparse regression models with strongly correlated variables. To handle correlated variables, the concept of clustering or grouping variables and then pursuing model fitting is widely accepted. When the dimension is…
Penalized regression is an attractive framework for variable selection problems. Often, variables possess a grouping structure, and the relevant selection problem is that of selecting groups, not individual variables. The group lasso has been proposed as a way of extending the ideas of the lasso to the problem of group…
Transformers can learn optimal variable selection in group-sparse classification.
Solar improves variable selection in high-dimensional data with complicated dependence structures.
Counterfactual reasoning is an important paradigm applicable in many fields, such as healthcare, economics, and education. In this work, we propose a novel method to address the issue of \textit{selection bias}. We learn two groups of latent random variables, where one group corresponds to variables that only cause sel…
ABM automates feature engineering and variable selection for loss-based models.
New method selects variables in groups with few nonzeros, improving support recovery.
New method corrects selection bias in post-selective inference for Group LASSO.
Proposes a robust optimization method for selecting grouped variables robustly.
Proposes a gradient-based variable selection method for binary classification in RKHS.
Proposes a boosting framework for sparsity in grouped covariates.
Group model selection is the problem of determining a small subset of groups of predictors (e.g., the expression data of genes) that are responsible for majority of the variation in a response variable (e.g., the malignancy of a tumor). This paper focuses on group model selection in high-dimensional linear models, in w…
Proposes a neural network framework for feature selection in high-dimensional settings.
IEN speeds up T-Rex+GVS for fast, efficient GWAS.
VC-PCR improves prediction by clustering correlated variables.
Study improves model estimation and variable selection using GANs with Lasso penalty.
This article considers the problem of multi-group classification in the setting where the number of variables is larger than the number of observations . Several methods have been proposed in the literature that address this problem, however their variable selection performance is either unknown or suboptimal to…
The paper develops a test for independence of selected Gaussian variables after thresholding correlations.
Variable selection for models including interactions between explanatory variables often needs to obey certain hierarchical constraints. The weak or strong structural hierarchy requires that the existence of an interaction term implies at least one or both associated main effects to be present in the model. Lately, thi…
Improves Group Lasso for categorical data by reducing dimensionality and selecting models.
Screening is the problem of finding a superset of the set of non-zero entries in an unknown p-dimensional vector β* given n noisy observations. Naturally, we want this superset to be as small as possible. We propose a novel framework for screening, which we refer to as Multiple Grouping (MuG), that groups variables, pe…
Develops MGQDA for multi-group classification with theoretical guarantees and practical applications.
Clustering analysis is one of the most widely used statistical tools in many emerging areas such as microarray data analysis. For microarray and other high-dimensional data, the presence of many noise variables may mask underlying clustering structures. Hence removing noise variables via variable selection is necessary…
New algorithm solves complex variable selection problems in high dimensions.
Flexible Cox model for time-dependent covariates with complex sparsity patterns.
Safe screening rule improves Group SLOPE efficiency.
SPPCSO addresses multicollinearity in high-dimensional data, improving model stability and predictive accuracy.
Method selects valid IVs from a large set using clustering and test of overidentifying restrictions.
We investigate structured sparsity methods for variable selection in regression problems where the target depends nonlinearly on the inputs. We focus on general nonlinear functions not limiting a priori the function space to additive models. We propose two new regularizers based on partial derivatives as nonlinear equi…
We consider the empirical risk minimization problem for linear supervised learning, with regularization by structured sparsity-inducing norms. These are defined as sums of Euclidean norms on certain subsets of variables, extending the usual -norm and the group -norm by allowing the subsets to overlap. T…
Feature selection aims to select the smallest feature subset that yields the minimum generalization error. In the rich literature in feature selection, information theory-based approaches seek a subset of features such that the mutual information between the selected features and the class labels is maximized. Despite …
We consider multi-task learning, which simultaneously learns related prediction tasks, to improve generalization performance. We factorize a coefficient matrix as the product of two matrices based on a low-rank assumption. These matrices have sparsities to simultaneously perform variable selection and learn and overlap…
Variable selection for high-dimensional linear models has received a lot of attention lately, mostly in the context of l1-regularization. Part of the attraction is the variable selection effect: parsimonious models are obtained, which are very suitable for interpretation. In terms of predictive power, however, these re…
Flexible co-data learning improves clinical prediction models.
In variable or graph selection problems, finding a right-sized model or controlling the number of false positives is notoriously difficult. Recently, a meta-algorithm called Stability Selection was proposed that can provide reliable finite-sample control of the number of false positives. Its benefits were demonstrated …
In machine learning, Feature Selection (FS) is a major part of efficient algorithm. It fuels the algorithm and is the starting block for our prediction. In this paper, we present a new method, called Optimal Coordinate Ascent (OCA) that allows us selecting features among block and individual features. OCA relies on coo…
We consider the problem of sparse variable selection in nonparametric additive models, with the prior knowledge of the structure among the covariates to encourage those variables within a group to be selected jointly. Previous works either study the group sparsity in the parametric setting (e.g., group lasso), or addre…
Identifying homogeneous subgroups of variables can be challenging in high dimensional data analysis with highly correlated predictors. We propose a new method called Hexagonal Operator for Regression with Shrinkage and Equality Selection, HORSES for short, that simultaneously selects positively correlated variables and…
Proposes a group-splicing algorithm for efficient BSGS in high-dimensional settings.
In this paper we consider the problem of grouped variable selection in high-dimensional regression using regularization (), which can be viewed as a natural generalization of the regularization (the group Lasso). The key condition is that the dimensionality can…
Proposes novel wSVMs for sparse learning and accurate probability estimation.
Extends causal discovery to group variables, improving performance in real-world applications.
Exclusive Lasso improves survival prediction in cancer datasets.
Genome-wide association studies (GWAS) have achieved great success in the genetic study of Alzheimer's disease (AD). Collaborative imaging genetics studies across different research institutions show the effectiveness of detecting genetic risk factors. However, the high dimensionality of GWAS data poses significant cha…
Bayesian method selects subsets for LMMs with structured dependence.
Proposes HDBEN for heteroscedastic regression with improved sparsity and variance modeling.