Proposes a two-stage method for selecting correlated predictors in high-dimensional data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Bayesian networks, and especially their structures, are powerful tools for representing conditional independencies and dependencies between random variables. In applications where related variables form a priori known groups, chosen to represent different "views" to or aspects of the same entities, one may be more inte…
Using ensemble methods for regression has been a large success in obtaining high-accuracy prediction. Examples are Bagging, Random forest, Boosting, BART (Bayesian additive regression tree), and their variants. In this paper, we propose a new perspective named variable grouping to enhance the predictive performance. Th…
Transformers can learn optimal variable selection in group-sparse classification.
In genomic analysis, biomarker discovery, image recognition, and other systems involving machine learning, input variables can often be organized into different groups by their source or semantic category. Eliminating some groups of variables can expedite the process of data acquisition and avoid over-fitting. Research…
In this paper, we introduce Adaptive Cluster Lasso(ACL) method for variable selection in high dimensional sparse regression models with strongly correlated variables. To handle correlated variables, the concept of clustering or grouping variables and then pursuing model fitting is widely accepted. When the dimension is…
This article considers the problem of multi-group classification in the setting where the number of variables is larger than the number of observations . Several methods have been proposed in the literature that address this problem, however their variable selection performance is either unknown or suboptimal to…
Proposes a boosting framework for sparsity in grouped covariates.
We study a norm for structured sparsity which leads to sparse linear predictors whose supports are unions of prede ned overlapping groups of variables. We call the obtained formulation latent group Lasso, since it is based on applying the usual group Lasso penalty on a set of latent variables. A detailed analysis of th…
Factor analysis provides linear factors that describe relationships between individual variables of a data set. We extend this classical formulation into linear factors that describe relationships between groups of variables, where each group represents either a set of related variables or a data set. The model also na…
sgboost reduces variable selection bias in boosting with balanced group selection.
Bayesian method models binary response and covariates for two groups, estimating causal relationships.
Penalized regression is an attractive framework for variable selection problems. Often, variables possess a grouping structure, and the relevant selection problem is that of selecting groups, not individual variables. The group lasso has been proposed as a way of extending the ideas of the lasso to the problem of group…
New method selects variables in groups with few nonzeros, improving support recovery.
Screening is the problem of finding a superset of the set of non-zero entries in an unknown p-dimensional vector β* given n noisy observations. Naturally, we want this superset to be as small as possible. We propose a novel framework for screening, which we refer to as Multiple Grouping (MuG), that groups variables, pe…
MultiDendrograms is a Java-written application that computes agglomerative hierarchical clusterings of data. Starting from a distances (or weights) matrix, MultiDendrograms is able to calculate its dendrograms using the most common agglomerative hierarchical clustering methods. The application implements a variable-gro…
Solar improves variable selection in high-dimensional data with complicated dependence structures.
Proposes a robust optimization method for selecting grouped variables robustly.
Extends causal discovery to group variables, improving performance in real-world applications.
Complex systems may contain heterogeneous types of variables that interact in a multi-level and multi-scale manner. In this context, high-level layers may considered as groups of variables interacting in lower-level layers. This is particularly true in biology, where, for example, genes are grouped in pathways and two …
IEN speeds up T-Rex+GVS for fast, efficient GWAS.
When Daan Krammer and Stephen Bigelow independently proved that braid groups are linear, they used the Lawrence-Krammer-Bigelow representation for generic values of its variables q and t. The t variable is closely connected to the traditional Garside structure of the braid group and plays a major role in Krammer's alge…
ABM automates feature engineering and variable selection for loss-based models.
New algorithm groups variables by ancestral relationships to improve causal graph estimation accuracy.
Categorical regressor variables are usually handled by introducing a set of indicator variables, and imposing a linear constraint to ensure identifiability in the presence of an intercept, or equivalently, using one of various coding schemes. As proposed in Yuan and Lin [J. R. Statist. Soc. B, 68 (2006), 49-67], the gr…
Group model selection is the problem of determining a small subset of groups of predictors (e.g., the expression data of genes) that are responsible for majority of the variation in a response variable (e.g., the malignancy of a tumor). This paper focuses on group model selection in high-dimensional linear models, in w…
Counterfactual reasoning is an important paradigm applicable in many fields, such as healthcare, economics, and education. In this work, we propose a novel method to address the issue of \textit{selection bias}. We learn two groups of latent random variables, where one group corresponds to variables that only cause sel…
We introduce a factor analysis model that summarizes the dependencies between observed variable groups, instead of dependencies between individual variables as standard factor analysis does. A group may correspond to one view of the same set of objects, one of many data sets tied by co-occurrence, or a set of alternati…
Develops MGQDA for multi-group classification with theoretical guarantees and practical applications.
New model clusters cells and individuals, revealing genetic influences on cell types.
Clustering analysis is one of the most widely used statistical tools in many emerging areas such as microarray data analysis. For microarray and other high-dimensional data, the presence of many noise variables may mask underlying clustering structures. Hence removing noise variables via variable selection is necessary…
Develops analysis of Hölder continuous mappings on Heisenberg groups.
GTBO uses group testing to optimize high-dimensional functions efficiently.
New method uses latent variables to estimate treatment effects from single-arm trials.
This paper is devoted to the specific class of pseudoconformal mappings of quaternion and octonion variables. Normal families of functions are defined and investigated. Four criteria of a family being normal are proven. Then groups of pseudoconformal diffeomorphisms of quaternion and octonion manifolds are investigated…
We show that there is an infinite group of special automorphisms of the deformed group of diffeomorphisms, which describes parallel transports in Riemannian spaces of any variable curvature. Generators of translations of such group contain covariant derivatives, and structure functions - the curvature tensor.
Proposes a gradient-based variable selection method for binary classification in RKHS.
Fixed points found in cluster modular groups under specific conditions.
New algorithm solves complex variable selection problems in high dimensions.
We consider the empirical risk minimization problem for linear supervised learning, with regularization by structured sparsity-inducing norms. These are defined as sums of Euclidean norms on certain subsets of variables, extending the usual -norm and the group -norm by allowing the subsets to overlap. T…
VC-PCR improves prediction by clustering correlated variables.
In this paper we consider the problem of grouped variable selection in high-dimensional regression using regularization (), which can be viewed as a natural generalization of the regularization (the group Lasso). The key condition is that the dimensionality can…
We consider the problem of predicting several response variables using the same set of explanatory variables. This setting naturally induces a group structure over the coefficient matrix, in which every explanatory variable corresponds to a set of related coefficients. Most of the existing methods that utilize this gro…
Researchers derive -series for and groups.
Safe screening rule improves Group SLOPE efficiency.
The manifold hypothesis states that many kinds of high-dimensional data are concentrated near a low-dimensional manifold. If the topology of this data manifold is non-trivial, a continuous encoder network cannot embed it in a one-to-one manner without creating holes of low density in the latent space. This is at odds w…
Variable selection for models including interactions between explanatory variables often needs to obey certain hierarchical constraints. The weak or strong structural hierarchy requires that the existence of an interaction term implies at least one or both associated main effects to be present in the model. Lately, thi…
A new method clusters mixed-type data tables effectively.