Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

192384576768 · Jun 202019922001200920182026
48 results for Structured Variable Selection

VC-PCR improves prediction by clustering correlated variables.

problem Decreased prediction accuracy due to cluster structure in predictor variables.
method Supervised variable selection and clustering to integrate cluster information into a sparse modeling process.
result VC-PCR achieves better prediction, variable selection, and clustering performance.

Stability Selection improves structured variable selection but requires careful tuning.

problem Finding a right-sized model or controlling false positives in structured selection problems.
method Stability Selection applied to group lasso and structured input-output lasso.
result Stability Selection often increases power but reduces error control reliability in structured settings.

Solar improves variable selection in high-dimensional data with complicated dependence structures.

problem Variable selection in ultrahigh dimensional data with severe multicollinearity and grouping effect issues.
method Subsample-ordered least angle regression (Solar) for ultrahigh dimensional data.
result Solar yields substantial improvements in sparsity, stability, and accuracy of variable selection compared to traditional methods.

Proposes a two-stage method for selecting correlated predictors in high-dimensional data.

problem Selecting correlated predictors in high-dimensional data with unknown group structures.
method Two-stage approach: variable clustering followed by group selection.
result The two-stage method improves prediction accuracy and active predictor selection.

New methods for variable selection in nonlinear regression problems.

problem Nonlinear regression problems where target depends on inputs.
method Structured sparsity methods based on partial derivatives, reformulated as finite dimensional equivalent.
result Favourable properties of structured-sparsity models and algorithm in terms of prediction and variable selection accuracy.

The study improves the perceptron's storage capacity by optimizing variable selection.

problem Distinguishing genuine structure from random correlations in high-dimensional data.
method Replica method from statistical mechanics for optimal variable selection.
result Optimal variable selection can surpass the Cover--Gardner bound for pattern classification.

Local learning method selects covariates for causal effect estimation in the presence of latent variables.

problem Estimating causal effects from nonexperimental data with latent variables.
method Local learning approach that identifies valid adjustment sets for causal relationships.
result Ensures soundness and completeness of causal effect estimation under standard assumptions.

When applying the support vector machine (SVM) to high-dimensional classification problems, we often impose a sparse structure in the SVM to eliminate the influences of the irrelevant predictors. The lasso and other variable selection techniques have been successfully used in the SVM to perform automatic variable selec…

2007-10-02abs ↗pdf ↗

Flexible Cox model for time-dependent covariates with complex sparsity patterns.

problem Lack of flexibility in enforcing specific sparsity patterns in time-dependent Cox models.
method Proposes a flexible framework for variable selection in time-dependent Cox models, accommodating complex selection rules.
result Achieves accurate estimation with low false alarm rates for complex covariate structures.

Variable selection for models including interactions between explanatory variables often needs to obey certain hierarchical constraints. The weak or strong structural hierarchy requires that the existence of an interaction term implies at least one or both associated main effects to be present in the model. Lately, thi…

2014-11-17abs ↗pdf ↗

Solar algorithm selects variables faster and more accurately in high-dimensional data.

problem Variable selection in high-dimensional data with high accuracy and stability.
method Subsample-ordered least-angle regression (solar) and its coordinate descent generalization (solar-cd) using L0L_0 norm solution path averaging.
result Solar selects variables with high accuracy and stability, reducing redundant variable selection.

New method reduces high-dimensional data to key features.

problem Challenges of high-dimensional data analysis and interpretability.
method Randomized search to produce subspaces, ensemble of models for variable selection.
result Outperforms existing methods in prediction and variable selection.

A novel method optimizes variable-stiffness structures for better strength and weight.

problem Optimizing variable-stiffness structures for higher strength and lighter weight.
method A novel multi-stage concurrent topology optimization scheme combining DMO, S-BPTO, and CFAO.
result The method ensures better fibre angle convergence and stable optimization.

CCI algorithm handles cycles, latent variables, and selection bias in causal discovery.

problem Cycles, latent variables, and selection bias in causal processes.
method CCI algorithm using a conditional independence oracle for cyclic, latent, and selection bias cases.
result CCI outperforms existing algorithms in cyclic cases and rivals them in acyclic cases.

Sparse estimation methods are aimed at using or obtaining parsimonious representations of data or models. While naturally cast as a combinatorial optimization problem, variable or feature selection admits a convex relaxation through the regularization by the 1\ell_1-norm. In this paper, we consider situations where we…

2011-09-12abs ↗pdf ↗

This thesis studies two problems in modern statistics. First, we study selective inference, or inference for hypothesis that are chosen after looking at the data. The motiving application is inference for regression coefficients selected by the lasso. We present the Condition-on-Selection method that allows for valid s…

2015-06-30abs ↗pdf ↗

PliableBVS extends Bayesian lasso for modeling interactions with modifying variables.

problem Modeling interactions between large and small sets of variables, especially in omics studies.
method Bayesian variable selection with spike-and-slab priors and hierarchical structure.
result PliableBVS outperforms pliable lasso in identifying active main and interaction effects.

Learning Markov blanket (MB) structures has proven useful in performing feature selection, learning Bayesian networks (BNs), and discovering causal relationships. We present a formula for efficiently determining the number of MB structures given a target variable and a set of other variables. As expected, the number of…

2014-07-09abs ↗pdf ↗

Transformers can learn optimal variable selection in group-sparse classification.

problem Understanding how transformers leverage attention to select relevant variables in group-sparse classification.
method Training a one-layer transformer using gradient descent to select variables from one group of input variables.
result A one-layer transformer can correctly leverage the attention mechanism to select variables, disregarding irrelevant ones.

Study identifies key ESG variables for assessing financial risk.

problem Assessing financial risk from ESG data with many variables.
method Proposed framework for hierarchical ESG data, selecting relevant variables.
result Selected ESG variables are more relevant to financial risk than aggregated scores.

Regularised canonical correlation analysis was recently extended to more than two sets of variables by the multiblock method Regularised generalised canonical correlation analysis (RGCCA). Further, Sparse GCCA (SGCCA) was proposed to address the issue of variable selection. However, for technical reasons, the variable …

2016-10-29abs ↗pdf ↗

The paper improves multi-task learning by selecting variables and grouping tasks.

problem Improving generalization performance in multi-task learning.
method Factorizes a coefficient matrix into two matrices with sparsity for variable selection and overlapping group structure among tasks. Minimized using alternating optimization methods.
result Validated the effectiveness of the method on both synthetic and real-world datasets.

New deep learning model interprets tabular data with variable selection and explainability.

problem Deep learning models lack interpretability and variable selection.
method Proposes a new network architecture that combines deep learning with generalized linear models.
result The model provides superior predictive power and interpretable results.

SPPCSO addresses multicollinearity in high-dimensional data, improving model stability and predictive accuracy.

problem Multicollinearity in high-dimensional data leads to unstable estimation and reduced predictive accuracy.
method SPPCSO integrates principal component regression and L1 regularization to adaptively adjust shrinkage factors.
result SPPCSO achieves stable and reliable estimation in high-noise settings, distinguishing signal variables from noise.

New algorithm selects relevant variables in high-dimensional graphical models.

problem Automatic selection of relevant variables in high-dimensional graphical models.
method Extends Chow and Liu's algorithm using mutual information and entropy coefficient of determination.
result Outperforms existing methods in selecting variables with explanatory power.

Unified framework for variable selection in model-based clustering with missing data.

problem Challenges in identifying relevant variables and handling missing data in model-based clustering.
method Unified framework incorporating a data-driven penalty matrix and a mechanism for missingness modeling.
result Achieves both asymptotic consistency and selection consistency in the presence of missing data.

Bayesian framework selects features and lags for time series forecasting.

problem Variable selection and lagged error term identification in time series models.
method Hierarchical Bayesian models with spike-and-slab priors, two-stage MCMC algorithm.
result Posterior selection consistency under mild conditions, improved predictive performance.

Paper distinguishes causal structures under latent confounding and selection bias.

problem Distinguishing causal relationships when latent variables and selection bias are present.
method Formulated selected-marginalized directed graphs (smDGs) to distinguish causal structures.
result Two causal structures are indistinguishable if they have the same selected-marginalized directed graph.

Energy trees handle complex data structures with multiple variable types.

problem Handling intricate data structures with various types of covariates.
method Energy trees, a regression and classification model, use energy statistics to accommodate structured covariates of different types.
result Energy trees maintain statistical foundations, interpretability, and robustness to overfitting.

Deep P-Spline automates DNN structure selection for complex regression problems.

problem Challenges in selecting optimal network structures for DNNs.
method Linking neuron selection to knot placement in basis expansion techniques, introducing a difference penalty for automated knot selection.
result Deep P-Spline extends model class and forms a latent variable modeling framework with theoretical guarantees.

Feature selection improves learning from mixtures of discrete variables.

problem Learning mixtures of discrete random variables, especially in unreliable crowdsourcing.
method Algorithm based on mutual information to rank workers and induce a low-order statistical model.
result Improvement in real data sets can be substantial.

A hybrid model for Bayesian optimization handles mixed variables using MCTS for categorical and GP for continuous.

problem Optimizing functions with mixed variable types (continuous, integer, categorical).
method Merges MCTS for categorical and GP for continuous variables, integrates UCTS search strategy, and dynamically selects kernels.
result Hybrid models outperform traditional methods in Bayesian optimization.

Stable specification search now handles latent variables.

problem Discovering causal relationships between latent variables.
method Extended S3C to S3C-Latent, combining stability selection and multi-objective optimization.
result S3C-Latent outperformed PC-MIMBuild on simulated and real-world data.

FREEtree improves tree-based methods for correlated longitudinal data.

problem Poor performance of Random Forests in high dimensional longitudinal data with correlated features.
method FREEtree uses a piecewise random effects model and clustering with WGCNA to select features and maintain interpretability.
result FREEtree outperforms other tree-based methods in prediction and feature selection accuracy.

New methods for selecting variables in complex biomedical data.

problem Selecting important variables in multivariate, functional, and complex biomedical data.
method Optimization-based variable selection methods for various regression models.
result Outperforms state-of-the-art methods in accuracy and speed.

The paper uses graph learning to detect valid instruments in high-dimensional data for house pricing.

problem Endogeneity bias and invalid instrument validation in high-dimensional data.
method Merge variable selection algorithms and probabilistic graphs to estimate house prices and causal structure.
result Efficient data-driven instrument selection and invalid instrument purge in high-dimensional data.

This paper studies simultaneous feature selection and extraction in supervised and unsupervised learning. We propose and investigate selective reduced rank regression for constructing optimal explanatory factors from a parsimonious subset of input features. The proposed estimators enjoy sharp oracle inequalities, and w…

2014-03-25abs ↗pdf ↗