New algorithm finds best subset in high-dimensional data models.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New suboptimal algorithm for best subset selection in high-dimensional data.
A fast algorithm selects best subsets in high-dimensional models.
Efficient algorithm solves best subset selection problem.
abess efficiently solves various machine learning problems quickly.
BWS selects best window subsets for efficient data pruning.
Optimizes subset selection in sparse learning problems.
We study a seemingly unexpected and relatively less understood overfitting aspect of a fundamental tool in sparse linear modeling - best subset selection, which minimizes the residual sum of squares subject to a constraint on the number of nonzero coefficients. While the best subset selection procedure is often perceiv…
Researchers expand on best subset selection theory, identifying key complexities.
EBBS integrates expert assessments into MIO best-subsets problem.
Identifies the best-performing algorithm from a set of candidates.
Proposes a group-splicing algorithm for efficient BSGS in high-dimensional settings.
Dimensionality reduction is a first step of many machine learning pipelines. Two popular approaches are principal component analysis, which projects onto a small number of well chosen but non-interpretable directions, and feature selection, which selects a small number of the original features. Feature selection can be…
This paper defines a generalized column subset selection problem which is concerned with the selection of a few columns from a source matrix A that best approximate the span of a target matrix B. The paper then proposes a fast greedy algorithm for solving this problem and draws connections to different problems that ca…
In this paper we discuss the variable selection method from \ell0-norm constrained regression, which is equivalent to the problem of finding the best subset of a fixed size. Our study focuses on two aspects, consistency and computation. We prove that the sparse estimator from such a method can retain all of the importa…
Novel optimization method detects change points in Gaussian data.
This paper uses Reinforcement Learning to select features from a large dataset.
Bayesian method selects subsets for LMMs with structured dependence.
New algorithm reduces high-dimensional data processing costs and achieves true sparsity.
Bayesian approach selects subsets of variables for interpretable prediction and identifies key factors in educational outcomes.
A new method for sparse linear bandits reduces exploration-exploitation tradeoff.
New MCMC algorithm reduces subset selection passes to 2 for optimal -dimensional subspace approximation.
Algorithm selects variables and bandwidths for geographically weighted regression.
A DP method selects best sparse models in high dimensions efficiently.
We consider the problem of matrix column subset selection, which selects a subset of columns from an input matrix such that the input can be well approximated by the span of the selected columns. Column subset selection has been applied to numerous real-world data applications such as population genetics summarization,…
In this article, we propose a new algorithm for supervised learning methods, by which one can both capture the non-linearity in data and also find the best subset model. To produce an enhanced subset of the original variables, an ideal selection method should have the potential of adding a supplementary level of regres…
This paper uses MIO to select features for kernel SVM classification.
Interactive RL and DT feedback improve feature selection efficiency.
In life sciences, the experts generally use empirical knowledge to recode variables, choose interactions and perform selection by classical approach. The aim of this work is to perform automatic learning algorithm for variables selection which can lead to know if experts can be help in they decision or simply replaced …
Transportation modes prediction is a fundamental task for decision making in smart cities and traffic management systems. Traffic policies designed based on trajectory mining can save money and time for authorities and the public. It may reduce the fuel consumption and commute time and moreover, may provide more pleasa…
In the last twenty-five years (1990-2014), algorithmic advances in integer optimization combined with hardware improvements have resulted in an astonishing 200 billion factor speedup in solving Mixed Integer Optimization (MIO) problems. We present a MIO approach for solving the classical best subset selection problem o…
We consider the problem of identifying any out of the best arms in an -armed stochastic multi-armed bandit. Framed in the PAC setting, this particular problem generalises both the problem of `best subset selection' and that of selecting `one out of the best m' arms [arcsk 2017]. In applications such as crowd…
The paper introduces a framework to select efficient datasets for preserving model rankings.
AutoFS combines trainers to improve feature selection efficiency and effectiveness.
Lasso and other regularization procedures are attractive methods for variable selection, subject to a proper choice of shrinkage parameter. Given a set of potential subsets produced by a regularization algorithm, a consistent model selection criterion is proposed to select the best one among this preselected set. The a…
We develop a simple stock selection model to explain why active equity managers tend to underperform a benchmark index. We motivate our model with the empirical observation that the best performing stocks in a broad market index often perform much better than the other stocks in the index. Randomly selecting a subset o…
BPASGM uses sparse graphical models to optimize portfolio selection.
We connect high-dimensional subset selection and submodular maximization. Our results extend the work of Das and Kempe (2011) from the setting of linear regression to arbitrary objective functions. For greedy feature selection, this connection allows us to obtain strong multiplicative performance bounds on several meth…
We study the problem of selecting a subset of k random variables from a large set, in order to obtain the best linear prediction of another variable of interest. This problem can be viewed in the context of both feature selection and sparse approximation. We analyze the performance of widely used greedy heuristics, usi…
We study the column subset selection problem with respect to the entrywise -norm loss. It is known that in the worst case, to obtain a good rank- approximation to a matrix, one needs an arbitrarily large number of columns to obtain a -approximation to the best entrywise -norm low ra…
Paper introduces DP methods for high-dimensional variable selection.
SCS identifies a range of plausible equally weighted portfolios, quantifying selection uncertainty.
Optimal design for multinomial logit models improves assortment selection efficiency.
Convolutional neural networks (CNNs) have been successfully applied to many recognition and learning tasks using a universal recipe; training a deep model on a very large dataset of supervised examples. However, this approach is rather restrictive in practice since collecting a large set of labeled images is very expen…
reval package selects best clustering solutions via stability-based validation.
We develop an efficient algorithm for low-rank approximation with improved approximation guarantees.
This paper improves volatility forecasting using dynamic subset selection in genetic programming.
The amount of information in the form of features and variables avail- able to machine learning algorithms is ever increasing. This can lead to classifiers that are prone to overfitting in high dimensions, high di- mensional models do not lend themselves to interpretable results, and the CPU and memory resources necess…