The paper is devoted to generalization of well-known Michael's Selection theorem on the case of extension dimension.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Variable selection and dimension reduction are two commonly adopted approaches for high-dimensional data analysis, but have traditionally been treated separately. Here we propose an integrated approach, called sparse gradient learning (SGL), for variable selection and dimension reduction via learning the gradients of t…
An efficient algorithm selects the correct number of latent dimensions in multidimensional probit models.
Study provides bounds for estimating intrinsic dimension using Gaussian kernels.
Introduces greedy feature selection for classifier-dependent feature ranking.
Efficient auto-tuning for DR hyperparameters with BO.
Paper proposes adaptive parameter selection for KGD algorithms.
Paper proposes an unsupervised feature selection algorithm with stability guarantees.
FedSel uses local differential privacy to protect data privacy in federated SGD.
SkMM selects data for finetuning by balancing bias and variance.
Paper improves Lasso for S&P500 index tracking with post-selection inference.
We study the problem of detecting change points (CPs) that are characterized by a subset of dimensions in a multi-dimensional sequence. A method for detecting those CPs can be formulated as a two-stage method: one for selecting relevant dimensions, and another for selecting CPs. It has been difficult to properly contro…
Study reduces dimensions for -means clustering for better accuracy.
Unsupervised dimension selection is an important problem that seeks to reduce dimensionality of data, while preserving the most useful characteristics. While dimensionality reduction is commonly utilized to construct low-dimensional embeddings, they produce feature spaces that are hard to interpret. Further, in applica…
Scalable feature selection improves GBDT model training speed.
Bayesian non-parametric model selects latent dimensions automatically.
New method selects relevant dimensions for better prediction in mixtures.
Semi-supervised learning improves classification in high dimensions.
High-dimensional data in many machine learning applications leads to computational and analytical complexities. Feature selection provides an effective way for solving these problems by removing irrelevant and redundant features, thus reducing model complexity and improving accuracy and generalization capability of the…
Sparse PCA selects variables with FDR control for improved performance.
Dimensionality reduction (DR) on the manifold includes effective methods which project the data from an implicit relational space onto a vectorial space. Regardless of the achievements in this area, these algorithms suffer from the lack of interpretation of the projection dimensions. Therefore, it is often difficult to…
This paper deals with a new filter algorithm for selecting the smallest subset of features carrying all the information content of a data set (i.e. for removing redundant features). It is an advanced version of the fractal dimension reduction technique, and it relies on the recently introduced Morisita estimator of Int…
We extend the Bayesian Information Criterion (BIC), an asymptotic approximation for the marginal likelihood, to Bayesian networks with hidden variables. This approximation can be used to select models given large samples of data. The standard BIC as well as our extension punishes the complexity of a model according to …
AgFlow speeds up model selection in penalized PCA.
Feature selection is an important challenge in machine learning. It plays a crucial role in the explainability of machine-driven decisions that are rapidly permeating throughout modern society. Unfortunately, the explosion in the size and dimensionality of real-world datasets poses a severe challenge to standard featur…
A new algorithm reduces regret in high-dimensional online learning problems.
In this paper, we propose a test, called Flagged-1-Bit (F1B) test, to study the intrinsic capability of recurrent neural networks in sequence learning. Four different recurrent network models are studied both analytically and experimentally using this test. Our results suggest that in general there exists a conflict be…
New algorithms for clustering and dimension reduction using relative von Neumann entropy.
New algorithms improve model selection in linear bandits with optimal regret.
Algorithm learns from both labeled and arbitrary test examples, giving guarantees for bounded VC dimension classes.
New BO method efficiently optimizes high-dimensional functions by automatically selecting variables.
This work develops rigorous theoretical basis for the fact that deep Bayesian neural network (BNN) is an effective tool for high-dimensional variable selection with rigorous uncertainty quantification. We develop new Bayesian non-parametric theorems to show that a properly configured deep BNN (1) learns the variable im…
We introduce a new approach to variable selection, called Predictive Correlation Screening, for predictor design. Predictive Correlation Screening (PCS) implements false positive control on the selected variables, is well suited to small sample sizes, and is scalable to high dimensions. We establish asymptotic bounds f…
It is now known that an extended Gaussian process model equipped with rescaling can adapt to different smoothness levels of a function valued parameter in many nonparametric Bayesian analyses, offering a posterior convergence rate that is optimal (up to logarithmic factors) for the smoothness class the true function be…
Model selection is indispensable to high-dimensional sparse modeling in selecting the best set of covariates among a sequence of candidate models. Most existing work assumes implicitly that the model is correctly specified or of fixed dimensions. Yet model misspecification and high dimensionality are common in real app…
This paper improves model selection with cross-validation using domain knowledge.
As an emerging research direction, online streaming feature selection deals with sequentially added dimensions in a feature space while the number of data instances is fixed. Online streaming feature selection provides a new, complementary algorithmic methodology to enrich online feature selection, especially targets t…
Efficient knockoffs for large-scale feature selection.
New method selects better graphs for GGM inference in small sample sizes.
Variable screening is a fast dimension reduction technique for assisting high dimensional feature selection. As a preselection method, it selects a moderate size subset of candidate variables for further refining via feature selection to produce the final model. The performance of variable screening depends on both com…
In this study, we propose an automatic learning method for variables selection based on Lasso in epidemiology context. One of the aim of this approach is to overcome the pretreatment of experts in medicine and epidemiology on collected data. These pretreatment consist in recoding some variables and to choose some inter…
Novel privatization framework for high-dimensional variable selection with differential privacy.
Dimension reduction and variable selection are performed routinely in case-control studies, but the literature on the theoretical aspects of the resulting estimates is scarce. We bring our contribution to this literature by studying estimators obtained via L1 penalized likelihood optimization. We show that the optimize…
The current study proposes a dimension reduction method, stepwise support vector machine (SVM), to reduce the dimensions of large p small n datasets. The proposed method is compared with other dimension reduction methods, namely, the Pearson product difference correlation coefficient (PCCs), recursive feature eliminati…
Variationality of conformal geodesics fails in higher dimensions.
New method corrects Laplace/BIC errors in singular models, revealing effective dimension.
This work introduces a protocol to automatically select the correct range of scales for meaningful Intrinsic Dimension estimation.
Novel unsupervised feature selection method using multi-step Markov transition probability.