Regularized M-estimators are used in diverse areas of science and engineering to fit high-dimensional models with some low-dimensional structure. Usually the low-dimensional structure is encoded by the presence of the (unknown) parameters in some low-dimensional model subspace. In such settings, it is desirable for est…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper proves conditions for estimating precision matrices with Laplacian constraints.
Proposes a neural network method to improve consistencies in high dimensional data analysis.
Consistent selection of predictors in high-dimensional binary models with misspecified parameters.
The paper studies multi-curve interest rate models and their consistency and finite-dimensional realizations.
Lasso proves consistent model selection for high-dimensional Ising models.
The paper explores trading off consistency and dimensionality in convex surrogates for multiclass classification.
Principal binets generalize curvature line surfaces to square lattices and are a discrete integrable system.
Study shows -NN regressor consistency in complex survey designs.
Consistency of k-NN rule proven in sigma-finite dimensional metric spaces.
We show that minimum-norm interpolation in the Reproducing Kernel Hilbert Space corresponding to the Laplace kernel is not consistent if input dimension is constant. The lower bound holds for any choice of kernel bandwidth, even if selected based on data. The result supports the empirical observation that minimum-norm …
We show that the two-stage adaptive Lasso procedure (Zou, 2006) is consistent for high-dimensional model selection in linear and Gaussian graphical models. Our conditions for consistency cover more general situations than those accomplished in previous work: we prove that restricted eigenvalue conditions (Bickel et al.…
Paper proposes a new sparsity scheme for high-dimensional VAR models.
Study reveals how high-dimensional models are vulnerable to consistent adversarial attacks.
New theorem for generalized group sparsity improves consistency and convergence rates.
We introduce a new discrete system that arises from ellipsoidal billiards and is closely related to the double reflection nets. The system is defined on the lattice of a uniform honeycomb consisting of rectified hypercubes and cross polytopes. In the -dimensional case, the lattice is regular and it incorporates dyna…
Paper proposes robust LAD estimators for 2D sinusoidal model, proving consistency and normality.
High-dimensional regression models struggle with resampling methods.
In this paper, we consider the exact triangles consisting of stable vector bundles on one-dimensional complex tori, and give a geometric interpretation of them in terms of the corresponding Fukaya category via the homological mirror symmetry.
This manuscript studies statistical properties of linear classifiers obtained through minimization of an unregularized convex risk over a finite sample. Although the results are explicitly finite-dimensional, inputs may be passed through feature maps; in this way, in addition to treating the consistency of logistic reg…
We consider the problem of extracting a low-dimensional, linear latent variable structure from high-dimensional random variables. Specifically, we show that under mild conditions and when this structure manifests itself as a linear space that spans the conditional means, it is possible to consistently recover the struc…
ARGEN method improves variable selection and regularization in high-dimensional sparse models.
The paper classifies Morse functions on 3-manifolds with specific level sets.
We consider the least-square linear regression problem with regularization by the -norm, a problem usually referred to as the Lasso. In this paper, we first present a detailed asymptotic analysis of model consistency of the Lasso in low-dimensional settings. For various decays of the regularization parameter, w…
New algorithm finds best subset in high-dimensional data models.
Affine manifolds linked to integrable equations and geometric structures.
Under proportional transaction costs, a price process is said to have a consistent price system, if there is a semimartingale with an equivalent martingale measure that evolves within the bid-ask spread. We show that a continuous, multi-asset price process has a consistent price system, under arbitrarily small proporti…
Variable screening is a fast dimension reduction technique for assisting high dimensional feature selection. As a preselection method, it selects a moderate size subset of candidate variables for further refining via feature selection to produce the final model. The performance of variable screening depends on both com…
The paper optimizes hyperplanes for binary classification in high-dimensional data with latent Gaussian mixtures.
Spectral density matrix estimation of multivariate time series is a classical problem in time series and signal processing. In modern neuroscience, spectral density based metrics are commonly used for analyzing functional connectivity among brain regions. In this paper, we develop a non-asymptotic theory for regularize…
Bayesian model infers factor dimensionality and sparse loading matrix adaptively.
The study defines divergence for multivector fields on infinite-dimensional manifolds.
In this work, we develop a novel principal component analysis (PCA) for semimartingales by introducing a suitable spectral analysis for the quadratic variation operator. Motivated by high-dimensional complex systems typically found in interest rate markets, we investigate correlation in high-dimensional high-frequency …
Estimates CATEs using high-dimensional linear regression models.
Learning rule consistency tied to non-existence of real-valued measurable cardinals.
Maximum Variance Unfolding is one of the main methods for (nonlinear) dimensionality reduction. We study its large sample limit, providing specific rates of convergence under standard assumptions. We find that it is consistent when the underlying submanifold is isometric to a convex subset, and we provide some simple e…
In this project we further investigate the idea of reducing the dimensionality of datasets using a Borel isomorphism with the purpose of subsequently applying supervised learning algorithms, as originally suggested by my supervisor V. Pestov (in 2011 Dagstuhl preprint). Any consistent learning algorithm, for example kN…
Analysis of cross-validation for early-stopped gradient descent in high-dimensional regression.
There is an increasing body of evidence suggesting that exact nearest neighbour search in high-dimensional spaces is affected by the curse of dimensionality at a fundamental level. Does it necessarily mean that the same is true for k nearest neighbours based learning algorithms such as the k-NN classifier? We analyse t…
The problem of learning forest-structured discrete graphical models from i.i.d. samples is considered. An algorithm based on pruning of the Chow-Liu tree through adaptive thresholding is proposed. It is shown that this algorithm is both structurally consistent and risk consistent and the error probability of structure …
Model selection is crucial to high-dimensional learning and inference for contemporary big data applications in pinpointing the best set of covariates among a sequence of candidate interpretable models. Most existing work assumes implicitly that the models are correctly specified or have fixed dimensionality. Yet both …
We consider three-dimensional Lorentzian metrics that locally admit four independent Killing vectors. Their classification is summarized, and conditions for characterizing them are found. These consist of algebraic classification of the traceless Ricci tensor, and other conditions satisfied by the curvature and its der…
The paper examines the consistency of Lasso regression applied to signature analysis of time series data.
CMPE improves SBI efficiency and accuracy.
Improved neural network surrogates for ICF using manifold and cycle consistency.
Determining how to appropriately select the tuning parameter is essential in penalized likelihood methods for high-dimensional data analysis. We examine this problem in the setting of penalized likelihood methods for generalized linear models, where the dimensionality of covariates p is allowed to increase exponentiall…
Random Tessellation Process improves multi-dimensional data analysis.
This paper improves adversarial robustness of deep learning models.