Study shows exponential convergence in classification errors using random features and SGD.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We consider stochastic gradient descent and its averaging variant for binary classification problems in a reproducing kernel Hilbert space. In the traditional analysis using a consistency property of loss functions, it is known that the expected classification error converges more slowly than the expected risk even whe…
Paper improves privacy in SGD with low noise, achieving optimal risk rates.
This paper studies the classification of high-dimensional Gaussian signals from low-dimensional noisy, linear measurements. In particular, it provides upper bounds (sufficient conditions) on the number of measurements required to drive the probability of misclassification to zero in the low-noise regime, both for rando…
Improved SGD with AdaGrad stepsizes adapts to unknown parameters and unbounded gradients.
New algorithms learn from comparisons to classify data robustly to noise.
New method reduces memorization in diffusion models without sacrificing image quality.
This paper considers the classification of linear subspaces with mismatched classifiers. In particular, we assume a model where one observes signals in the presence of isotropic Gaussian noise and the distribution of the signals conditioned on a given class is Gaussian with a zero mean and a low-rank covariance matrix.…
Study generalizes matrix completion with side info in low noise settings.
We consider binary classification problems with positive definite kernels and square loss, and study the convergence rates of stochastic gradient methods. We show that while the excess testing loss (squared loss) converges slowly to zero as the number of observations (and thus iterations) goes to infinity, the testing …
Density mode clustering is a nonparametric clustering method. The clusters are the basins of attraction of the modes of a density estimator. We study the risk of mode-based clustering. We show that the clustering risk over the cluster cores --- the regions where the density is high --- is very small even in high dimens…
Improves score estimation for noised targets using known clean scores.
We consider two variables that are related to each other by an invertible function. While it has previously been shown that the dependence structure of the noise can provide hints to determine which of the two variables is the cause, we presently show that even in the deterministic (noise-free) case, there are asymmetr…
New algorithm reduces feature count and accelerates error convergence.
This work shows how transformers use multi-concept word semantics for efficient in-context learning.
Paper provides label complexity guarantees for deep active learning.
This paper tackles deferral learning with multiple experts, providing strong theoretical guarantees.
The problem of devising learning strategies for discrete losses (e.g., multilabeling, ranking) is currently addressed with methods and theoretical analyses ad-hoc for each loss. In this paper we study a least-squares framework to systematically design learning algorithms for discrete losses, with quantitative character…
New findings show privacy affects generalization error in a non-monotonic way.
We consider high-dimensional binary classification by sparse logistic regression. We propose a model/feature selection procedure based on penalized maximum likelihood with a complexity penalty on the model size and derive the non-asymptotic bounds for the resulting misclassification excess risk. The bounds can be reduc…
Enhanced consistency bounds derived for classification under a new noise condition.
The successive projection algorithm (SPA) is a fast algorithm to tackle separable nonnegative matrix factorization (NMF). Given a nonnegative data matrix , SPA identifies an index set such that there exists a nonnegative matrix with . SPA has been successfully used as a…
Let $\cF$ be a set of classification procedures with values in . Given a loss function, we want to construct a procedure which mimics at the best possible rate the best procedure in $\cF$. This fastest rate is called optimal rate of aggregation. Considering a continuous scale of loss functions with various …
Bayesian regression underestimates parameter uncertainties in noisy models.
A new method reduces sample complexity for meta-learning.
The paper tackles inverse uncertainty quantification in neutron noise analysis.
We present a simple noise-robust margin-based active learning algorithm to find homogeneous (passing the origin) linear separators and analyze its error convergence when labels are corrupted by noise. We show that when the imposed noise satisfies the Tsybakov low noise condition (Mammen, Tsybakov, and others 1999; Tsyb…
A method merges two pretrained diffusion experts to improve image quality and likelihood.
In this paper, we propose a provably correct algorithm for convolutive nonnegative matrix factorization (CNMF) under separability assumptions. CNMF is a convolutive variant of nonnegative matrix factorization (NMF), which functions as an NMF with additional sequential structure. This model is useful in a number of appl…
Optimizes binary regression models with gradient ascent-descent methods.
This paper proposes a deep speech enhancement method which exploits the high potential of residual connections in a wide neural network architecture, a topology known as Wide Residual Network. This is supported on single dimensional convolutions computed alongside the time domain, which is a powerful approach to proces…
We provide new conditions for the Strong Atiyah conjecture to lift to finite group extensions. In particular, we show cocompact special groups satisfy these conditions, so the Strong Atiyah conjecture holds for virtually cocompact special groups.
Study confirms equivalence in Heisenberg groups between curvature-dimension conditions and strong Brunn-Minkowski inequalities.
Paper finds a counterexample showing rectangle condition doesn't detect strong irreducibility.
Sparse multinomial logistic regression for multiclass classification with feature selection.
Paper relaxes stability and generalization assumptions for SGD.
The study examines different types of equilibria for stopping problems in one-dimensional diffusion processes.
Two new methods reduce random forest latency and improve accuracy.
We consider families of strongly consistent multivariate conditional risk measures. We show that under strong consistency these families admit a decomposition into a conditional aggregation function and a univariate conditional risk measure as introduced Hoffmann et al. (2016). Further, in analogy to the univariate cas…
We prove that, under low noise assumptions, the support vector machine with random features (RFSVM) can achieve the learning rate faster than on a training set with samples when an optimized feature map is used. Our work extends the previous fast rate analysis of random features method from…
Study on interest rate model with jumps, proving strong convergence in simulations.
New examples show strong Kato limits can be branching and not satisfy known conditions.
Online L2D algorithm for multiclass classification with varying experts.
Efficient algorithm for contextual bandits with first-order guarantees.
We show that, under some technical conditions, the Strong Slope Conjecture proposed by Kalfagianni and Tran is closed under connect sums and cabling. As an application, we establish the Strong Slope Conjecture for graph knots.
We provide new results concerning label efficient, polynomial time, passive and active learning of linear separators. We prove that active learning provides an exponential improvement over PAC (passive) learning of homogeneous linear separators under nearly log-concave distributions. Building on this, we provide a comp…
Random Fourier features classification achieves fast learning rates with fewer features.
The stability of statistical analysis is an important indicator for reproducibility, which is one main principle of scientific method. It entails that similar statistical conclusions can be reached based on independent samples from the same underlying population. In this paper, we introduce a general measure of classif…