Gradient descent implicitly follows regularization for general losses.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This work extends implicit bias analysis to multiclass classification using a new loss framework.
Unified framework approximates gradient descent's implicit bias in high dimensions.
Mirror flow optimizes separable data problems, converging to a maximum margin classifier.
New model captures time-varying volatility with stochastic exponential tails.
A fast method for training linear classifiers maximizes margins.
New algorithms improve stopping time for best arm identification.
Large stepsize GD for logistic regression converges faster than expected.
Stochastic Gradient Descent (SGD) is a central tool in machine learning. We prove that SGD converges to zero loss, even with a fixed (non-vanishing) learning rate - in the special case of homogeneous linear classifiers with smooth monotone loss functions, optimized on linearly separable data. Previous works assumed eit…
We provide a detailed study on the implicit bias of gradient descent when optimizing loss functions with strictly monotone tails, such as the logistic loss, over separable datasets. We look at two basic questions: (a) what are the conditions on the tail of the loss function under which gradient descent converges in the…
New loss function restores importance weighting in overparameterized models.
We present an empirical study of the subordination hypothesis for a stochastic time series of a stock price. The fluctuating rate of trading is identified with the stochastic variance of the stock price, as in the continuous-time random walk (CTRW) framework. The probability distribution of the stock price changes (log…
Analysis of gradient descent on wide neural networks reveals strong generalization.
We consider a priori generalization bounds developed in terms of cross-validation estimates and the stability of learners. In particular, we first derive an exponential Efron-Stein type tail inequality for the concentration of a general function of n independent random variables. Next, under some reasonable notion of s…
This paper shows that the implicit bias of gradient descent on linearly separable data is exactly characterized by the optimal solution of a dual optimization problem given by a smoothed margin, even for general losses. This is in contrast to prior results, which are often tailored to exponentially-tailed losses. For t…
Study on U-statistics with heavy-tailed samples, providing tail bounds and LDP.
Efficiently estimates covariance for sub-Weibull vectors with sub-Gaussian rate.
A random walk on a separable, geodesic hyperbolic metric space converges to the boundary with probability one when the step distribution supports two independent loxodromics. In particular, the random walk makes positive linear progress. Progress is known to be linear with exponential decay when …
We derive exponential tail inequalities for sums of random matrices with no dependence on the explicit matrix dimensions. These are similar to the matrix versions of the Chernoff bound and Bernstein inequality except with the explicit matrix dimensions replaced by a trace quantity that can be small even when the dimens…
Study shows wealth distribution tails near criticality are not universal.
New bounds for non-convex estimators without Bernstein condition.
Gradient methods avoid overfitting on separable data.
This work extends diffusion models to handle heavy-tailed targets, improving score estimation and sampling guarantees.
We prove near-tight concentration of measure for polynomial functions of the Ising model under high temperature. For any degree , we show that a degree- polynomial of a -spin Ising model exhibits exponential tails that scale as at radius . Our concentration radius is opti…
There is accumulating evidence in the literature that stability of learning algorithms is a key characteristic that permits a learning algorithm to generalize. Despite various insightful results in this direction, there seems to be an overlooked dichotomy in the type of stability-based generalization bounds we have in …
We provide upper bounds of the expected Wasserstein distance between a probability measure and its empirical version, generalizing recent results for finite dimensional Euclidean spaces and bounded functional spaces. Such a generalization can cover Euclidean spaces with large dimensionality, with the optimal dependence…
Concentration inequalities form an essential toolkit in the study of high dimensional (HD) statistical methods. Most of the relevant statistics literature in this regard is based on sub-Gaussian or sub-exponential tail assumptions. In this paper, we first bring together various probabilistic inequalities for sums of in…
Paper introduces machine learning for time series data, improving nowcasting accuracy.
Bayesian posterior contraction rates improve with decreasing tails
We study random walks on groups with the feature that, roughly speaking, successive positions of the walk tend to be "aligned". We formalize and quantify this property by means of the notion of deviation inequalities. We show that deviation inequalities have several consequences including Central Limit Theorems, the lo…
Paper establishes convergence rates and concentration bounds for stochastic approximation and reinforcement learning with Markovian noise.
Study examines implied volatility behavior in Bachelier model.
Paper provides tail bounds for stochastic mirror descent in heavy-tailed noise.
This work achieves exponential concentration in heavy-tailed data over CAT(κ) spaces using the Fréchet median.
New method improves calibration of neural networks by targeting robust margins and local smoothness.
We study an equivalence of (i) deterministic pathwise statements appearing in the online learning literature (termed \emph{regret bounds}), (ii) high-probability tail bounds for the supremum of a collection of martingales (of a specific form arising from uniform laws of large numbers for martingales), and (iii) in-expe…
Study robust linear regression without distributional assumptions for heavy-tailed responses.
We introduce a new statistical tool (the TP-statistic and TE-statistic) designed specifically to compare the behavior of the sample tail of distributions with power-law and exponential tails as a function of the lower threshold u. One important property of these statistics is that they converge to zero for power laws o…
Paper establishes universal lower bounds and optimal rates for clustering sub-exponential mixture models.
We investigate the Heston model with stochastic volatility and exponential tails as a model for the typical price fluctuations of the Brazilian São Paulo Stock Exchange Index (IBOVESPA). Raw prices are first corrected for inflation and a period spanning 15 years characterized by memoryless returns is chosen for the ana…
The well-known theorem of Dybvig, Ingersoll and Ross shows that the long zero-coupon rate can never fall. This result, which, although undoubtedly correct, has been regarded by many as surprising, stems from the implicit assumption that the long-term discount function has an exponential tail. We revisit the problem in …
Improved concentration inequalities for sub-Weibull variables enhance statistical and machine learning applications.
Unified RMOT framework for non-modelable risk factors reduces audit bounds.
We study the cross-correlation matrix of inventory variations of the most active individual and institutional investors in an emerging market to understand the dynamics of inventory variations. We find that the distribution of cross-correlation coefficient has a power-law form in the bulk followed by …
Paper develops a TR-SSQP method for noisy optimization with heavy-tailed noise.
We study density estimation for classes of shift-invariant distributions over . A multidimensional distribution is "shift-invariant" if, roughly speaking, it is close in total variation distance to a small shift of it in any direction. Shift-invariance relaxes smoothness assumptions commonly used in non-p…
Efficiently estimates covariance matrix for elliptical distributions under strong contamination.
Soft diamond regularizers improve deep learning performance and sparsity.