New analysis shows halting time is predictable for large models, improving optimization efficiency.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Deep ROC analysis improves model selection and interpretation in medical and AI applications.
The Lasso performs well in ultra-sparse linear models with finite support size.
Average teaching complexity for locating target regions among halfspace intersections is Θ(d).
Unified analysis of finite weight averaging methods in deep learning.
New framework analyzes SGD dynamics in large samples and dimensions.
The study uses statistical methods to analyze nuclear mass models.
KL regularization helps RL algorithms by implicitly averaging q-values.
Study optimizes step size for Metropolis algorithm in non-identifiable cases.
New insights into using momentum for non-convex optimization.
Noise-resilient method improves Hurst exponent estimation accuracy in noisy data.
This paper investigates the average-case time complexity of certifying RIP matrices.
This paper studies the lower bound complexity for the optimization problem whose objective function is the average of individual smooth convex functions. We consider the algorithm which gets access to gradient and proximal oracle for each individual component. For the strongly-convex case, we prove such an algorith…
We consider in this work a system of two stochastic differential equations named the perturbed compositional gradient flow. By introducing a separation of fast and slow scales of the two equations, we show that the limit of the slow motion is given by an averaged ordinary differential equation. We then demonstrate that…
Paper analyzes how EMA improves SGD in linear regression.
Improved TD learning with tail averaging and regularization achieves optimal convergence rates.
The paper analyzes MACD using operator theory.
This paper studies non-asymptotic model selection for the general case of arbitrary design matrices and arbitrary nonzero entries of the signal. In this regard, it generalizes the notion of incoherence in the existing literature on model selection and introduces two fundamental measures of coherence---termed as the wor…
Statistical uncertainty of different filtration techniques for market network analysis is studied. Two measures of statistical uncertainty are discussed. One is based on conditional risk for multiple decision statistical procedures and another one is based on average fraction of errors. It is shown that for some import…
Proposes a new framework for balancing average- and worst-case performance in machine learning.
In this paper, we consider the Tensor Robust Principal Component Analysis (TRPCA) problem, which aims to exactly recover the low-rank and sparse components from their sum. Our model is based on the recently proposed tensor-tensor product (or t-product). Induced by the t-product, we first rigorously deduce the tensor sp…
Sharp analysis of out-of-distribution error in overparameterized models with importance weights.
Unified analysis of Federated Averaging and Nesterov FedAvg for linear speedup.
This work characterizes the benefits of averaging schemes widely used in conjunction with stochastic gradient descent (SGD). In particular, this work provides a sharp analysis of: (1) mini-batching, a method of averaging many samples of a stochastic gradient to both reduce the variance of the stochastic gradient estima…
Stochastic variance reduction algorithms have recently become popular for minimizing the average of a large, but finite, number of loss functions. In this paper, we propose a novel Riemannian extension of the Euclidean stochastic variance reduced gradient algorithm (R-SVRG) to a compact manifold search space. To this e…
Orthogonal Matching Pursuit (OMP) has long been considered a powerful heuristic for attacking compressive sensing problems; however, its theoretical development is, unfortunately, somewhat lacking. This paper presents an improved Restricted Isometry Property (RIP) based performance guarantee for T-sparse signal reconst…
In recent years, stochastic variance reduction algorithms have attracted considerable attention for minimizing the average of a large but finite number of loss functions. This paper proposes a novel Riemannian extension of the Euclidean stochastic variance reduced gradient (R-SVRG) algorithm to a manifold search space.…
The paper offers precise bounds for averaged LSA iterates in linear systems.
Bayesian model averaging (BMA) is the state of the art approach for overcoming model uncertainty. Yet, especially on small data sets, the results yielded by BMA might be sensitive to the prior over the models. Credal Model Averaging (CMA) addresses this problem by substituting the single prior over the models by a set …
This paper analyzes stability and generalization of Markov chain stochastic gradient methods.
The paper studies the convergence of SAA for systemic risk measures.
Learning a similarity metric has gained much attention recently, where the goal is to learn a function that maps input patterns to a target space while preserving the semantic distance in the input space. While most related work focused on images, we focus instead on learning a similarity metric for neuroimages, such a…
New model-free RL algorithm tackles robust average-reward problems with finite sample complexity analysis.
Empirical evidence is given for a significant difference in the collective trend of the share prices during the stock index rising and falling periods. Data on the Dow Jones Industrial Average and its stock components are studied between 1991 and 2008. Pearson-type correlations are computed between the stocks and avera…
Robustly computes intrinsic coordinates on point clouds using resampling and averaging.
Algorithm improves online canonical correlation analysis.
We study least squares linear regression over uncorrelated Gaussian features that are selected in order of decreasing variance. When the number of selected features is at most the sample size , the estimator under consideration coincides with the principal component regression estimator; when , the esti…
In a recent Nature paper, Gabaix et al. \cite{Gabaix03} presented a theory to explain the power law tail of price fluctuations. The main points of their theory are that volume fluctuations, which have a power law tail with exponent roughly -1.5, are modulated by the average market impact function, which describes the r…
OMD and DA perform similarly in static settings but OMD is inferior under dynamic learning rates.
Machine learning models outperform traditional technical analysis in Bitcoin trading.
Simplified analysis of SGD for linear regression with weight averaging.
In this paper, we investigate trading strategies based on exponential moving averages (ExpMAs) of an underlying risky asset. We study both logarithmic utility maximization and long-term growth rate maximization problems and find closed-form solutions when the drift of the underlying is modeled by either an Ornstein-Uhl…
Paper provides tail bounds for stochastic mirror descent in heavy-tailed noise.
Stochastic Gradient Descent (SGD) is one of the simplest and most popular stochastic optimization methods. While it has already been theoretically studied for decades, the classical analysis usually required non-trivial smoothness assumptions, which do not apply to many modern applications of SGD with non-smooth object…
In mixture model-based clustering applications, it is common to fit several models from a family and report clustering results from only the `best' one. In such circumstances, selection of this best model is achieved using a model selection criterion, most often the Bayesian information criterion. Rather than throw awa…
We use variational Gaussian approximations to analyze parametric models with unknown data-generating distributions.
New method quantifies uncertainty in distributed regression.
We investigate the average-case complexity of decision problems for finitely generated groups, in particular the word and membership problems. Using our recent results on ``generic-case complexity'' we show that if a finitely generated group has the word problem solvable in subexponential time and has a subgroup of…