New algorithm samples superlinearly growing log-gradient distributions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new tamed stochastic gradient Hamiltonian Monte Carlo algorithm for superlinearly growing stochastic gradients.
Developing a scientific understanding of cities in a fast urbanizing world is essential for planning sustainable urban systems. Recently, it was shown that income and wealth creation follow increasing returns, scaling superlinearly with city size. We study scaling of per capita incomes for separate census defined incom…
RELTA-SGLD stabilizes nonconvex SGLD updates with a lighter taming scheme.
New method samples from non-log-concave distributions with weak dissipativity.
Zeroth-order optimization methods lack inherent privacy guarantees.
Study proves convergence of interest rate model approximations.
Deep neural networks struggle with numerical instability during training.
In this work, we present a globalized stochastic semismooth Newton method for solving stochastic optimization problems involving smooth nonconvex and nonsmooth convex terms in the objective function. We assume that only noisy gradient and Hessian information of the smooth part of the objective function is available via…
The L1-regularized Gaussian maximum likelihood estimator (MLE) has been shown to have strong statistical guarantees in recovering a sparse inverse covariance matrix, or alternatively the underlying graph structure of a Gaussian Markov Random Field, from very limited samples. We propose a novel algorithm for solving the…
Measures of wealth and production have been found to scale superlinearly with the population of a city. Therefore, it makes economic sense for humans to congregate together in dense settlements. A recent model of population dynamics showed that population growth can become superexponential due to the superlinear scalin…
This paper accelerates distributed convex optimization by mitigating ill-conditioning issues.
Accelerated gradient method's stability deteriorates exponentially with steps.
Model shows AI adoption amplifies financial market risk through prediction, herding, and cognitive dependency.
Gaussian processes (GPs) provide a powerful non-parametric framework for reasoning over functions. Despite appealing theory, its superlinear computational and memory complexities have presented a long-standing challenge. State-of-the-art sparse variational inference methods trade modeling accuracy against complexity. H…
Study on neural networks' performance under different normalizations as N grows.
New method corrects bias in estimating entropic risk for better decision-making.
Faster weak supervision framework using triplet methods.
Let M be a closed orientable 3-manifold with a negatively curved Riemannian metric. Let {M_i} be a collection of finite regular covers with degree d_i. (1) If the Heegaard genus of M_i grows more slowly than the square root of d_i, then M_i has positive first Betti number for all sufficiently large i. (2) The strong He…
Over the past few years, neural networks were proven vulnerable to adversarial images: targeted but imperceptible image perturbations lead to drastically different predictions. We show that adversarial vulnerability increases with the gradients of the training objective when viewed as a function of the inputs. Surprisi…
We develop a progressive training approach for neural networks which adaptively grows the network structure by splitting existing neurons to multiple off-springs. By leveraging a functional steepest descent idea, we derive a simple criterion for deciding the best subset of neurons to split and a splitting gradient for …
Minimal graphs grow slowly on curved spaces, proving constant solutions.
Gradient flow expands curves to round shapes.
Time series data constitutes a distinct and growing problem in machine learning. As the corpus of time series data grows larger, deep models that simultaneously learn features and classify with these features can be intractable or suboptimal. In this paper, we present feature learning via long short term memory (LSTM) …
Gradient descent with growing learning rate enables learning non-linear features in neural networks.
Backpropagation-free trunk training improves model performance on various benchmarks.
Cosine similarity can force points to grow in magnitude, causing convergence issues.
Training very deep networks is an important open problem in machine learning. One of many difficulties is that the norm of the back-propagated error gradient can grow or decay exponentially. Here we show that training very deep feed-forward networks (FFNs) is not as difficult as previously thought. Unlike when back-pro…
The paper proves inequalities and growth rates for Schouten solitons.
Thinking LLMs struggle with stock prediction, especially as data complexity increases.
Improved optimization technique reduces training complexity for non-convex problems.
This paper examines a novel gradient boosting framework for regression. We regularize gradient boosted trees by introducing subsampling and employ a modified shrinkage algorithm so that at every boosting stage the estimate is given by an average of trees. The resulting algorithm, titled Boulevard, is shown to converge …
We would like to congratulate the authors of "A Bayesian Conjugate Gradient Method" on their insightful paper, and welcome this publication which we firmly believe will become a fundamental contribution to the growing field of probabilistic numerical methods and in particular the sub-field of Bayesian numerical methods…
Classical stochastic gradient methods for optimization rely on noisy gradient approximations that become progressively less accurate as iterates approach a solution. The large noise and small signal in the resulting gradients makes it difficult to use them for adaptive stepsize selection and automatic stopping. We prop…
We present and analyze several strategies for improving the performance of stochastic variance-reduced gradient (SVRG) methods. We first show that the convergence rate of these methods can be preserved under a decreasing sequence of errors in the control variate, and use this to derive variants of SVRG that use growing…
We investigate Liouville theorems and dimension estimates for the space of exponentially growing holomorphic functions on complete Kähler manifolds. While our work is motivated by the study of gradient Ricci solitons in the theory of Ricci flow, the most general results we prove here do not require any knowledge of cur…
New bounds show BBVI's gradient variance matches SGD conditions, improving parameterization efficiency.
Cross-regularization adapts model complexity during training.
ES and FD gradients converge as optimization dimension grows.
Deep neural networks have enabled progress in a wide variety of applications. Growing the size of the neural network typically results in improved accuracy. As model sizes grow, the memory and compute requirements for training these models also increases. We introduce a technique to train deep neural networks using hal…
DPZero fine-tunes large models privately without backpropagation.
Recent work has established an empirically successful framework for adapting learning rates for stochastic gradient descent (SGD). This effectively removes all needs for tuning, while automatically reducing learning rates over time on stationary problems, and permitting learning rates to grow appropriately in non-stati…
GradaGrad adapts learning rate non-monotonically, overcoming AdaGrad's step size decrease.
New approach uses 'growth' and 'harvesting' concepts to improve deep learning models.
While robust parameter estimation has been well studied in parametric density estimation, there has been little investigation into robust density estimation in the nonparametric setting. We present a robust version of the popular kernel density estimator (KDE). As with other estimators, a robust version of the KDE is u…
New method stabilizes saddle-point optimization with unbounded gradients.
Optimal gradient quantization reduces communication costs in distributed deep learning.
Direct proof shows adaptive gradient descent converges near-linearly for convex functions.