Two parallel samplers enhance image quality in limited denoising steps.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New insights into SGD and SGD-M in high dimensions.
The paper analyzes fixed step-size SA schemes on Riemannian manifolds.
Proposes a neural network for learning step-size policies for L-BFGS optimization.
Gradually Truncated Log-normal distribution - Size distribution of firms Abstract Many natural and economical phenomena are described through power law or log- normal distributions. In these cases, probability decreases very slowly with step size compared to normal distribution. Thus it is essential to cut-off these di…
New SAGA algorithm with decreasing step for stochastic optimization.
In high dimensions, find paths connecting points with intermediate steps in a dense set.
High-dimensional SGD limits show surprising dynamics and phase transitions.
Paper improves CLT and bootstrap approximations for LSA with decreasing step size.
Paper proposes a dual-level approach for multi-step forecasting of dynamical systems.
LOBDIF predicts limit order book events using a diffusion model.
Applying standard Markov chain Monte Carlo (MCMC) algorithms to large data sets is computationally infeasible. The recently proposed stochastic gradient Langevin dynamics (SGLD) method circumvents this problem in three ways: it generates proposed moves using only a subset of the data, it skips the Metropolis-Hastings a…
We establish decoupled functional CLTs for two-time-scale stochastic approximation.
We use a generalization of the Gibbons-Hawking ansatz to study the behavior of certain non-compact Calabi-Yau manifolds in the large complex structure limit. This analysis provides an intermediate step toward proving the metric collapse conjecture for toric hypersurfaces and complete intersections.
The mean-field variant of the model of limit order driven market introduced recently by Maslov is formulated and solved. The agents do not have any strategies and the memory of the system is kept within the order book. We show that he evolution of the order book is governed by a matrix multiplicative process. The resul…
We present a new microscopic stochastic model for an ensemble of interacting investors that buy and sell stocks in discrete time steps via limit orders based on individual forecasts about the price of the stock. These orders determine the supply and demand fixing after each round (time step) the new price of the stock …
The paper discusses polynomial convergence to conical Kähler-Einstein metrics.
The goal of score following is to track a musical performance, usually in the form of audio, in a corresponding score representation. Established methods mainly rely on computer-readable scores in the form of MIDI or MusicXML and achieve robust and reliable tracking results. Recently, multimodal deep learning methods h…
To improve the resilience of distributed training to worst-case, or Byzantine node failures, several recent approaches have replaced gradient averaging with robust aggregation methods. Such techniques can have high computational costs, often quadratic in the number of compute nodes, and only have limited robustness gua…
Community detection in hypergraphs is explored. Under a generative hypergraph model called "d-wise hypergraph stochastic block model" (d-hSBM) which naturally extends the Stochastic Block Model from graphs to d-uniform hypergraphs, the asymptotic minimax mismatch ratio is characterized. For proving the achievability, w…
Negative step sizes improve second-order methods for neural networks.
Stochastic gradient descent (SGD) has been widely used in machine learning due to its computational efficiency and favorable generalization properties. Recently, it has been empirically demonstrated that the gradient noise in several deep learning settings admits a non-Gaussian, heavy-tailed behavior. This suggests tha…
In this paper, we introduce a method for adapting the step-sizes of temporal difference (TD) learning. The performance of TD methods often depends on well chosen step-sizes, yet few algorithms have been developed for setting the step-size automatically for TD learning. An important limitation of current methods is that…
Simpler one-step distributional RL framework for control.
The study examines how limited liability and haircut affect a bank's loan portfolio's liquidity risk.
Few-step protein backbone generators reduce sampling time by over 20x.
Certified training improves robustness against adversarial attacks.
Deep unfolding is a promising deep-learning technique in which an iterative algorithm is unrolled to a deep network architecture with trainable parameters. In the case of gradient descent algorithms, as a result of the training process, one often observes the acceleration of the convergence speed with learned non-const…
Best arm identification (or, pure exploration) in multi-armed bandits is a fundamental problem in machine learning. In this paper we study the distributed version of this problem where we have multiple agents, and they want to learn the best arm collaboratively. We want to quantify the power of collaboration under limi…
We propose a novel time discretization for the log-normal SABR model which is a popular stochastic volatility model that is widely used in financial practice. Our time discretization is a variant of the Euler-Maruyama scheme. We study its asymptotic properties in the limit of a large number of time steps under a certai…
FastVoiceGrad speeds up VC to one step, matching or surpassing quality.
Distributed statistical inference has recently attracted enormous attention. Many existing work focuses on the averaging estimator. We propose a one-step approach to enhance a simple-averaging based distributed estimator. We derive the corresponding asymptotic properties of the newly proposed estimator. We find that th…
New insights into neural network feature learning through multi-step gradient descent.
We construct a global homeomorphism from any 3D Ricci limit space to a smooth manifold, that is locally bi-Holder. This extends the recent work of Miles Simon and the second author, and we build upon their techniques. A key step in our proof is the construction of local "pyramid Ricci flows", existing on uniform region…
Study examines how two-layer networks learn features after one gradient step.
Let be a hyperbolic surface of finite topological type, such that the Fuchsian group is non-elementary, and consider any generating set of . When sampling by an -step random walk in with each step given by an element…
We provide larger step-size restrictions for which gradient descent based algorithms (almost surely) avoid strict saddle points. In particular, consider a twice differentiable (non-convex) objective function whose gradient has Lipschitz constant L and whose Hessian is well-behaved. We prove that the probability of init…
Polyak step size GD reaches final radius of convergence after log iterations.
Proposes QDF to improve multi-step time-series forecasting.
In this paper, we propose new listwise learning-to-rank models that mitigate the shortcomings of existing ones. Existing listwise learning-to-rank models are generally derived from the classical Plackett-Luce model, which has three major limitations. (1) Its permutation probabilities overlook ties, i.e., a situation wh…
In deep latent Gaussian models, the latent variable is generated by a time-inhomogeneous Markov chain, where at each time step we pass the current state through a parametric nonlinear map, such as a feedforward neural net, and add a small independent Gaussian perturbation. This work considers the diffusion limit of suc…
We show that any grafting ray in Teichmüller space is (strongly) asymptotic to some Teichmüller geodesic ray. As an intermediate step we introduce surfaces that arise as limits of these degenerating Riemann surfaces. Given a grafting ray, the proof involves a Teichmüller ray with a conformally equivalent limit, and bui…
EM Distillation simplifies diffusion models to one-step generators.
In this paper, we show that one can naturally associate a limiting dynamical system on an -tree to any degenerating sequence of rational maps $f_n: \hat\C \longrightarrow \hat\C$ of fixed degree. The construction of is in steps: first we use barycentric extension to get $\E f_n : \Hy…
Although reinforcement learning has made great strides recently, a continuing limitation is that it requires an extremely high number of interactions with the environment. In this paper, we explore the effectiveness of reusing experience from the experience replay buffer in the Deep Q-Learning algorithm. We test the ef…
This paper explores the limits of large batch sizes in deep learning.
Study dynamics of alternating minimization for bilinear regression under large system limits.
Improved conformalized quantile regression for adaptive prediction intervals.