A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Study on L2-boosting behavior as learning rate approaches zero.
problem Understanding the asymptotic behavior of L2-boosting algorithms with vanishing learning rates.
method Analyzes L2-boosting for regression with linear base learners, proving a deterministic limit and characterizing it as a solution to a linear differential equation.
result Proves the existence of a unique solution to the limit problem and analyzes the training and test error.
We obtain expressions for the shear and the vorticity tensors of perfect-fluid spacetimes, in terms of the divergence of the Weyl tensor. For such spacetimes, we prove that if the gradient of the energy density is parallel to the velocity, then either the expansion rate is zero, or the vorticity vanishes. This statemen…
Stochastic Gradient Descent (SGD) is a central tool in machine learning. We prove that SGD converges to zero loss, even with a fixed (non-vanishing) learning rate - in the special case of homogeneous linear classifiers with smooth monotone loss functions, optimized on linearly separable data. Previous works assumed eit…
The Weil-Petersson metric for the moduli space of Riemann surfaces has negative sectional curvature. Surfaces represented in the complement of a compact set in the moduli space have short geodesics. At such surfaces the Weil-Petersson metric is approximately a product metric. An almost product metric has sections with …
Learning rates in stochastic neural network training are currently determined a priori to training, using expensive manual or automated iterative tuning. This study proposes gradient-only line searches to resolve the learning rate for neural network training algorithms. Stochastic sub-sampling during training decreases…
The gradient-based optimization method for deep machine learning models suffers from gradient vanishing and exploding problems, particularly when the computational graph becomes deep. In this work, we propose the tangent-space gradient optimization (TSGO) for the probabilistic models to keep the gradients from vanishin…
Generative Adversarial Networks (GAN) have shown promising results on a wide variety of complex tasks. Recent experiments show adversarial training provides useful gradients to the generator that helps attain better performance. In this paper, we intend to theoretically analyze whether supervised learning with adversar…
This paper contains some vanishing theorems for L2 harmonic forms on complete Riemannian manifolds with a weighted Poincaré inequality and a certain lower bound of the curvature. The results are in the spirit of Li-Wang and Lam, but without assumptions of sign and growth rate of the weight function, so they can be a…
We formulate and study a general family of (continuous-time) stochastic dynamics for accelerated first-order minimization of smooth convex functions. Building on an averaging formulation of accelerated mirror descent, we propose a stochastic variant in which the gradient is contaminated by noise, and study the resultin…
We study utility indifference prices and optimal purchasing quantities for a contingent claim, in an incomplete semi-martingale market, in the presence of vanishing hedging errors and/or risk aversion. Assuming that the average indifference price converges to a well defined limit, we prove that optimally taken position…
In this paper, we formalise order-robust optimisation as an instance of online learning minimising simple regret, and propose Vroom, a zero'th order optimisation algorithm capable of achieving vanishing regret in non-stationary environments, while recovering favorable rates under stochastic reward-generating processes.…
Square metrics F=α(α+β)2 are a special class of Finsler metrics. It is the rate kind of metric category to be of excellent geometrical properties. In this paper, we discuss the so-called singular square metrics F=α(bα+β)2. A characterization for such metrics to be of vanishing Douglas curvature is p…
Stochastic gradient methods are dominant in nonconvex optimization especially for deep models but have low asymptotical convergence due to the fixed smoothness. To address this problem, we propose a simple yet effective method for improving stochastic gradient methods named predictive local smoothness (PLS). First, we …
The paper explores growth rates and Perron numbers in Coxeter systems with low-dimensional Davis complexes.
problem Investigating growth rates and specific types of numbers in Coxeter systems with Davis complexes of low dimension.
method Examining Coxeter systems with Davis complexes of dimension at most 2, focusing on growth rates and specific types of numbers.
result The growth rate of Coxeter systems with Davis complexes of dimension at most 2 are either Salem or Pisot numbers, depending on the Euler characteristic.
The main aim of this paper is to provide an analysis of gradient descent (GD) algorithms with gradient errors that do not necessarily vanish, asymptotically. In particular, sufficient conditions are presented for both stability (almost sure boundedness of the iterates) and convergence of GD with bounded, (possibly) non…
We give another proof for a result of Brick stating that the simple connectivity at infinity is a geometric property of finitely presented groups. This allows us to define the rate of vanishing of $\p1i$ for those groups which are simply connected at infinity. Further we show that this rate is linear for cocompact latt…
This paper investigates the efficacy of a regularized multi-task learning (MTL) framework based on SVM (M-SVM) to answer whether MTL always provides reliable results and how MTL outperforms independent learning. We first find that M-SVM is Bayes risk consistent in the limit of large sample size. This implies that despi…
The paper provides theoretical guarantees for optimized sampling in compressed sensing, showing error vanishes with more measurements.
problem Theoretical and practical improvements in compressed sensing with optimized sampling schemes.
method Theoretical analysis and empirical experiments with optimized sampling schemes for subsampled unitary matrices.
result The error caused by measurement noise vanishes with an increasing number of measurements for optimized sampling schemes, assuming Gaussian noise.
We derive explicit formulas for time decay, for the European call and put options at expiry, and use them to calculate analytical approximations to the price of the American put and early exercise boundary near expiry. We show that for many families of non-Gaussian processes used in empirical studies of financial marke…
We use freeness assumptions of random matrix theory to analyze the dynamical behavior of inference algorithms for probabilistic models with dense coupling matrices in the limit of large systems. For a toy Ising model, we are able to recover previous results such as the property of vanishing effective memories and the a…
Study shows how learning and analytical models affect reneging and jockeying in a dual M/M/1 system.
problem How do learning and analytical models affect reneging and jockeying in a dual M/M/1 system?
method Analytical and online trained actor-critic models were used to study reneging and jockeying in a dual M/M/1 system.
result Both analytical and online trained actor-critic models yield the same asymptotic limits for reneging and jockeying, but differ in practical sizes.