New KSDs control moments in approximations, improving diagnostics and tests.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This study shows the moment-SOS hierarchy converges in polynomial optimization over product of spheres.
For a GJR-GARCH specification with a generic innovation distribution we derive analytic expressions for the first four conditional moments of the forward and aggregated returns and variances. Moment for the most commonly used GARCH models are stated as special cases. We also the limits of these moments as the time hori…
The paper examines linking numbers in grid models and finds polynomial moments.
Adaptive gradient methods such as Adam have been shown to be very effective for training deep neural networks (DNNs) by tracking the second moment of gradients to compute the individual learning rates. Differently from existing methods, we make use of the most recent first moment of gradients to compute the individual …
ADOPT optimizes Adam to converge with any β2 without bounded noise.
COS method convergence conditions expanded for heavy-tailed distributions.
Study resolvent convergence for random matrices with general covariance profiles.
Samplets and multiwavelets constructed from scattered data converge to specific densities in the limit.
Optimal Transport (OT) problems arise in a wide range of applications, from physics to economics. Getting numerical approximate solution of these problems is a challenging issue of practical importance. In this work, we investigate the relaxation of the OT problem when the marginal constraints are replaced by some mome…
We associate certain probability measures on to geodesics in the space $\H_L$ of positively curved metrics on a line bundle , and to geodesics in the finite dimensional symmetric space of hermitian norms on . We prove that the measures associated to the finite dimensional spaces converge weakly to t…
The learning of domain-invariant representations in the context of domain adaptation with neural networks is considered. We propose a new regularization method that minimizes the discrepancy between domain-specific latent feature representations directly in the hidden activation space. Although some standard distributi…
DualAdam improves generalization of Adam by integrating its update mechanisms.
Learning rate needs to decrease with higher data moments for effective ICA in high dimensions.
Derives a series expansion for Asian option pricing with polynomial jump-diffusion moments.
Let G be a countable group which acts by isometries on a separable, but not necessarily proper, Gromov hyperbolic space X. We say the action of G is weakly hyperbolic if G contains two independent hyperbolic isometries. We show that a random walk on such G converges to the Gromov boundary almost surely. We apply the co…
Optimizes mixture models without parametrizing distributions using tensor decomposition.
Nonlinear SGD achieves high-probability rates in non-convex optimization with heavy-tailed noise.
Gradient descent on Hadamard manifolds converges to boundary points, solving optimization problems.
New approach for testable learning using moment matching and Rademacher complexity.
New adaptive stepsize method for stochastic approximation converges to target point.
Donaldson defined a parabolic flow on Kahler manifolds which arises from considering the action of a group of symplectomorphisms on the space of smooth maps between manifolds. One can define a moment map for this action, and then consider the gradient flow of the square of its norm. Chen discovered the same flow from a…
We consider least squares estimation in a general nonparametric regression model. The rate of convergence of the least squares estimator (LSE) for the unknown regression function is well studied when the errors are sub-Gaussian. We find upper bounds on the rates of convergence of the LSE when the errors have uniformly …
New SGMM algorithm for efficient estimation of moment restriction models.
Study on eigenvalue distribution of correlated time series, showing deformation of Marchenko-Pastur distribution.
MGD combines maximum entropy and diffusion methods for efficient sampling.
Paper proposes an efficient algorithm to handle high-order portfolio moments.
PMT uses public data moments to make DP feasible for unbounded data.
Gradient descent with random weights in linear regression analyzed for various noise types.
We present a detailed analysis of \emph{observable} moments based parameter estimators for the Heston SDEs jointly driving the rate of returns and the squared volatilities . Since volatilities are not directly observable, our parameter estimators are constructed from empirical moments of realized volatilitie…
AdamNX improves Adam's stability by adjusting its learning rate.
A new method for assessing Bayesian sampling quality, PSD, is proposed and shown to be more powerful and efficient.
Mixture modeling is a general technique for making any simple model more expressive through weighted combination. This generality and simplicity in part explains the success of the Expectation Maximization (EM) algorithm, in which updates are easy to derive for a wide class of mixture models. However, the likelihood of…
Adaptive moment methods have been remarkably successful in deep learning optimization, particularly in the presence of noisy and/or sparse gradients. We further the advantages of adaptive moment techniques by proposing a family of double adaptive stochastic gradient methods~\textsc{DASGrad}. They leverage the complemen…
For a finite-dimensional (but possibly noncompact) symplectic manifold with a compact group acting with a proper moment map, we show that the square of the moment map is an equivariantly perfect Morse function in the sense of Kirwan, and that the set of critical points of the square of the moment map is a countable dis…
The study analyzes convergence of adaptive optimizers under low-precision training.
Asynchronous stochastic gradient descent (ASGD) is a popular parallel optimization algorithm in machine learning. Most theoretical analysis on ASGD take a discrete view and prove upper bounds for their convergence rates. However, the discrete view has its intrinsic limitations: there is no characterization of the optim…
This paper proposes a Lasso-type estimator for a high-dimensional sparse parameter identified by a single index conditional moment restriction (CMR). In addition to this parameter, the moment function can also depend on a nuisance function, such as the propensity score or the conditional choice probability, which we es…
A new method for generating samples without training, using smoothed score matching.
In this paper we investigate the convergence properties of the upwards gradient flow of the norm-square of a moment map on the space of representations of a quiver. The first main result gives a necessary and sufficient algebraic criterion for a complex group orbit to intersect the unstable set of a given critical poin…
Study on convergence rates for optimal transport with regularization.
The paper strengthens the classical result of MLE convergence to a Gaussian distribution.
This paper considers the problem of implementing large-scale gradient descent algorithms in a distributed computing setting in the presence of {\em straggling} processors. To mitigate the effect of the stragglers, it has been previously proposed to encode the data with an erasure-correcting code and decode at the maste…
Paper proposes a policy gradient method for confounded POMDPs.
We extend some properties of random walks on hyperbolic groups to random walks on convergence groups. In particular we prove that if a convergence group acts on a compact metrizable space with the convergence property then we can provide with a compact topology such that random walks on converge a…
We construct a canonical Hausdorff complex analytic moduli space of Fano manifolds with Kähler-Ricci solitons. This naturally enlarges the moduli space of Fano manifolds with Kähler-Einstein metrics, which was constructed by Odaka and Li-Wang-Xu. We discover a moment map picture for Kähler-Ricci solitons, and give comp…
A financial model without short-selling shows deviations from normality.
QEM uses parallel importance weighting for fast approximate Bayesian inference.