Gradient descent dynamics in quadratic regression models are analyzed, revealing five phases: monotonic, catapult, periodic, chaotic, and divergent.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We provide geometric conditions on a pair of hyperplanes of a CAT(0) cube complex that imply divergence bounds for the cube complex. As an application, we classify all right-angled Coxeter groups with quadratic divergence and show right-angled Coxeter groups cannot exhibit a divergence function between quadratic and cu…
We show that both Teichmuller space (with the Teichmuller metric) and the mapping class group (with a word metric) have geodesic divergence that is intermediate between the linear rate of flat spaces and the exponential rate of hyperbolic spaces. For every two geodesic rays in Teichmuller space, we find that their dive…
The set of directions from a quadratic differential that diverge on average under Teichmuller geodesic flow has Hausdorff dimension exactly equal to one-half.
Let W be a 2-dimensional right-angled Coxeter group. We characterise such W with linear and quadratic divergence, and construct right-angled Coxeter groups with divergence polynomial of arbitrary degree. Our proofs use the structure of walls in the Davis complex.
Square percolation determines threshold for group divergence in random graphs.
In this paper we address the problem of artist style transfer where the painting style of a given artist is applied on a real world photograph. We train our neural networks in adversarial setting via recently introduced quadratic potential divergence for stable learning process. To further improve the quality of genera…
We study the asymptotic geometry of Teichmueller geodesic rays. We show that when the transverse measures to the vertical foliations of the quadratic differentials determining two different rays are topologically equivalent, but are not absolutely continuous with respect to each other, then the rays diverge in Teichmue…
We study the limits of holonomy representations of complex projective structures on a compact Riemann surface in the Morgan-Shalen compactification of the character variety. We show that the dual R-trees of the quadratic differentials associated to a divergent sequence of projective structures determine the Morgan-Shal…
The paper rethinks the use of exponential averaging in machine learning optimization.
Study on divergence and thickness for Coxeter groups, generalizing previous work.
Study geodesics in curved spaces, counts ambiguous paths, confirms number theory conjectures.
New bound matches exact generalization error for quadratic Gaussian problem.
Efficiently reduces rank of non-negative matrices with quadratic time complexity.
New formulations for comparing metric measure spaces with arbitrary positive measures.
This paper provides efficient algorithms for computing entropy and KL divergence in Bayesian networks.
We consider the limit set in Thurston's compactification PMF of Teichmueller space of some Teichmueller geodesics defined by quadratic differentials with minimal but not uniquely ergodic vertical foliations. We show that a) there are quadratic differentials so that the limit set of the geodesic is a unique point, b) th…
In this paper we explore relationships between divergence and thick groups, and with the same techniques we estimate lengths of shortest conjugators. We produce examples, for every positive integer n, of CAT(0) groups which are thick of order n and with polynomial divergence of order n+1, both these phenomena are new. …
Adapts RKHS methods to estimate density ratios with optimal error.
Optimal transport and information geometry both study geometric structures on spaces of probability distributions. Optimal transport characterizes the cost-minimizing movement from one distribution to another, while information geometry originates from coordinate-invariant properties of statistical inference. Their con…
Study measures complexity of surfaces using a new graph to prove group properties.
A typical goal of supervised dimension reduction is to find a low-dimensional subspace of the input space such that the projected input variables preserve maximal information about the output variables. The dependence maximization approach solves the supervised dimension reduction problem through maximizing a statistic…
In this paper, we develop an approach to recursively estimate the quadratic risk for matrix recovery problems regularized with spectral functions. Toward this end, in the spirit of the SURE theory, a key step is to compute the (weak) derivative and divergence of a solution with respect to the observations. As such a so…
EGMU optimizes portfolios using KL divergence, ensuring positive solutions.
A new ParVI framework improves particle-based variational inference methods.
The Bregman divergence (Bregman distance, Bregman measure of distance) is a certain useful substitute for a distance, obtained from a well-chosen function (the "Bregman function"). Bregman functions and divergences have been extensively investigated during the last decades and have found applications in optimization, o…
New bounds found for optimizing non-convex functions with noisy data.
A new method speeds up computation of Sinkhorn divergences to linear time.
Threshold found for hyperbolicity in random Coxeter groups.
We study some analytical and geometric properties of a two-dimensional nonlinear sigma model with gravitino which comes from supersymmetric string theory. When the action is critical w.r.t. variations of the various fields including the gravitino, there is a symmetric, traceless and divergence-free energy-momentum tens…
Improved analysis for diffusion models reduces KL divergence error dependence on data dimension and discretization step size.
On four-dimensional closed manifolds we introduce a class of canonical Riemannian metrics, that we call weak harmonic Weyl metrics, defined as critical points in the conformal class of a quadratic functional involving the norm of the divergence of the Weyl tensor. This class includes Einstein and, more in general, harm…
New method uses neural networks to solve complex PDEs from optimal control theory.
One-pass SGD dynamics in overparameterized quadratic networks show slow escape from poor solutions.
We consider the minimization of composite objective functions composed of the expectation of quadratic functions and an arbitrary convex function. We study the stochastic dual averaging algorithm with a constant step-size, showing that it leads to a convergence rate of O(1/n) without strong convexity assumptions. This …
New method accelerates energetic variational inference using particle dynamics.
Reinforcement Learning(RL) with sparse rewards is a major challenge. We propose \emph{Hindsight Trust Region Policy Optimization}(HTRPO), a new RL algorithm that extends the highly successful TRPO algorithm with \emph{hindsight} to tackle the challenge of sparse rewards. Hindsight refers to the algorithm's ability to l…
The paper extends results on Bach-flat solitons to new types.
Index tracking is a popular form of asset management. Typically, a quadratic function is used to define the tracking error of a portfolio and the look back approach is applied to solve the index tracking problem. We argue that a forward looking approach is more suitable, whereby the tracking error is expressed as expec…
Selective removal of data subsets can efficiently unlearn unwanted distributions.
The study proves uniqueness of large isoperimetric sets in specific noncompact manifolds.
New insights into how large learning rates affect transformer training dynamics.
We study two global structural properties of a graph , denoted AS and CFS, which arise in a natural way from geometric group theory. We study these properties in the Erdös--Rényi random graph model G(n,p), proving a sharp threshold for a random graph to have the AS property asymptotically almost surely, and giving f…
Graph-based methods provide a powerful tool set for many non-parametric frameworks in Machine Learning. In general, the memory and computational complexity of these methods is quadratic in the number of examples in the data which makes them quickly infeasible for moderate to large scale datasets. A significant effort t…
MAP inference for general energy functions remains a challenging problem. While most efforts are channeled towards improving the linear programming (LP) based relaxation, this work is motivated by the quadratic programming (QP) relaxation. We propose a novel MAP relaxation that penalizes the Kullback-Leibler divergence…
New algorithms optimize risk for large datasets, improving efficiency.
The paper explores how the Gauss curvature of Riemannian surfaces can be represented as the divergence of a vector field.
We study the long-time behaviour of nonnegative solutions of the Porous Medium Equation posed on Cartan-Hadamard manifolds having very large negative curvature, more precisely when the sectional or Ricci curvatures diverge at infinity more than quadratically in terms of the geodesic distance to the pole. We find an une…