A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
In this paper, we analyze the effects of depth and width on the quality of local minima, without strong over-parameterization and simplification assumptions in the literature. Without any simplification assumption, for deep nonlinear neural networks with the squared loss, we theoretically show that the quality of local…
Consider the 1-dimensional Hurwitz space parameterizing covers of P^1 branched at four points. We study its intersection with divisor classes on the moduli space of curves. As an application, we calculate the slope of the Teichmuller curve parameterizing square-tiled cyclic covers and recover the sum of its Lyapunov ex…
We show that the gradient descent algorithm provides an implicit regularization effect in the learning of over-parameterized matrix factorization models and one-hidden-layer neural networks with quadratic activations. Concretely, we show that given O~(dr2) random linear measurements of a rank r positive s…
This paper proves SGD converges to global minimum for over-parameterized ReLU networks.
problem Theoretical understanding of implicit neural networks is limited.
method Gradient flow analysis of ReLU activated implicit neural networks.
result Randomly initialized gradient descent converges to global minimum at a linear rate for square loss function in over-parameterized ReLU networks.
We propose randomized least-squares value iteration (RLSVI) -- a new reinforcement learning algorithm designed to explore and generalize efficiently via linearly parameterized value functions. We explain why versions of least-squares value iteration that use Boltzmann or epsilon-greedy exploration can be highly ineffic…
The paper studies implicit regularization in over-parameterized models for high-dimensional data.
problem Understanding implicit regularization in over-parameterized models for high-dimensional data.
method The paper designs regularization-free algorithms for the high-dimensional single index model and provides theoretical guarantees for the induced implicit regularization phenomenon.
result The proposed methods achieve minimax optimal statistical rates of convergence and outperform classical methods with explicit regularization.
Recently, several studies have proven the global convergence and generalization abilities of the gradient descent method for two-layer ReLU networks. Most studies especially focused on the regression problems with the squared loss function, except for a few, and the importance of the positivity of the neural tangent ke…
In this short note, we prove that the space of all admissible piecewise linear metrics parameterized by length square on a triangulated manifolds is a convex cone. We further study Regge's Einstein-Hilbert action and give a much more reasonable definition of discrete Einstein metric than our former version in \cite{G}.…
R. Schwartz's inequality provides an upper bound for the Schwarzian derivative of a parameterization of a circle in the complex plane and on the potential of Hill's equation with coexisting periodic solutions. We prove a discrete version of this inequality and obtain a version of the planar Blaschke-Santalo inequality …
We consider the optimization problem associated with training simple ReLU neural networks of the form x↦∑i=1kmax{0,wi⊤x} with respect to the squared loss. We provide a computer-assisted proof that even if the input distribution is standard Gaussian, even if the dime…
In this paper, we construct polynomial growth harmonic maps from once-punctured Riemann surfaces of any finite genus to any even-sided, regular, ideal polygon in the hyperbolic plane. We also establish their uniqueness within a class of maps which differ by exponentially decaying variations. Previously, harmonic maps f…
The "double descent" risk curve was proposed to qualitatively describe the out-of-sample prediction accuracy of variably-parameterized machine learning models. This article provides a precise mathematical analysis for the shape of this curve in two simple data models with the least squares/least norm predictor. Specifi…
We present a general-purpose method to train Markov chain Monte Carlo kernels, parameterized by deep neural networks, that converge and mix quickly to their target distribution. Our method generalizes Hamiltonian Monte Carlo and is trained to maximize expected squared jumped distance, a proxy for mixing speed. We demon…
We introduce a Gaussian process model of functions which are additive. An additive function is one which decomposes into a sum of low-dimensional functions, each depending on only a subset of the input variables. Additive GPs generalize both Generalized Additive Models, and the standard GP models which use squared-expo…
We present an efficient second-order algorithm with O~(η1T) regret for the bandit online multiclass problem. The regret bound holds simultaneously with respect to a family of loss functions parameterized by η, for a range of η restricted by the norm of the competitor. The family of loss funct…
Let α(s) be an arc on a connected oriented surface S in E3, parameterized by arc length s, with torsion τ and length l. The total square torsion F of α is defined by T=\int_{0}^{l}τ^{2}ds\ $. . The arc α is called a relaxed elastic line of second kind if it is an extremal for the variational problem of minimizing the v…
We consider the task of low-multilinear-rank functional regression, i.e., learning a low-rank parametric representation of functions from scattered real-valued data. Our first contribution is the development and analysis of an efficient gradient computation that enables gradient-based optimization procedures, including…
Let α be an arc on a connected oriented surface S in Minkowski 3-space, parameterized by arc length s, with torsion τ and length l. The total square torsion H of α is defined by . The arc is called a relaxed elastic line of second kind if it is an extremal for the variational prob…