A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Study geometric properties of loss functions to understand neural network performance.
problem Understanding the geometric properties of high-dimensional loss functions to improve neural network performance.
method Combine concepts from high-dimensional probability and differential geometry to study curvature properties in lower-dimensional loss representations.
result Mean curvature in the original loss space determines if saddle points appear as minima, maxima, or flat regions.
This paper establishes minimax rates for online regression with arbitrary classes of functions and general losses. We show that below a certain threshold for the complexity of the function class, the minimax rates depend on both the curvature of the loss function and the sequential complexities of the class. Above this…
The speed at which one can minimize an expected loss using stochastic methods depends on two properties: the curvature of the loss and the variance of the gradients. While most previous works focus on one or the other of these properties, we explore how their interaction affects optimization speed. Further, as the ulti…
The total curvature of complex hypersurfaces in $\bC^{n+1}$ and its variation in families appear to depend not only on singularities but also on the behaviour in the neighbourhood of infinity. We find the asymptotic loss of total curvature towards infinity and we express the total curvature and the Gauss-Bonnet defect …
We explore geometric aspects of bubble convergence for harmonic maps. More precisely, we show that the formation of bubbles is characterised by the local excess of curvature on the target manifold. We give a universal estimate for curvature concentration masses at each bubble point and show that there is no curvature l…
Neural network training relies on our ability to find "good" minimizers of highly non-convex loss functions. It is well-known that certain network architecture designs (e.g., skip connections) produce loss functions that train easier, and well-chosen training parameters (batch size, learning rate, optimizer) produce mi…
The local geometry of high dimensional neural network loss landscapes can both challenge our cherished theoretical intuitions as well as dramatically impact the practical success of neural network training. Indeed recent works have observed 4 striking local properties of neural loss landscapes on classification tasks: …
State-of-the-art classifiers have been shown to be largely vulnerable to adversarial perturbations. One of the most effective strategies to improve robustness is adversarial training. In this paper, we investigate the effect of adversarial training on the geometry of the classification landscape and decision boundaries…
The classical asymptotic theory for parametric M-estimators guarantees that, in the limit of infinite sample size, the excess risk has a chi-square type distribution, even in the misspecified case. We demonstrate how self-concordance of the loss allows to characterize the critical sample size sufficient to guarantee …
This manuscript provides optimization guarantees, generalization bounds, and statistical consistency results for AdaBoost variants which replace the exponential loss with the logistic and similar losses (specifically, twice differentiable convex losses which are Lipschitz and tend to zero on one side). The heart of the…
In this paper, we establish the non-positivity of the second eigenvalue of the Schrödinger operator −div(Pr∇⋅)−Wr2 on a closed hypersurface Σn of Rn+1, where Wr is a power of the (r+1)-th mean curvature of Σn. In the case that this eigenvalue is null we have a…
There are many surprising and perhaps counter-intuitive properties of optimization of deep neural networks. We propose and experimentally verify a unified phenomenological model of the loss landscape that incorporates many of them. High dimensionality plays a key role in our model. Our core idea is to model the loss la…
Automatic segmentation of auditory ossicles from CT images using Ricci curvature.
problem Automatic diagnosis of ossicles' diseases from 3D CT images of the head.
method Proposes a completely automatic method that locates and segments ossicles without manual labels or templates, using Ricci curvature in an energy function.
result Performance of the proposed method using discrete Forman-Ricci curvature is superior to state-of-the-art methods.
Stochastic Gradient Descent (SGD) based training of neural networks with a large learning rate or a small batch-size typically ends in well-generalizing, flat regions of the weight space, as indicated by small eigenvalues of the Hessian of the training loss. However, the curvature along the SGD trajectory is poorly und…
This paper examines the role and efficiency of the non-convex loss functions for binary classification problems. In particular, we investigate how to design a simple and effective boosting algorithm that is robust to the outliers in the data. The analysis of the role of a particular non-convex loss for prediction accur…
Burq-Gérard-Tzvetkov and Hu established Lp estimates (2≤p≤∞) for the restriction of eigenfunctions to submanifolds. The estimates are sharp, except for the log loss at the endpoint L2 estimates for submanifolds of codimension 2. It has long been believed that the log loss at the endpoint can be remov…
In [LW], we construct examples of two-dimensional Hamiltonian stationary self-shrinkers and self-expanders for Lagrangian mean curvature flows, which are asymptotic to the union of two Schoen-Wolfson cones. These self-shrinkers and self-expanders can be glued together to yield solutions of the Brakke flow - a weak form…
Differentiable Architecture Search (DARTS) has attracted a lot of attention due to its simplicity and small search costs achieved by a continuous relaxation and an approximation of the resulting bi-level optimization problem. However, DARTS does not work robustly for new problems: we identify a wide range of search spa…