New research challenges the flatness-generalization link in deep neural networks.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The article examines in some detail the convergence rate and mean-square-error performance of momentum stochastic gradient methods in the constant step-size and slow adaptation regime. The results establish that momentum methods are equivalent to the standard stochastic gradient method with a re-scaled (larger) step-si…
New research shows shrinkage methods re-scale portfolio efficient frontiers under distributional misspecification.
In this paper we introduce three methods for re-scaling data sets aiming at improving the likelihood of clustering validity indexes to return the true number of spherical Gaussian clusters with additional noise features. Our method obtains feature re-scaling factors taking into account the structure of a given data set…
Boosting is a learning scheme that combines weak prediction rules to produce a strong composite estimator, with the underlying intuition that one can obtain accurate prediction rules by combining "rough" ones. Although boosting is proved to be consistent and overfitting-resistant, its numerical convergence rate is rela…
In this paper, we study the backward Ricci flow on locally homogeneous 3-manifolds. We describe the long time behavior and show that, typically and after a proper re-scaling, there is convergence to a sub-Riemannian geometry. A similar behavior was observed by the authors in the case of the cross curvature flow.
We compute the analytic expression of the probability distributions F{FTSE100,+} and F{FTSE100,-} of the normalized positive and negative FTSE100 (UK) index daily returns r(t). Furthermore, we define the alpha re-scaled FTSE100 daily index positive returns r(t)^alpha and negative returns (-r(t))^alpha that we call, aft…
Layer normalization (LayerNorm) has been successfully applied to various deep neural networks to help stabilize training and boost model convergence because of its capability in handling re-centering and re-scaling of both inputs and weight matrix. However, the computational overhead introduced by LayerNorm makes these…
ASAM improves deep neural network generalization by adapting sharpness to scale.
In terms of the stock exchange returns, we compute the analytic expression of the probability distributions F{DAX,+} and F{DAX,-} of the normalized positive and negative DAX (Germany) index daily returns r(t). Furthermore, we define the alpha re-scaled DAX daily index positive returns r(t)^alpha and negative returns (-…
The first order behavior of multivariate heavy-tailed random vectors above large radial thresholds is ruled by a limit measure in a regular variation framework. For a high dimensional vector, a reasonable assumption is that the support of this measure is concentrated on a lower dimensional subspace, meaning that certai…
Paper proves BGW tau-function can be represented as Q-polynomials.
The Yokonuma-Hecke algebras are quotients of the modular framed braid group and they support Markov traces. In this paper, which is sequel to Juyumaya and Lambropoulou (2007), we explore further the structures of the -adic framed braids and the -adic Yokonuma-Hecke algebras constructed in Juyumaya and Lambropoulo…
We introduce a prototype model in an attempt to capture some aspects of market dynamics simulating a trading mechanism. The model description starts with a discrete-space, continuous-time Markov process describing arrival and movement of orders with different prices. We then perform a re-scaling procedure leading to a …
We give a formal and complete characterization of the explicit regularizer induced by dropout in deep linear networks with squared loss. We show that (a) the explicit regularizer is composed of an -path regularizer and other terms that are also re-scaling invariant, (b) the convex envelope of the induced regula…
This paper highlights the size-dependency of income distributions, i.e. the income distribution curves versus the population of a country systematically. By using the generalized Lotka-Volterra model to fit the empirical income data in the United States during 1996-2007, we found an important parameter can scale wi…
Let be a principal bundle. Consider a sequence of metrics on obtained by re-scaling the fibers to points. The Gromov-Hausdorff limit of the tangent bundles over these principal bundles with their Sasaki metric is seen herein to be a locally trivial fiber bundle containing the tangent space to the base as a…
Layer normalization (LayerNorm) is a technique to normalize the distributions of intermediate layers. It enables smoother gradients, faster training, and better generalization accuracy. However, it is still unclear where the effectiveness stems from. In this paper, our main contribution is to take a step further in und…
Proves a formula for Kontsevich-Witten tau-function using Schur Q-polynomials.
A stochastic theory for the toppling activity in sandpile models is developed, based on a simple mean-field assumption about the toppling process. The theory describes the process as an anti-persistent Gaussian walk, where the diffusion coefficient is proportional to the activity. It is formulated as a generalization o…
A new method combines predictors and their lags using supervised PCA for dynamic forecasting.
Study eigenvalues of Laplace operator on specific 3D manifolds under Ricci flow.
We consider first order gradient methods for effectively optimizing a composite objective in the form of a sum of smooth and, potentially, non-smooth functions. We present accelerated and adaptive gradient methods, called FLAG and FLARE, which can offer the best of both worlds. They can achieve the optimal convergence …
If a sequence of Riemannian manifolds, , converges in the pointed Gromov-Hausdorff sense to a limit space, , and if are vector bundles over endowed with metrics of Sasaki-type with a uniform upper bound on rank, then a subsequence of the converges in the pointed Gromov-Hausdorff sense t…
Defines magnitude for length spaces with measures, agreeing with finite spaces' magnitude.
Variational Bayesian neural networks combine the flexibility of deep learning with Bayesian uncertainty estimation. However, inference procedures for flexible variational posteriors are computationally expensive. A recently proposed method, noisy natural gradient, is a surprisingly simple method to fit expressive poste…
Deformed holomorphic Chern-Simons theory yields new instantons.
We construct a sequence of compact embedded minimal disks in the unit ball in Euclidean 3-space whose boundaries are in the boundary of the ball and where the curvatures blow up at every point of a line segment of the vertical axis, extending from the origin. We further study the transversal structure of the minimal li…
We analyze dropout in deep networks with rectified linear units and the quadratic loss. Our results expose surprising differences between the behavior of dropout and more traditional regularizers like weight decay. For example, on some simple data sets dropout training produces negative weights even though the output i…
Study minimax rates for density estimation under Huber contamination and Besov IPM losses.
A characterization of the C-projective vector fields on a Randers spaces is presented in terms of a recently introduced non-Riemannian quantity defined by Z. Shen and denoted by ; It is proved that the quantity is invariant for C-projective vector fields. Therefore, the dimension of the algebra of the …
New insights into simple kernel smoothing reveal surprising asymptotics.
Kernel-based L2-boosting with structure constraints improves regression efficiency.
Volatility estimation based on high-frequency data is key to accurately measure and control the risk of financial assets. A Lévy process with infinite jump activity and microstructure noise is considered one of the simplest, yet accurate enough, models for financial data at high-frequency. Utilizing this model, we prop…
The paper studies deep neural networks with Gaussian weights and finds their asymptotic behavior.
The paper examines VI for overparameterized BNNs, revealing a trade-off between likelihood and KL terms.
We prove that if is a CW-complex and is a 0-cell of , then the crossed module does not depend on the cellular decomposition of up to free products with , where is the 1-skeleton of . From this it follows that if is a finite crossed module and is finite, the…
Categorical regressor variables are usually handled by introducing a set of indicator variables, and imposing a linear constraint to ensure identifiability in the presence of an intercept, or equivalently, using one of various coding schemes. As proposed in Yuan and Lin [J. R. Statist. Soc. B, 68 (2006), 49-67], the gr…
New method estimates sparse canonical vectors efficiently.
We prove that if is a CW-complex and is its 1-skeleton then the crossed module depends only on the homotopy type of as a space, up to free products, in the category of crossed modules, with . From this it follows that, if is a finite crossed module and is finite, then th…
The paper studies conformal invariants of Riemannian manifolds and proves vanishing theorems and inequalities.
In the semantic segmentation of street scenes the reliability of the prediction and therefore uncertainty measures are of highest interest. We present a method that generates for each input image a hierarchy of nested crops around the image center and presents these, all re-scaled to the same size, to a neural network …
Motivated by string topology and the arc operad, we introduce the notion of quasi-operads and consider four (quasi)-operads which are different varieties of the operad of cacti. These are cacti without local zeros (or spines) and cacti proper as well as both varieties with fixed constant size one of the constituting lo…
Using the large deviation principle (LDP) for a re-scaled fractional Brownian motion where the rate function is defined via the reproducing kernel Hilbert space, we compute small-time asymptotics for a correlated fractional stochastic volatility model of the form $dS_t=S_tσ(Y_t) (\barρ dW_t +ρdB_t), \,dY_t=dB^H…
The paper stratifies positive scalar curvature metrics on manifolds.
To keep up with increasing dataset sizes and model complexity, distributed training has become a necessity for large machine learning tasks. Parameter servers ease the implementation of distributed parameter management---a key concern in distributed training---, but can induce severe communication overhead. To reduce c…
Deep learning estimates time-varying Markov model parameters.
Training-free model learns SDE dynamics without training, accelerating parameter studies.