Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

94188282376 · Jun 202019922001200920172026
48 results for parameter re-scaling

New research challenges the flatness-generalization link in deep neural networks.

problem The correlation between flatness of the loss landscape and generalization in deep neural networks is questioned.
method The study examines various flatness measures and popular SGD variants, finding some break the flatness-generalization link. It proposes using logP(f)\log P(f), a global quantity, as a predictor of generalization.
result The log of Bayesian prior upon initialization, logP(f)\log P(f), is a significantly more robust predictor of generalization than flatness measures.

The article examines in some detail the convergence rate and mean-square-error performance of momentum stochastic gradient methods in the constant step-size and slow adaptation regime. The results establish that momentum methods are equivalent to the standard stochastic gradient method with a re-scaled (larger) step-si…

2016-03-14abs ↗pdf ↗

New research shows shrinkage methods re-scale portfolio efficient frontiers under distributional misspecification.

problem Poor performance of mean-variance portfolio decisions under distributional assumptions.
method Investigation of shrinkage methods under different distributional assumptions (auto-correlation, skewness, excess kurtosis).
result Shrinkage methods re-scale the sample efficient frontier, implying standard comparison methods are flawed.

Boosting is a learning scheme that combines weak prediction rules to produce a strong composite estimator, with the underlying intuition that one can obtain accurate prediction rules by combining "rough" ones. Although boosting is proved to be consistent and overfitting-resistant, its numerical convergence rate is rela…

2015-05-06abs ↗pdf ↗

In this paper, we study the backward Ricci flow on locally homogeneous 3-manifolds. We describe the long time behavior and show that, typically and after a proper re-scaling, there is convergence to a sub-Riemannian geometry. A similar behavior was observed by the authors in the case of the cross curvature flow.

2008-10-18abs ↗pdf ↗

We compute the analytic expression of the probability distributions F{FTSE100,+} and F{FTSE100,-} of the normalized positive and negative FTSE100 (UK) index daily returns r(t). Furthermore, we define the alpha re-scaled FTSE100 daily index positive returns r(t)^alpha and negative returns (-r(t))^alpha that we call, aft…

2010-04-07abs ↗pdf ↗

Layer normalization (LayerNorm) has been successfully applied to various deep neural networks to help stabilize training and boost model convergence because of its capability in handling re-centering and re-scaling of both inputs and weight matrix. However, the computational overhead introduced by LayerNorm makes these…

2019-10-16abs ↗pdf ↗

ASAM improves deep neural network generalization by adapting sharpness to scale.

problem Fixed-radius sharpness measure is sensitive to parameter scaling, weakening its connection to generalization.
method Introduces adaptive sharpness, a scale-invariant measure, and proposes ASAM for deep learning.
result ASAM significantly improves model generalization performance across various datasets.

In terms of the stock exchange returns, we compute the analytic expression of the probability distributions F{DAX,+} and F{DAX,-} of the normalized positive and negative DAX (Germany) index daily returns r(t). Furthermore, we define the alpha re-scaled DAX daily index positive returns r(t)^alpha and negative returns (-…

2010-04-07abs ↗pdf ↗

The first order behavior of multivariate heavy-tailed random vectors above large radial thresholds is ruled by a limit measure in a regular variation framework. For a high dimensional vector, a reasonable assumption is that the support of this measure is concentrated on a lower dimensional subspace, meaning that certai…

2019-06-26abs ↗pdf ↗

The Yokonuma-Hecke algebras are quotients of the modular framed braid group and they support Markov traces. In this paper, which is sequel to Juyumaya and Lambropoulou (2007), we explore further the structures of the pp-adic framed braids and the pp-adic Yokonuma-Hecke algebras constructed in Juyumaya and Lambropoulo…

2009-05-22abs ↗pdf ↗

We introduce a prototype model in an attempt to capture some aspects of market dynamics simulating a trading mechanism. The model description starts with a discrete-space, continuous-time Markov process describing arrival and movement of orders with different prices. We then perform a re-scaling procedure leading to a …

2012-01-22abs ↗pdf ↗

We give a formal and complete characterization of the explicit regularizer induced by dropout in deep linear networks with squared loss. We show that (a) the explicit regularizer is composed of an 2\ell_2-path regularizer and other terms that are also re-scaling invariant, (b) the convex envelope of the induced regula…

2019-05-28abs ↗pdf ↗

This paper highlights the size-dependency of income distributions, i.e. the income distribution curves versus the population of a country systematically. By using the generalized Lotka-Volterra model to fit the empirical income data in the United States during 1996-2007, we found an important parameter λλ can scale wi…

2010-12-09abs ↗pdf ↗

Let PMP\to M be a principal bundle. Consider a sequence of metrics on PP obtained by re-scaling the fibers to points. The Gromov-Hausdorff limit of the tangent bundles over these principal bundles with their Sasaki metric is seen herein to be a locally trivial fiber bundle containing the tangent space to the base as a…

2015-03-31abs ↗pdf ↗

Layer normalization (LayerNorm) is a technique to normalize the distributions of intermediate layers. It enables smoother gradients, faster training, and better generalization accuracy. However, it is still unclear where the effectiveness stems from. In this paper, our main contribution is to take a step further in und…

2019-11-16abs ↗pdf ↗

Study eigenvalues of Laplace operator on specific 3D manifolds under Ricci flow.

problem Analyze eigenvalues of Laplace operator with potential under backward Ricci flow.
method Use backward Ricci flow on locally homogeneous 3-manifolds, derive bounds and convergence results.
result Eigenvalue λ+(t)λ^{+}(t) approaches zero as flow converges to sub-Riemannian geometry.

We consider first order gradient methods for effectively optimizing a composite objective in the form of a sum of smooth and, potentially, non-smooth functions. We present accelerated and adaptive gradient methods, called FLAG and FLARE, which can offer the best of both worlds. They can achieve the optimal convergence …

2016-05-26abs ↗pdf ↗

If a sequence of Riemannian manifolds, XiX_i, converges in the pointed Gromov-Hausdorff sense to a limit space, XX_\infty, and if EiE_i are vector bundles over XiX_i endowed with metrics of Sasaki-type with a uniform upper bound on rank, then a subsequence of the EiE_i converges in the pointed Gromov-Hausdorff sense t…

2010-11-02abs ↗pdf ↗

Variational Bayesian neural networks combine the flexibility of deep learning with Bayesian uncertainty estimation. However, inference procedures for flexible variational posteriors are computationally expensive. A recently proposed method, noisy natural gradient, is a surprisingly simple method to fit expressive poste…

2018-11-30abs ↗pdf ↗

Deformed holomorphic Chern-Simons theory yields new instantons.

problem Deforming classical holomorphic Chern-Simons theory on Calabi-Yau manifolds.
method Deformation of complex structure by a parameter \( h \) leading to new instanton solutions.
result Existence of instanton solutions invariant under re-scalings of \( h \) and their connection to \( G_2 \)-instantons.

We analyze dropout in deep networks with rectified linear units and the quadratic loss. Our results expose surprising differences between the behavior of dropout and more traditional regularizers like weight decay. For example, on some simple data sets dropout training produces negative weights even though the output i…

2016-02-14abs ↗pdf ↗

Study minimax rates for density estimation under Huber contamination and Besov IPM losses.

problem Minimax convergence rates of nonparametric density estimation under Huber contamination model with outliers.
method Re-scaled thresholding wavelet series estimator and GAN architectures.
result Achieves minimax optimal convergence rates under Besov IPM losses.

A characterization of the C-projective vector fields on a Randers spaces is presented in terms of a recently introduced non-Riemannian quantity defined by Z. Shen and denoted by Ξ{\bfΞ}; It is proved that the quantity Ξ{\bfΞ} is invariant for C-projective vector fields. Therefore, the dimension of the algebra of the …

2018-11-06abs ↗pdf ↗

The paper studies deep neural networks with Gaussian weights and finds their asymptotic behavior.

problem Understanding the behavior of deep neural networks with large width.
method Function-space perspective, Gaussian process analysis, weak convergence in large-width limit.
result Deep neural networks with large width converge to a continuous Gaussian process.

The paper examines VI for overparameterized BNNs, revealing a trade-off between likelihood and KL terms.

problem Critical issue in mean-field VI training for overparameterized BNNs.
method Theoretical and empirical study of overparameterized two-layer BNNs using VI.
result A trade-off between likelihood and KL terms in overparameterized regime, with KL scaling crucial.

We prove that if MM is a CW-complex and * is a 0-cell of MM, then the crossed module Π2(M,M1,)Π_2(M,M^1,*) does not depend on the cellular decomposition of MM up to free products with Π2(D2,S1,)Π_2(D^2,S^1,*), where M1M^1 is the 1-skeleton of MM. From this it follows that if GG is a finite crossed module and MM is finite, the…

2005-07-12abs ↗pdf ↗

We prove that if MM is a CW-complex and M1M^1 is its 1-skeleton then the crossed module Π2(M,M1)Π_2(M,M^1) depends only on the homotopy type of MM as a space, up to free products, in the category of crossed modules, with Π2(D2,S1)Π_2(D^2,S^1). From this it follows that, if GG is a finite crossed module and MM is finite, then th…

2008-01-25abs ↗pdf ↗

The paper studies conformal invariants of Riemannian manifolds and proves vanishing theorems and inequalities.

problem Analyzing conformal invariants of Riemannian manifolds and their implications.
method Defining new conformal invariants and proving vanishing theorems and inequalities.
result Established inequalities relating conformal invariants to other geometric invariants.

Motivated by string topology and the arc operad, we introduce the notion of quasi-operads and consider four (quasi)-operads which are different varieties of the operad of cacti. These are cacti without local zeros (or spines) and cacti proper as well as both varieties with fixed constant size one of the constituting lo…

2002-09-11abs ↗pdf ↗

Using the large deviation principle (LDP) for a re-scaled fractional Brownian motion BtHB^H_t where the rate function is defined via the reproducing kernel Hilbert space, we compute small-time asymptotics for a correlated fractional stochastic volatility model of the form $dS_t=S_tσ(Y_t) (\barρ dW_t +ρdB_t), \,dY_t=dB^H…

2016-10-27abs ↗pdf ↗

To keep up with increasing dataset sizes and model complexity, distributed training has become a necessity for large machine learning tasks. Parameter servers ease the implementation of distributed parameter management---a key concern in distributed training---, but can induce severe communication overhead. To reduce c…

2020-02-03abs ↗pdf ↗

Deep learning estimates time-varying Markov model parameters.

problem Estimating time-dependent parameters in Markov models.
method Reframes parameter estimation as an optimization problem using maximum likelihood.
result Real solution close to SDE with neural network-derived parameters under specific conditions.

Training-free model learns SDE dynamics without training, accelerating parameter studies.

problem High computational cost of simulating parameter-dependent SDEs.
method Training-free conditional diffusion model with joint kernel-weighted Monte Carlo estimator.
result Accurate approximation of conditional distributions across varying parameter values.