Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

316192122 · Jun 202019922001200920172026
48 results for He initialization

Initialization of parameters in deep neural networks has been shown to have a big impact on the performance of the networks (Mishkin & Matas, 2015). The initialization scheme devised by He et al, allowed convolution activations to carry a constrained mean which allowed deep networks to be trained effectively (He et al.…

2017-02-21abs ↗pdf ↗

A hybrid model combines BPH and HE distributions for better heavy-tailed distribution approximation.

problem Accurate modeling of heavy-tailed distributions in various applications.
method A hybrid model of Bernstein phase-type and hyperexponential distributions with optimized parameters.
result Significant improvement in capturing both body and tail of heavy-tailed distributions.

The Residual Network (ResNet), proposed in He et al. (2015), utilized shortcut connections to significantly reduce the difficulty of training, which resulted in great performance boosts in terms of both training and generalization error. It was empirically observed in He et al. (2015) that stacking more layers of resid…

2016-11-03abs ↗pdf ↗

This paper analyzes the Lipschitz constants of deep neural networks with random weights.

problem Estimating the Lipschitz constants of deep neural networks with random parameters.
method High probability upper and lower bounds derived for ReLU neural networks with He initialization.
result The behavior of the Lipschitz constant varies significantly between p[1,2)p \in [1,2) and p[2,]p \in [2,\infty].

The paper establishes principles for initializing and designing GNNs with ReLU activations to avoid oversmoothing and correlation collapse.

problem Oversmoothing and correlation collapse in deep ReLU GNNs.
method The paper derives and validates three principles for initialization and architecture selection in finite width graph neural networks with ReLU activations.
result Correct initialization, residual aggregation operators, and residual connections significantly improve early training dynamics in deep ReLU GNNs.

Consider an agent who enters a financial market on day t = 0 with an initial capital amount x. He invests this amount on stocks and the money market, and by day t = T, has generated a wealth W . He is given a convex class of probability measures (called scenarios) and a real-valued function (or floors) corresponding to…

2006-01-25abs ↗pdf ↗

New method stabilizes deep neural networks by setting Lyapunov exponent to zero.

problem Stability issues in deep neural networks with low width.
method Lyapunov initialization method to set Lyapunov exponent to zero.
result Lyapunov exponent governs stability of deep networks; standard methods fail for low width.

Our study analyzes how neural network initialization affects privacy and utility in overparameterized models.

problem Privacy and utility trade-off in overparameterized neural networks.
method Analytical proof of KL divergence privacy bound, focusing on initialization, width, and depth.
result Privacy bound improvement with increasing depth under certain initializations, degradation under others.

We prove that two-layer (Leaky)ReLU networks initialized by e.g. the widely used method proposed by He et al. (2015) and trained using gradient descent on a least-squares loss are not universally consistent. Specifically, we describe a large class of one-dimensional data-generating distributions for which, with high pr…

2020-02-12abs ↗pdf ↗

Unified learning-rate scale for CNNs and ResNets, avoiding depth imbalance.

problem Challenges in choosing an appropriate learning rate for deep networks, especially as depth increases.
method Introduces Arithmetic-Mean μμP (AM-μμP), constraining network-wide average pre-activation second moment to a constant scale, combined with residual-aware He fan-in initialization.
result Demonstrates a 3/2-3/2 scaling law for learning rates across depths, enabling zero-shot learning-rate transfer.

Optimizes deep neural network initialization variance for better performance.

problem Improving deep neural network performance through optimal initialization variance.
method Using SGD dynamics and Fokker-Planck equations, we study the relationship between initialization and expected loss function.
result An optimal condition for initialization variance that leads to lower training loss and higher test accuracy.

Mini-Hes improves LFA model performance on HDI tasks with missing data.

problem Effective representation of high-dimensional, incomplete data for user behavior understanding.
method Proposes Mini-Hes, a parallelizable second-order LFA model using mini-block diagonal Hessian-free optimization.
result Mini-Hes outperforms state-of-the-art models in missing data estimation tasks on recommender system datasets.

We prove that on Fano manifolds, the Kähler-Ricci flow produces a "most destabilising" degeneration, with respect to a new stability notion related to the H-functional. This answers questions of Chen-Sun-Wang and He. We give two applications of this result. Firstly, we give a purely algebro-geometric formula for the su…

2016-12-21abs ↗pdf ↗

Consider an American option that pays G(X^*_t) when exercised at time t, where G is a positive increasing function, X^*_t := \sup_{s\le t}X_s, and X_s is the price of the underlying security at time s. Assuming zero interest rates, we show that the seller of this option can hedge his position by trading in the underlyi…

2011-08-20abs ↗pdf ↗

A Bayesian agent learns about the structure of a stationary process from ob- serving past outcomes. We prove that his predictions about the near future become ap- proximately those he would have made if he knew the long run empirical frequencies of the process.

2014-06-25abs ↗pdf ↗

Adversarial training (AT) is one of the most effective defenses against adversarial attacks for deep learning models. In this work, we advocate incorporating the hypersphere embedding (HE) mechanism into the AT procedure by regularizing the features onto compact manifolds, which constitutes a lightweight yet effective …

2020-02-20abs ↗pdf ↗

This paper studies bounds for the Lipschitz constant of random neural networks.

problem Quantifying the worst-case robustness of neural networks against adversarial perturbations.
method Analyzes upper and lower bounds for the Lipschitz constant of random ReLU neural networks under specific initialization conditions.
result For deep networks, the upper bound is larger than the lower bound by a logarithmic factor in width.

Proves initial data on big bang singularities for Einstein-nonlinear scalar field equations lead to unique solutions.

problem Initial data on big bang singularities for Einstein equations.
method Geometric formulation of initial data, proving existence and uniqueness of solutions.
result Initial data on the singularity for the Einstein-nonlinear scalar field equations in 4 spacetime dimensions lead to a unique development of the data.

We review some ideas of Grothendieck and others on actions of the absolute Galois group Γ Q of Q (the automorphism group of the tower of finite extensions of Q), related to the geometry and topology of surfaces (mapping class groups, Teichm{ü}ller spaces and moduli spaces of Riemann surfaces). Grothendieck's motivation…

2016-03-10abs ↗pdf ↗

We solve the equivalence problem for the orthogonally separable webs on the three-sphere under the action of the isometry group. This continues a classical project initiated by Olevsky in which he solved the corresponding canonical forms problem. The solution to the equivalence problem together with the results by Olev…

2010-09-22abs ↗pdf ↗

In this paper we consider a modification of the classical Merton portfolio optimization problem. Namely, an investor can trade in financial asset and consume his capital. He is additionally endowed with a one unit of an indivisible asset which he can sell at any time. We give a numerical example of calculating the opti…

2014-03-13abs ↗pdf ↗

Study on the complexity of 1D ReLU neural networks, proving growth in linear regions.

problem Understanding the complexity and expressivity of 1D ReLU neural networks.
method Analyzing the number of linear regions in randomly initialized, fully connected 1D ReLU networks in the infinite-width limit.
result The expected number of linear regions grows as a function of the number of neurons in each layer.

Theory explains deep nonlinear networks' plateaus and transitions.

problem Understanding long plateaus and feature acquisition transitions in deep nonlinear networks.
method Derived an exact identity for Frobenius norms, classified activation functions, and reduced matrix flow to a scalar ODE.
result Escape time law τ=Θ(ε(r2))τ_\star = Θ(\varepsilon^{-(r-2)}) for deep nonlinear networks, where rr is the number of bottleneck layers.

In a recent work, Galloway [9] proved a local foliation theorem by MOTSs for a 3-dimensional initial data set (M,g,K)(M,g,K) with mean curvature τ0τ\le0 in a 4-dimensional spacetime (M,g)(\overline M,\overline g) when (under suitable assumptions) MM has a stable spherical MOTS ΣΣ which achieves an upper bound for the area. H…

2015-04-25abs ↗pdf ↗

Herbert Gr{ö}tzsch is the main founder of the theory of quasicon-formal mappings. We review five of his papers, written between 1928 and 1932, that show the progress of his work from conformal to quasiconformal geometry. This will give an idea of his motivation for introducing quasicon-formal mappings, of the problems …

2019-12-17abs ↗pdf ↗

The paper proves rigidity and uniformization theorems for infinite circle patterns and convex polyhedra in hyperbolic 3-space.

problem Characterize infinite circle patterns and convex polyhedra in hyperbolic 3-space.
method Extends techniques from previous work to prove rigidity and uniformization theorems for infinite circle patterns and convex polyhedra.
result Establishes existence and rigidity of infinite regular circle patterns and convex trivalent polyhedra.

We study pricing and (super)hedging for American options in an imperfect market model with default, where the imperfections are taken into account via the nonlinearity of the wealth dynamics. The payoff is given by an RCLL adapted process (ξt)(ξ_t). We define the {\em seller's superhedging price} of the American option a…

2017-08-29abs ↗pdf ↗

In this paper, the author considers the numerical computation of CVA for large systems by Mote Carlo methods. He introduces two types of stochastic mesh methods for the computations of CVA. In the first method, stochastic mesh method is used to obtain the future value of the derivative contracts. In the second method, …

2015-10-15abs ↗pdf ↗

In 2003, S.-s. Chern began a study of almost-complex structures on the 6-sphere, with the idea of exploiting the special properties of its well-known almost-complex structure invariant under the exceptional group G2G_2. While he did not solve the (currently still open) problem of determining whether there exists an int…

2014-05-14abs ↗pdf ↗