Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

4183124165 · Jun 202019922001200920182026
48 results for Identity initialization

Gradient descent with identity initialization learns positive definite linear transformations efficiently.

problem Learning positive definite linear transformations using gradient descent.
method Gradient descent with identity initialization, analyzing the population quadratic loss.
result Gradient descent with identity initialization efficiently learns positive definite linear transformations.

Conditions for conformal Killing vectors in vacuum spacetimes.

problem Finding conditions for conformal Killing vectors in vacuum spacetimes.
method Classical argument to identify a suitable propagation identity and check well-posedness of the initial value problem.
result Necessary and sufficient conditions for conformal Killing initial data (CKID) are found, extending known Killing initial data (KID).

A new neural network initialization method is proposed for faster and more accurate training.

problem Efficient initialization for training multi-layer feedforward neural networks.
method Initialization based on Stein's identity, using eigenvectors of cross-moment matrix.
result The SteinGLM method is faster and more accurate than other initialization methods.

Maximizing withdrawal success in a pooled annuity fund with multiple annuitants.

problem Optimizing withdrawal success in a pooled annuity fund with homogeneous annuitants.
method Maximizing the probability of completing withdrawals until death over portfolio weight functions.
result Increasing the number of annuitants can significantly increase the maximum probability of withdrawal success.

Inverse curvature flow studied in anti-de Sitter-Schwarzschild manifold.

problem Analyzing curvature flow in a specific spacetime geometry.
method Inverse hessian quotient curvature flow with star-shaped initial hypersurface.
result The solution exists for all time and converges exponentially fast to the identity.

New findings show that minimal feature diversity at neural network initialization is harmful but can be mitigated with noise.

problem The importance of feature diversity at neural network initialization.
method A series of experiments comparing different initialization schemes, including adding noise.
result Minimal feature diversity is harmful but can be mitigated with noise, even standard GPU noise is sufficient.

New method shows random, diverse initializations are not essential for deep neural networks.

problem The necessity of random, diverse initializations in deep neural networks.
method Constructed a deep convolutional network with identical features by initializing weights to 0, enabling signal propagation and stable gradients.
result Random, diverse initializations are not necessary for training neural networks.

The paper establishes principles for initializing and designing GNNs with ReLU activations to avoid oversmoothing and correlation collapse.

problem Oversmoothing and correlation collapse in deep ReLU GNNs.
method The paper derives and validates three principles for initialization and architecture selection in finite width graph neural networks with ReLU activations.
result Correct initialization, residual aggregation operators, and residual connections significantly improve early training dynamics in deep ReLU GNNs.

We investigate unification of two systems of identical elements having different dimensions which may be of interest for both physics and economics. Characteristic parameters as well as explicit formulae for the temperature (in economics - capital turnover) and dimension of the united system are obtained as functions o…

2015-10-03abs ↗pdf ↗

The closed string model in the background gravity field is considered as a bi-Hamiltonian system in assumption that string model is the integrable model for particular kind of the background fields. The dual nonlocal Poisson brackets(PB), depending of the background fields and of their derivatives, are obtained. The in…

2004-11-24abs ↗pdf ↗

Zero Initialization improves short-term load forecasting accuracy.

problem Improving the learning speed and accuracy of neural networks for load forecasting.
method Proposed and tested Zero Initialization (ZI) for weights of a single layer network, comparing with Xavier, He, and Identity initialization.
result ZI reduces the number of epochs and improves accuracy in short-term load forecasting.

Given asymptotically flat initial data on M^3 for the vacuum Einstein field equation, and given a bounded domain in M, we construct solutions of the vacuum constraint equations which agree with the original data inside the given domain, and are identical to that of a suitable Kerr slice (or identical to a member of som…

2003-01-21abs ↗pdf ↗

Sharp stability estimate for tensor tomography in non-positive curvature.

problem Stability estimate for tensor tomography on manifolds with non-positive curvature.
method Pestov identity with localized frequency boundary term.
result Stability estimate of the form L2HT1/2L^2\mapsto H^{1/2}_{T}.

New sampling and identity-testing methods for mixtures of distributions that don't satisfy approximate tensorization of entropy.

problem Sampling and identity-testing for mixtures of distributions that don't satisfy approximate tensorization of entropy.
method Fast mixing of Glauber dynamics and efficient identity-testers in the coordinate-conditional sampling access model.
result Efficient identity-testers for mixtures of ATE distributions in the coordinate-conditional sampling access model.

Study on memorization vs. generalization in overparameterized networks.

problem Understanding the trade-off between memorization and generalization in neural networks.
method Examined fully-connected and convolutional networks trained to minimize reconstruction error.
result Different architectures exhibit distinct inductive biases, affecting generalization from a single training example.

Theory explains deep nonlinear networks' plateaus and transitions.

problem Understanding long plateaus and feature acquisition transitions in deep nonlinear networks.
method Derived an exact identity for Frobenius norms, classified activation functions, and reduced matrix flow to a scalar ODE.
result Escape time law τ=Θ(ε(r2))τ_\star = Θ(\varepsilon^{-(r-2)}) for deep nonlinear networks, where rr is the number of bottleneck layers.

Deep learning predicts user identity, activity, and location from Wi-Fi signals.

problem Privacy concerns and need for non-invasive user authentication, activity classification, and tracking.
method End-to-end deep learning framework using passive Wi-Fi signals.
result System autonomously predicts user identity, activity, and location without user intervention.

Paper finds infinitely many solutions changing sign for critical fractional equations.

problem Critical fractional equations with sign-changing solutions.
method Reduction to equivalent problem on sphere, blow-up arguments, Pohozaev's identity, regularity results, symmetries of sphere.
result Unbounded sequence of sign-changing solutions for critical problems.

Layer normalization with activations prevents Gram matrix rank collapse at initialization.

problem Rank collapse in Gram matrices at initialization slows training in deep networks.
method Proved that layer normalization, with activation layers, biases Gram matrix towards identity matrix at exponential rate.
result Layer normalization with activations biases Gram matrix towards identity matrix at exponential rate with depth at initialization.

Framework captures neural network learning in large-width limit.

problem Understanding learning dynamics in large neural networks.
method Developed a rigorous framework for multilayer neural networks in mean field limit.
result Global convergence guarantees for various network architectures and initializations.

Quantum circuit models learn better with specific initialization strategies.

problem Understanding and improving the optimization landscape of IQP-based generative models.
method Proved barren plateaus for random initialization, established lower bounds, and developed data-dependent initialization.
result Data-dependent initialization leads to faster convergence and better minimums.

We prove that torsion-free G_2 structures are (weakly) dynamically stable along the Laplacian flow for closed G_2 structures. More precisely, given a torsion-free G_2 structure φ\varphi on a compact 7-manifold, the Laplacian flow with initial value cohomologous and sufficiently close to φ\varphi will converge to a to…

2015-04-29abs ↗pdf ↗

We consider two functions on Sp(g,R) with values in the cyclic group of order four {1,-1,i,-i}. One was defined by Lion and Vergne. The other is -i raised to the power given by an integer valued function defined by Masbaum and the author (initially on the mapping class group of a surface). We identify these functions w…

2013-08-05abs ↗pdf ↗

We extend the lifecycle model (LCM) of consumption over a random horizon (a.k.a. the Yaari model) to a world in which (i.) the force of mortality obeys a diffusion process as opposed to being deterministic, and (ii.) a consumer can adapt their consumption strategy to new information about their mortality rate (a.k.a. h…

2012-05-10abs ↗pdf ↗

We prove that on a Kähler manifold admitting an extremal metric ωω and for any Kähler potential φ0\varphi_0 close to ωω, the Calabi flow starting at φ0\varphi_0 exists for all time and the modified Calabi flow starting at φ0\varphi_0 will always be close to ωω. Furthermore, when the initial data is invariant under t…

2010-07-26abs ↗pdf ↗

This paper examines the impact of random initialization in neural networks using NTK theory.

problem Understanding the impact of random initialization in neural networks using NTK theory.
method Analyzes the convergence of training dynamics and generalization error of wide neural networks with random initialization.
result The generalization error of wide neural networks trained by gradient descent is \( \Omega(n^{-\frac{3}{d+3}}) \), highlighting the benefits of mirror initialization and suggesting limitations of NTK theory.

The paper explores conditions for homothetic Killing vectors on spacetime hypersurfaces.

problem Conditions for the existence of homothetic Killing vectors on spacetime hypersurfaces.
method General identities relating deformation tensor and tensor on hypersurfaces, applied to specific settings.
result Necessary and sufficient conditions for homothetic Killing vectors on spacetime hypersurfaces.

When the vacuum Einstein equations are cast in the form of hamiltonian evolution equations, the initial data lie in the cotangent bundle of the manifold MΣ of riemannian metrics on a Cauchy hypersurface Σ. As in every lagrangian field theory with symmetries, the initial data must satisfy constraints. But, unlike those …

2010-03-15abs ↗pdf ↗