Gradient descent with identity initialization learns positive definite linear transformations efficiently.
problem Learning positive definite linear transformations using gradient descent.
method Gradient descent with identity initialization, analyzing the population quadratic loss.
result Gradient descent with identity initialization efficiently learns positive definite linear transformations.
Conditions for conformal Killing vectors in vacuum spacetimes.
problem Finding conditions for conformal Killing vectors in vacuum spacetimes.
method Classical argument to identify a suitable propagation identity and check well-posedness of the initial value problem.
result Necessary and sufficient conditions for conformal Killing initial data (CKID) are found, extending known Killing initial data (KID).
A new neural network initialization method is proposed for faster and more accurate training.
problem Efficient initialization for training multi-layer feedforward neural networks.
method Initialization based on Stein's identity, using eigenvectors of cross-moment matrix.
result The SteinGLM method is faster and more accurate than other initialization methods.
Mimetic initialization improves Transformer training on small datasets.
problem Difficulty in training Transformers on small datasets.
method Initialize self-attention layers to look like pre-trained models.
result Vanilla Transformers trained with mimetic initialization achieve higher accuracy.
Maximizing withdrawal success in a pooled annuity fund with multiple annuitants.
problem Optimizing withdrawal success in a pooled annuity fund with homogeneous annuitants.
method Maximizing the probability of completing withdrawals until death over portfolio weight functions.
result Increasing the number of annuitants can significantly increase the maximum probability of withdrawal success.
Batch normalization makes deep residual networks train faster.
problem Training deep residual networks with large depths.
method Downscaling the residual branch by a normalizing factor early in training.
result Normalized residual blocks compute functions close to the identity function early in training.
Inverse curvature flow studied in anti-de Sitter-Schwarzschild manifold.
problem Analyzing curvature flow in a specific spacetime geometry.
method Inverse hessian quotient curvature flow with star-shaped initial hypersurface.
result The solution exists for all time and converges exponentially fast to the identity.
Gradient descent converges globally in deep linear residual networks with ZAS initialization.
problem Optimizing deep linear residual networks for convergence.
method Zero-asymmetric (ZAS) initialization for gradient descent.
result Gradient descent converges to an ε-optimal point in O(L^3 log(1/ε)) iterations.
MBML improves multi-task RL by inferring task identity from state-action pairs.
problem Multi-task batch reinforcement learning with unseen tasks.
method MBML uses triplet loss and relabeling to robustify task inference.
result Significantly faster convergence on unseen tasks compared to random initialization.
New findings show that minimal feature diversity at neural network initialization is harmful but can be mitigated with noise.
problem The importance of feature diversity at neural network initialization.
method A series of experiments comparing different initialization schemes, including adding noise.
result Minimal feature diversity is harmful but can be mitigated with noise, even standard GPU noise is sufficient.
New method shows random, diverse initializations are not essential for deep neural networks.
problem The necessity of random, diverse initializations in deep neural networks.
method Constructed a deep convolutional network with identical features by initializing weights to 0, enabling signal propagation and stable gradients.
result Random, diverse initializations are not necessary for training neural networks.
SupSup model learns thousands of tasks without forgetting, using randomly initialized subnetworks.
problem Sequentially learning many tasks without forgetting.
method Randomly initialized base network with task-specific subnetworks (supermasks).
result Gradient-based optimization can identify the correct subnetwork for new tasks.
The paper establishes principles for initializing and designing GNNs with ReLU activations to avoid oversmoothing and correlation collapse.
problem Oversmoothing and correlation collapse in deep ReLU GNNs.
method The paper derives and validates three principles for initialization and architecture selection in finite width graph neural networks with ReLU activations.
result Correct initialization, residual aggregation operators, and residual connections significantly improve early training dynamics in deep ReLU GNNs.
New equations connect unit Killing vectors to initial data.
problem Characterizing initial data for Einstein vacuum with unit Killing vectors.
method Developed new equations (uKID) by eliminating scaling and using propagation identity.
result Found equations that are finite type and characterize unit normalized Killing vectors.
We investigate unification of two systems of identical elements having different dimensions which may be of interest for both physics and economics. Characteristic parameters as well as explicit formulae for the temperature (in economics - capital turnover) and dimension of the united system are obtained as functions o…
Barren plateaus are not an average-case phenomenon, but a highly non-unique problem.
problem Avoiding barren plateaus in neural network training
method First-moment framework for initialization strategies
result Many families of inequivalent initialization strategies can avoid concentration
The closed string model in the background gravity field is considered as a bi-Hamiltonian system in assumption that string model is the integrable model for particular kind of the background fields. The dual nonlocal Poisson brackets(PB), depending of the background fields and of their derivatives, are obtained. The in…
Zero Initialization improves short-term load forecasting accuracy.
problem Improving the learning speed and accuracy of neural networks for load forecasting.
method Proposed and tested Zero Initialization (ZI) for weights of a single layer network, comparing with Xavier, He, and Identity initialization.
result ZI reduces the number of epochs and improves accuracy in short-term load forecasting.
Given asymptotically flat initial data on M^3 for the vacuum Einstein field equation, and given a bounded domain in M, we construct solutions of the vacuum constraint equations which agree with the original data inside the given domain, and are identical to that of a suitable Kerr slice (or identical to a member of som…
Sharp stability estimate for tensor tomography in non-positive curvature.
problem Stability estimate for tensor tomography on manifolds with non-positive curvature.
method Pestov identity with localized frequency boundary term.
result Stability estimate of the form L2↦HT1/2. Geometrically interpolates rigid body motions with initial and terminal twists.
problem Finding spatial trajectories between prescribed initial and terminal poses.
method Derives solutions for k-IV-TIP and k-BV-TIP for k=1,...,4.
result Automatic cubic interpolation identical to minimum acceleration curve when twists are zero.
Proves existence of multi-phase flows from arbitrary initial data.
problem Non-uniqueness issue in Brakke flows.
method Global existence proof for multi-phase mean curvature flow.
result Validates explicit identity for evolving grain volumes.
New sampling and identity-testing methods for mixtures of distributions that don't satisfy approximate tensorization of entropy.
problem Sampling and identity-testing for mixtures of distributions that don't satisfy approximate tensorization of entropy.
method Fast mixing of Glauber dynamics and efficient identity-testers in the coordinate-conditional sampling access model.
result Efficient identity-testers for mixtures of ATE distributions in the coordinate-conditional sampling access model.
Study on memorization vs. generalization in overparameterized networks.
problem Understanding the trade-off between memorization and generalization in neural networks.
method Examined fully-connected and convolutional networks trained to minimize reconstruction error.
result Different architectures exhibit distinct inductive biases, affecting generalization from a single training example.
Theory explains deep nonlinear networks' plateaus and transitions.
problem Understanding long plateaus and feature acquisition transitions in deep nonlinear networks.
method Derived an exact identity for Frobenius norms, classified activation functions, and reduced matrix flow to a scalar ODE.
result Escape time law τ⋆=Θ(ε−(r−2)) for deep nonlinear networks, where r is the number of bottleneck layers. New method for initializing RBM weights without datasets.
problem No dataset-free weight-initialization for RBMs.
method Statistical mechanical analysis to derive Gaussian distribution with optimized standard deviation.
result Optimal weight initialization improves learning efficiency in RBMs.
Deep learning predicts user identity, activity, and location from Wi-Fi signals.
problem Privacy concerns and need for non-invasive user authentication, activity classification, and tracking.
method End-to-end deep learning framework using passive Wi-Fi signals.
result System autonomously predicts user identity, activity, and location without user intervention.
Paper finds infinitely many solutions changing sign for critical fractional equations.
problem Critical fractional equations with sign-changing solutions.
method Reduction to equivalent problem on sphere, blow-up arguments, Pohozaev's identity, regularity results, symmetries of sphere.
result Unbounded sequence of sign-changing solutions for critical problems.
Layer normalization with activations prevents Gram matrix rank collapse at initialization.
problem Rank collapse in Gram matrices at initialization slows training in deep networks.
method Proved that layer normalization, with activation layers, biases Gram matrix towards identity matrix at exponential rate.
result Layer normalization with activations biases Gram matrix towards identity matrix at exponential rate with depth at initialization.
A fast regime-split Black-Scholes implied volatility solver
problem Fast computation of implied volatility
method Analytical and numerical expansions
result Achieves near-machine precision with minimal iterations
Framework captures neural network learning in large-width limit.
problem Understanding learning dynamics in large neural networks.
method Developed a rigorous framework for multilayer neural networks in mean field limit.
result Global convergence guarantees for various network architectures and initializations.
Quantum circuit models learn better with specific initialization strategies.
problem Understanding and improving the optimization landscape of IQP-based generative models.
method Proved barren plateaus for random initialization, established lower bounds, and developed data-dependent initialization.
result Data-dependent initialization leads to faster convergence and better minimums.
We prove that torsion-free G_2 structures are (weakly) dynamically stable along the Laplacian flow for closed G_2 structures. More precisely, given a torsion-free G_2 structure φ on a compact 7-manifold, the Laplacian flow with initial value cohomologous and sufficiently close to φ will converge to a to…
Study shows bifurcation in optimal retirement planning.
problem Optimal consumption and retirement planning model.
method Cobb-Douglas utility, simple model with wealth bifurcation.
result Critical wealth level leads to a continuum of retirement trajectories.
Improved neural network training by coupled initialization reduces neuron count.
problem Training neural networks efficiently with fewer neurons.
method Coupled initialization of weights into pairs of identical Gaussian vectors.
result Significantly reduced number of neurons required for network convergence.
We consider two functions on Sp(g,R) with values in the cyclic group of order four {1,-1,i,-i}. One was defined by Lion and Vergne. The other is -i raised to the power given by an integer valued function defined by Masbaum and the author (initially on the mapping class group of a surface). We identify these functions w…
In this paper an approach is proposed to represent a class of dissipative mechanical systems by corresponding infinite-dimensional Hamiltonian systems. This approach is based upon the following structure: for any non-conservative classical mechanical system and arbitrary initial conditions, there exists a conservative …
MCE reduces embedding instability in nonlinear dimensionality reduction.
problem Embedding instability caused by random initialization.
method Median of multiple embeddings (MCE) based on large deviation theory.
result MCE achieves consistency at an exponential rate and effectively mitigates instability.
Randomly initialized transformers show extreme token preferences.
problem Structural biases in randomly initialized transformers.
method Dissection of transformer architecture at initialization.
result Initialization-induced biases persist throughout training.
We extend the lifecycle model (LCM) of consumption over a random horizon (a.k.a. the Yaari model) to a world in which (i.) the force of mortality obeys a diffusion process as opposed to being deterministic, and (ii.) a consumer can adapt their consumption strategy to new information about their mortality rate (a.k.a. h…
Bayesian model predicts online activity participation.
problem Predicting the number of new users initiating an activity.
method Simple Bayesian approach for online activity sample sizes.
result Effective in predicting sample size for online experiments.
We prove that on a Kähler manifold admitting an extremal metric ω and for any Kähler potential φ0 close to ω, the Calabi flow starting at φ0 exists for all time and the modified Calabi flow starting at φ0 will always be close to ω. Furthermore, when the initial data is invariant under t…
This paper examines the impact of random initialization in neural networks using NTK theory.
problem Understanding the impact of random initialization in neural networks using NTK theory.
method Analyzes the convergence of training dynamics and generalization error of wide neural networks with random initialization.
result The generalization error of wide neural networks trained by gradient descent is \( \Omega(n^{-\frac{3}{d+3}}) \), highlighting the benefits of mirror initialization and suggesting limitations of NTK theory.
Deep GCNII tackles over-smoothing problem in graph convolutional networks.
problem Over-smoothing problem in shallow graph convolutional networks.
method Proposes GCNII with initial residual and identity mapping techniques.
result Deep GCNII outperforms state-of-the-art methods on various tasks.
Proves well-posedness of gradient solitons on bundle gerbe.
problem Analyzing gradient generalized Ricci solitons on bundle gerbes.
method Analytic Cauchy problem solved for gradient solitons on abelian bundle gerbes.
result Initial data equations solved on compact Riemann surfaces.
The paper explores conditions for homothetic Killing vectors on spacetime hypersurfaces.
problem Conditions for the existence of homothetic Killing vectors on spacetime hypersurfaces.
method General identities relating deformation tensor and tensor on hypersurfaces, applied to specific settings.
result Necessary and sufficient conditions for homothetic Killing vectors on spacetime hypersurfaces.
Constructs extensions for Bartnik data to approach mass limits.
problem Finding mass limits for Bartnik data.
method Shi-Tam type metric construction and refined monotonicity.
result Mass of constructed extensions can be made arbitrarily close to half area radius.
When the vacuum Einstein equations are cast in the form of hamiltonian evolution equations, the initial data lie in the cotangent bundle of the manifold MΣ of riemannian metrics on a Cauchy hypersurface Σ. As in every lagrangian field theory with symmetries, the initial data must satisfy constraints. But, unlike those …