Zero Initialization improves short-term load forecasting accuracy.
problem Improving the learning speed and accuracy of neural networks for load forecasting.
method Proposed and tested Zero Initialization (ZI) for weights of a single layer network, comparing with Xavier, He, and Identity initialization.
result ZI reduces the number of epochs and improves accuracy in short-term load forecasting.
Study shows different initialization schemes for LoRA finetuning impact performance.
problem The impact of initialization schemes on LoRA finetuning performance.
method Compared two initialization schemes: B=0, A=random vs. A=0, B=random.
result First initialization scheme yields better performance on average.
Initial data with zero mass must be in pp-wave spacetimes.
problem Proving initial data with zero mass must be in pp-wave spacetimes.
method Spinorial methods combined with spacetime harmonic functions.
result Initial data with zero mass must be contained in pp-wave spacetimes.
New method stabilizes deep neural networks by setting Lyapunov exponent to zero.
problem Stability issues in deep neural networks with low width.
method Lyapunov initialization method to set Lyapunov exponent to zero.
result Lyapunov exponent governs stability of deep networks; standard methods fail for low width.
Randomly initialized wide neural networks with zero-mean activations are nearly independent, potentially solving AI interpretability limits.
problem Measuring the limits of AI interpretability.
method Randomly initialized neural networks with large width and zero-mean activation functions.
result Neural networks with zero-mean activations are nearly independent, solving the computational no-coincidence conjecture.
Paper optimizes neural network initialization using SMT solvers.
problem Improving neural network performance through better initialization.
method Reduces initialization to SMT problem solving.
result Proposed method achieves better performance than random initialization.
New method for initializing RBM weights without datasets.
problem No dataset-free weight-initialization for RBMs.
method Statistical mechanical analysis to derive Gaussian distribution with optimized standard deviation.
result Optimal weight initialization improves learning efficiency in RBMs.
We study singularities of Lagrangian mean curvature flow in $\C^n$ when the initial condition is a zero-Maslov class Lagrangian. We start by showing that, in this setting, singularities are unavoidable. More precisely, we construct Lagrangians with arbitrarily small Lagrangian angle and Lagrangians which are Hamiltonia…
Gradient descent converges linearly for neural networks with specific conditions.
problem Optimizing neural networks with fixed width and depth.
method Local Polyak-Lojasiewicz criterion for gradient flow and descent.
result Gradient descent converges to zero-loss solutions under certain conditions.
The paper uses curvature to assess stability of linear systems.
problem Assessing stability of linear time-invariant systems.
method Using the first curvature of trajectories to prove stability conditions.
result Conditions for stability and asymptotic stability are derived.
We study the problem of inviscid slightly compressible fluids in a bounded domain. We find a unique solution to the initial-boundary value problem and show that it is near the analogous solution for an incompressible fluid provided the initial conditions for the two problems are close. In particular, the divergence of …
The paper proves stability of the positive mass theorem in spherical symmetry.
problem The stability of the positive mass theorem in spherical symmetry when mass is small.
method Formulated a conjecture and proved it under spherical symmetry assumption.
result A sequence of asymptotically flat initial data converging to Minkowski space.
Gradient descent converges globally in deep linear residual networks with ZAS initialization.
problem Optimizing deep linear residual networks for convergence.
method Zero-asymmetric (ZAS) initialization for gradient descent.
result Gradient descent converges to an ε-optimal point in O(L^3 log(1/ε)) iterations.
We explore how neural networks train to zero loss, focusing on initial scale.
problem Understanding neural network training dynamics and zero loss.
method Macroscopic limits analysis of gradient descent dynamics.
result Gradient descent can drive deep neural networks to zero loss regardless of initialization.
Article strengthens initial data rigidity theorem to show unique spacetime extension.
problem Initial data rigidity in spacetime geometry.
method Showed initial data sets carry a lightlike parallel vector field, leading to unique spacetime extension.
result Local uniqueness of spacetimes extending initial data sets under dominant energy condition.
New bandit algorithms improve sparse reward learning.
problem Sparse rewards hinder learning efficiency in real-world bandit applications.
method Developed algorithms based on Upper Confidence Bound and Thompson Sampling for zero-inflated distributions.
result Empirical performance of new algorithms is superior to existing methods.
The Residual Network (ResNet), proposed in He et al. (2015), utilized shortcut connections to significantly reduce the difficulty of training, which resulted in great performance boosts in terms of both training and generalization error. It was empirically observed in He et al. (2015) that stacking more layers of resid…
On a 4-dimensional compact symplectic manifold, we consider a smooth family of compatible almost-complex structures such that at time zero the induced metric is Hermite-Einstein almost-Kähler metric with zero or negative Hermitian scalar curvature. We prove, under certain hypothesis, the existence of a smooth family of…
Study on when smooth Ricci flow remains smooth at the start.
problem When does a smooth Ricci flow remain smooth down to the initial time?
method Curvature estimates and lower Ricci bounds in three dimensions.
result Positive results for flows with lower curvature bounds, negative for others.
Pólya's theorem extended to meromorphic functions on Riemann surfaces.
problem Distribution of zeros of iterated derivatives of meromorphic functions.
method Recasting local arguments into translation surfaces and using flat metrics.
result Asymptotic distribution of zeros on compact Riemann surfaces.
Robust Lasso-Zero handles missing covariates and sparse corruptions.
problem Sparse corruptions and missing covariates in sparse linear models.
method Extension of Lasso-Zero to handle sparse corruptions, with theoretical guarantees on sign recovery.
result Robust Lasso-Zero can handle missing values without specifying a parametric model.
This paper examines weight initialization for 1-Lipschitz networks to improve robustness against adversarial attacks.
problem Improving the robustness of deep neural networks against adversarial attacks.
method Examined weight parametrization of AOL and SLL networks, calculated weight variance bounds, and demonstrated weight decay.
result Weight initialization causes deep 1-Lipschitz networks to decay to zero, and weight variance does not affect output variance distribution.
We prove that any smooth vacuum spacetime containing a compact Cauchy horizon with surface gravity that can be normalised to a non-zero constant admits a Killing vector field. This proves a conjecture by Moncrief and Isenberg from 1983 under the assumption on the surface gravity and generalises previous results due to …
Study of Bondi-Sachs formalism for massless scalar field with zero cosmological constant.
problem Analyzing the Bondi-Sachs formalism for Einstein's massless scalar field equations.
method Asymptotic expansions and peeling property for Bondi-Sachs metrics and scalar fields.
result Positivity of Bondi energy-momentum under specific conditions.
The paper studies Yamabe metrics and stability in Riemannian manifolds.
problem Existence of complete Yamabe metrics with zero scalar curvature.
method Yamabe flow and local L1-stability analysis. result Local L1-stability of the Yamabe flow on manifolds with non-negative Ricci curvature. Gradient descent learns ReLU functions with non-zero bias efficiently.
problem Learning ReLU functions with non-zero bias under Gaussian distributions.
method Gradient descent starting from random initialization.
result Gradient descent achieves near-optimal error with high probability.
Uncertainty sampling is explained as a gradient step on a smoothed loss, leading to better parameters.
problem Reducing the amount of data required to learn a classifier.
method Interprets uncertainty sampling as a preconditioned stochastic gradient step on a smoothed zero-one loss.
result Uncertainty sampling converges to stationary points of the smoothed population zero-one loss.
We investigate unification of two systems of identical elements having different dimensions which may be of interest for both physics and economics. Characteristic parameters as well as explicit formulae for the temperature (in economics - capital turnover) and dimension of the united system are obtained as functions o…
In this paper we study some global properties of static potentials on asymptotically flat 3-manifolds (M,g) in the nonvacuum setting. Heuristically, a static potential f represents the (signed) length along M of an irrotational timelike Killing vector field, which can degenerate on surfaces corresponding to the…
Study examines stability of MOTS in symmetric initial data sets.
problem Stability of marginally outer trapped surfaces in symmetric initial data sets.
method Characterizes instability through vector decomposition and properties of normal and tangent components.
result Characterizes instability of MOTS by the nature of zero sets and divergences.
Geometrically interpolates rigid body motions with initial and terminal twists.
problem Finding spatial trajectories between prescribed initial and terminal poses.
method Derives solutions for k-IV-TIP and k-BV-TIP for k=1,...,4.
result Automatic cubic interpolation identical to minimum acceleration curve when twists are zero.
Scale-free distributions and correlation functions found in financial data are reminiscent of the scale invariance of physical observables in the vicinity of a critical point. Here, we present empirical evidence for a transition phenomenon, accompanied by a symmetry breaking, in the investors' demand for stocks. We stu…
New assumptions and algorithm solve offline two-player zero-sum Markov games.
problem Solving offline two-player zero-sum Markov games under insufficient assumptions.
method Proposed unilateral concentration assumption and pessimism-type algorithm.
result Algorithm efficiently learns Nash equilibrium under unilateral concentration.
Study analyzes bond traders' views on equity market dynamics.
problem Understanding temporal shifts in equity market parameters.
method Utilizes Black-Derman-Toy model and zero-coupon bond pricing.
result Discovers correlations between risk-neutral probability and market variables.
We study the regularizing properties of complex Monge-Ampère flows on a Kähler manifold (X,ω) when the initial data are ω-psh functions with zero Lelong number at all points. We prove that the general Monge-Ampère flow has a solution which is immediately smooth. We also prove the uniqueness and stability of solutio…
SGD fails to converge for deep ReLU networks with limited random initializations.
problem SGD convergence in deep neural networks with limited random initializations.
method Analysis of four discretization parameters: network architecture, training data, gradient steps, and random initializations.
result SGD fails to converge for ReLU networks with depth much larger than width.
The Einstein equations in wave map gauge are a geometric second order system for a Lorentzian metric. To study existence of solutions of this hyperbolic quasi diagonal system with initial data on a characteristic cone which are not zero in a neighbourhood of the vertex one can appeal to theorems due to Cagnac and Dossa…
We extend the study of the vacuum Einstein constraint equations on manifolds with ends of cylindrical type initiated by Chruściel and Mazzeo by finding a class of solutions to the fully coupled system on such manifolds. We show that given a Yamabe positive metric g, which is conformally asymptotically cylindrical on ea…
New insights into how neural network initialization affects its performance.
problem Understanding how initialization impacts the generalization error of deep neural networks.
method By analyzing the NTK regime of DNNs, the study provides a quantitative answer to the impact of initialization on generalization error.
result Random initialization can increase generalization error, but antisymmetric initialization can mitigate this effect.
We study the strong predictable representation property in filtrations initially enlarged with a random variable L. We prove that the strong predictable representation property can always be transferred to the enlarged filtration as long as the classical density hypothesis of Jacod (1985) holds. This generalizes the ex…
Adversarial training achieves optimal test error for shallow networks.
problem Achieving optimal adversarial test error for general data distributions.
method Applying new Rademacher complexity bounds and properties of optimal adversarial predictors.
result Adversarial training can achieve optimal adversarial test error for general data distributions.
New technique trains deep neural networks without normalization or minibatch statistics.
problem Training deep neural networks at high learning rates without normalization.
method Channel-wise zero-mean initialization and gradient modification to maintain common mode rejection.
result Achieves higher accuracy compared to batch normalization and shows minibatches are unnecessary.
One of the mysteries in the success of neural networks is randomly initialized first order methods like gradient descent can achieve zero training loss even though the objective function is non-convex and non-smooth. This paper demystifies this surprising phenomenon for two-layer fully connected ReLU activated neural n…
We consider smooth complete solutions to Ricci flow with bounded curvature on manifolds without boundary in dimension three. Assuming an open ball at time zero of radius one has curvature bounded from below by -1, then we prove estimates which show that compactly contained subregions of this ball will be smoothed out b…
New approach to handle ranking function variation in zero-shot NAS.
problem Variation in ranking function outputs due to randomness.
method Viewing ranking function output as a random variable and constructing a stochastic ordering.
result Stochastic ordering boosts performance in neural architecture search.
Study rules out exotic S4 and #nCP2 construction using zero surgery homeomorphisms.
problem Tackles the possibility of constructing exotic S4 or #nCP2 using zero surgery homeomorphisms. method Uses zero surgery homeomorphisms to relate slice properties of knots stably after a connected sum with a 4-manifold.
result Rules out the possibility of constructing exotic S4 or #nCP2 using zero surgery homeomorphisms. We construct an algebra of smooth functions over the tangent groupoid associated to any Lie groupoid. This algebra is a field of algebras over the closed interval [0, 1] which fiber at zero is the algebra of Schwartz functions over the Lie algebroid, whereas any fiber out of zero is the convolution algebra of the initi…
The problem of resource allocation of nonlinear networked control systems is investigated, where, unlike the well discussed case of triggering for stability, the objective is optimal triggering. An approximate dynamic programming approach is developed for solving problems with fixed final times initially and then it is…