Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,738 papers · 148 categories

Trend · papers per month

13.0%26.0%38.9%51.9% · Jun 202019922001200920172026
48 results for Scale critical initial data

Constructs initial data leading to apparent horizons and tests Penrose Inequality.

problem Testing Penrose Inequality in dynamical spacetimes.
method Scale critical initial data for Einstein vacuum system, constructing Cauchy data.
result Penrose Inequality holds in an open region of the future of initial data.

Study shows how electromagnetic and gravitational waves can form trapped surfaces.

problem Formation of trapped surfaces from initial data with electromagnetic fields.
method Established a scale-critical semi-global existence result from past null infinity for the Einstein-Maxwell system.
result Generalized approach for studying Einstein vacuum equations and extended a result to scale-critical regime.

Constructs foliations of critical surfaces for Hawking energy in asymptotically flat initial data sets.

problem Positivity and rigidity of Hawking quasi-local energy in asymptotically flat spacetimes.
method Lyapunov-Schmidt reduction within a Willmore-foliation framework.
result Existence and uniqueness of foliations by Hawking surfaces, positivity and large-sphere limit of Hawking energy.

Scale-free distributions and correlation functions found in financial data are reminiscent of the scale invariance of physical observables in the vicinity of a critical point. Here, we present empirical evidence for a transition phenomenon, accompanied by a symmetry breaking, in the investors' demand for stocks. We stu…

2001-11-19abs ↗pdf ↗

This work studies scaling laws for low-precision training in high-dimensional linear regression.

problem Optimizing trade-off between model quality and training costs in high-dimensional linear regression.
method Theoretical study of scaling laws for low-precision training within a high-dimensional sketched linear regression framework, analyzing multiplicative and additive quantization.
result Multiplicative quantization maintains full-precision model size, while additive quantization reduces effective model size.

Proves long-time Ricci flow existence and topological rigidity for pinched integral curvature manifolds.

problem Proving long-time existence and topological rigidity for manifolds with pinched scale-invariant integral curvature.
method Proves long-time existence of Ricci flow for manifolds with bounded curvature and pinched scale-invariant integral curvature, converging to a flat metric.
result Flow converges to a flat metric, implying topological rigidity of the manifold.

SGD transitions between maxima and minima with varying time scales.

problem Understanding SGD's behavior near critical points in noisy landscapes.
method Analyzing SGD convergence and escape dynamics in 1D landscapes with infinite- and finite-variance noise.
result SGD reliably moves to the basin's minimum unless close to a local maximum, where it can linger.

High-dimensional SGD limits show surprising dynamics and phase transitions.

problem Understanding SGD in high dimensions and its scaling limits.
method Proving limit theorems for SGD trajectories in high dimensions, choosing summary statistics, initialization, and step-size.
result Critical scaling regime for step-size, new correction term, and complex diffusive limits.

We consider the mean curvature evolution of rotationally symmetric surfaces. Using numerical methods, we detect critical behavior at the threshold of singularity formation resembling the one of gravitational collapse. In particular, the mean curvature simulation of a one-parameter family of initial data reveals the exi…

2009-03-19abs ↗pdf ↗

This paper solves tensor robust principal component analysis via scaled gradient descent.

problem Extracting useful information from tensor data robust to corruptions and ill-conditioning.
method Directly recovers low-rank tensor factors via scaled gradient descent with adaptive thresholding.
result The proposed algorithm converges linearly to the true low-rank tensor at a constant rate independent of the condition number.

We reparametrize ReLU NNs as splines to understand their learning dynamics.

problem Understanding the learning dynamics and inductive bias of neural networks.
method Reparametrize ReLU NNs as continuous piecewise linear splines to study learning dynamics.
result Standard weight initializations yield very flat functions, leading to strength and type of implicit regularization.

We show some results for the L2L^2 curvature flow linked by the theme of addressing collapsing phenomena. First we show long time existence and convergence of the flow for SO(3)SO(3)-invariant initial data on S3S^3, as well as a long time existence and convergence statement for three-manifolds with initial L2L^2 norm of c…

2012-01-05abs ↗pdf ↗

Dropout schedules can be optimized to significantly reduce model test loss.

problem Improving model performance in neural networks.
method Developed a mean-field theory of dropout at the edge of chaos, proposing front-loaded dropout schedules.
result Front-loaded dropout schedules reduce test loss by 18-35% over constant dropout.

New method proves instability of naked singularity and censors it.

problem Proving instability and censoring naked singularity.
method Einstein-scalar field system, hyperbolic short-pulse method, non-perturbative elliptic arguments.
result Tiny anisotropic perturbation leads to anisotropic apparent horizon censoring the naked singularity.

We study the evolution of wormhole geometries under Ricci flow using numerical methods. Depending on values of initial data parameters, wormhole throats either pinch off or evolve to a monotonically growing state. The transition between these two behaviors exhibits a from of critical phenomena reminiscent of that obser…

2008-08-06abs ↗pdf ↗

This work aims to create a large-scale model for critical care time series data.

problem Lack of large-scale datasets and distribution shifts in critical care time series data.
method Harmonized dataset creation and transfer learning research.
result Established a foundation for large-scale multi-variate time series models in critical care.

We determine the critical batch size for large language models and find it scales with data size, not model size.

problem Determining the optimal batch size for large-scale model training.
method We propose a measure of critical batch size, pre-trained models, and systematic hyper-parameter sweeps.
result The critical batch size scales primarily with data size, not model size.

A neural network model predicts the critical point of the Ising phase transition.

problem Predicting the critical point of the Ising phase transition using supervised learning.
method Proposed a minimal one-free-parameter neural network model to describe the supervised learning problem for the Ising model.
result Just one free parameter is enough to describe the universal finite-size-scaling function in the network output.

Two new scalable K-means initialization methods proposed for large-scale clustering.

problem Efficient initialization for large-scale clustering problems.
method Divide-and-conquer approach and random projection method for multiple lower-dimensional subspaces.
result The proposed methods outperform state-of-the-art in large-scale clustering tasks.

We develop a gluing construction which adds scaled and truncated asymptotically Euclidean solutions of the Einstein constraint equations to compact solutions with potentially non-trivial cosmological constants. The result is a one-parameter family of initial data which has ordinary and scaled "point-particle" limits an…

2009-08-12abs ↗pdf ↗

New method trains shallow neural networks with subquadratic width scaling.

problem Training shallow neural networks with optimal width scaling.
method Polyak-Lojasiewicz condition, smoothness, standard data assumptions, random matrix theory.
result Subquadratic scaling on network width with standard initialization strategies.

In this paper we are concerned with the learnability of energies from data obtained by observing time evolutions of their critical points starting at random initial equilibria. As a byproduct of our theoretical framework we introduce the novel concept of mean-field limit of critical point evolutions and of their energy…

2019-11-01abs ↗pdf ↗

The paper studies how curves evolve under area constraints and converges to a critical point.

problem Evolution of plane curves with fixed area under elastic energy gradient.
method Local and global existence of the flow, simplicity assumption, Łojasiewicz--Simon inequality.
result The evolving curve's length remains bounded and converges to a critical point.

Study shows formation of Kerr black holes with complete apparent horizons and proves Penrose inequalities.

problem Formation of Kerr black holes and Penrose inequalities.
method Combining gravitational-collapse and Kerr stability results with new coordinate changes and elliptic arguments.
result Proves dynamical and spacetime Penrose inequalities in black hole formation spacetimes.

Learning rate needs to decrease with higher data moments for effective ICA in high dimensions.

problem Slower convergence of ICA in high-dimensional data with high-order moments.
method High-dimensional ODE analysis of ICA algorithm under controlled moment structure.
result Critical learning rate threshold for effective ICA when moments are high.

Gradient descent with large steps leads to chaotic parameter space and unpredictable outcomes.

problem Understanding the behavior of gradient descent with large step sizes in matrix factorization.
method Analyzing the fractal structure of the parameter space and deriving critical step sizes for convergence.
result Gradient descent with large steps exhibits chaotic behavior and sensitivity to initialization, creating a fractal boundary between converging and diverging minimizers.

Theory explains deep nonlinear networks' plateaus and transitions.

problem Understanding long plateaus and feature acquisition transitions in deep nonlinear networks.
method Derived an exact identity for Frobenius norms, classified activation functions, and reduced matrix flow to a scalar ODE.
result Escape time law τ=Θ(ε(r2))τ_\star = Θ(\varepsilon^{-(r-2)}) for deep nonlinear networks, where rr is the number of bottleneck layers.

In massive open online courses (MOOCs), peer grading serves as a critical tool for scaling the grading of complex, open-ended assignments to courses with tens or hundreds of thousands of students. But despite promising initial trials, it does not always deliver accurate results compared to human experts. In this paper,…

2013-07-09abs ↗pdf ↗

Two-layer CNNs can overfit well if initialized correctly.

problem Understanding the conditions for benign overfitting in over-parameterized CNNs.
method Extending analysis to fully trainable two-layer CNNs, examining initialization scaling effects.
result Initialization scaling of the output layer is crucial; large scales lead to fixed output behavior, small scales to complex interactions.

Study reveals how initialization scale affects training accuracy in linear networks.

problem Understanding implicit bias in linear classification models.
method Asymptotic analysis of gradient flow trajectories and training loss minimization.
result Implicit bias is more complex at reasonable initialization scales and training accuracies.

This work analyzes actor-critic methods for faster convergence.

problem Finite-time analysis and sample complexity of two-time-scale actor-critic methods.
method Non-asymptotic analysis under non-i.i.d. setting, proving convergence to first-order stationary point.
result Actor-critic method finds a first-order stationary point with ildeO(ε2.5)\mathcal{ ilde{O}}(ε^{-2.5}) sample complexity.

Critical points of scale-invariant curvature energies in 4D are analytic.

problem Analyzing critical points of curvature energies in 4D manifolds.
method Applying Noether's theorem to identify conservation laws and lower order elliptic system of PDEs, then using integrability by compensation and interpolation theory.
result Critical points of scale-invariant curvature energies in 4D are analytic.

Robots learn new tasks autonomously with minimal human intervention.

problem Lack of scalable data collection for robot learning.
method Multi-task imitation learning with autonomous data collection and one-shot generalization.
result Robots can continuously improve through autonomous data collection without reinforcement learning.

Global stability proved for Navier-Stokes equations on hyperbolic space.

problem Stability of the Navier-Stokes equations on hyperbolic space.
method Proved global stability with exponential decay rate for small initial data.
result Exponential decay rate of $μλ_\Def^{(3)}$ for Navier-Stokes equations on hyperbolic space.

We introduce a scalable measure of curvature for analyzing training dynamics of large language models.

problem Analyzing the training dynamics of large language models due to high computational cost of measuring Hessian sharpness.
method We introduce critical sharpness and relative critical sharpness as computationally efficient measures capturing Hessian sharpness phenomena.
result We provide the first demonstration of sharpness phenomena at scale up to 7B parameters.

New method avoids spurious critical points for low-rank matrix recovery.

problem Low-rank matrix recovery problems on Riemannian manifold.
method Riemannian gradient descent with random initialization.
result Riemannian gradient descent avoids spurious critical points and converges nearly linearly.

New method diagnoses criticality in deep neural networks, improving performance.

problem Improving theoretical understanding and practical initialization of deep neural networks.
method Introducing partial Jacobians and deriving recurrence relations for their norms to analyze criticality.
result Proper stacking of LayerNorm and residual connections leads to a critical architecture for any initialization.

Neural networks compress and sample WDN contamination dynamics efficiently.

problem Infrastructure monitoring of complex, networked systems like water distribution networks is expensive and challenging.
method Developed Graph Fourier Transform (GFT) operators and neural networks (NN) for efficient data collection and inference.
result High accuracy reconstruction of contamination dynamics using only 5-10% of the sample set.

Investigates energy minimizers and critical points of scale-invariant tangent-point energies for knots.

problem Finding and characterizing minimizers and critical points of scale-invariant tangent-point energies for closed curves.
method Develops convergence and regularity theories based on fractional Sobolev spaces and new energy functionals.
result Minimizing sequences converge to locally critical embeddings in all but finitely many points, and locally critical embeddings are regular.

Investigates multifractal scaling in critical dynamics of random surfaces.

problem Analyzing multifractal scaling in critical dynamics of random surfaces.
method Examined multifractal scaling in various conformal field theories on random surfaces.
result Higher moments of time variations of the order parameter exhibit multifractal scaling.

Changing initialization scale affects deep model generalization, leading to memorization or improved performance.

problem Understanding how initialization scale impacts deep model generalization and memorization.
method Experimental setup with varying initialization scales, analysis of activation and loss functions, and development of an alignment measure.
result Increasing initialization scale leads to memorization, and decreasing it improves generalization, depending on activation and loss functions.

We investigate the combination of actor-critic reinforcement learning algorithms with uniform large-scale experience replay and propose solutions for two challenges: (a) efficient actor-critic learning with experience replay (b) stability of off-policy learning where agents learn from other agents behaviour. We employ …

2019-09-25abs ↗pdf ↗

The paper develops a theory linking pretraining and fine-tuning in neural networks.

problem Understanding how initialization choices impact feature learning and generalization in neural networks.
method Analytical theory of diagonal linear networks, deriving generalization error as a function of initialization parameters and task statistics.
result Different initialization choices place networks into four fine-tuning regimes with varying abilities to support feature learning and generalization.

Constructs initial data for Einstein vacuum equations involving multiple localized gravitational sources.

problem Modeling the interaction of distant gravitational systems in general relativity.
method Time-symmetric initial data construction using gluing schemes and localized sources.
result Produces initial data sets with finite ADM mass and multiple Einstein-Rosen bridges.

Paper analyzes convergence rates of two time-scale AC and NAC algorithms.

problem Finite-sample convergence rate analysis of two time-scale AC and NAC algorithms.
method Developed novel techniques for bias error and convergence rate analysis.
result Established non-asymptotic convergence rates for two time-scale AC and NAC.