Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,694 papers · 148 categories

Trend · papers per month

3672108144 · May 202619922001200920172026
48 results for Sharp Edges

The study analyzes sharpness dynamics in neural networks, revealing mechanisms and conditions.

problem Understanding sharpness in neural network training.
method Fixed point analysis and edge of stability analysis in a simplified 2-layer linear network.
result Reveals mechanisms behind sharpness trends, conditions for edge of stability, and a period-doubling route to chaos.

Gradient descent at edge of stability stabilizes implicitly, following projected gradient descent.

problem Gradient descent's stability and sharpness behavior at the edge of instability.
method Cubic Taylor expansion analysis of gradient descent dynamics.
result Gradient descent at edge of stability implicitly follows projected gradient descent.

New findings show mini-batch SGD operates in a 'Edge of Stochastic Stability' regime.

problem Understanding the stability and convergence of mini-batch SGD.
method Analyzing the mini-batch Hessian and its directional curvature.
result Mini-batch SGD operates in a different stability regime (Edge of Stochastic Stability) compared to full-batch GD.

Gradient descent near stability threshold exhibits sharpness oscillations.

problem Understanding sharpness behavior near stability threshold in non-Euclidean norms.
method Interpreted EoS through Directional Smoothness and generalized sharpness under arbitrary norms.
result Non-Euclidean GD with generalized sharpness shows sharpness oscillations near 2/η2/η.

New insights into network generalization show learning rate affects both norm and sharpness.

problem Understanding the generalization of overparameterized networks.
method Empirical analysis and theoretical proof of the trade-off between norm and sharpness.
result Learning rate influences both norm and sharpness, neither alone minimizes generalization error.

Cohen et al. (2021) show GD trajectories align on a bifurcation diagram.

problem Understanding the Edge of Stability (EoS) phenomenon in gradient descent.
method Empirical studies and rigorous mathematical proofs for two-layer networks and single-neuron networks.
result GD trajectories align on a specific bifurcation diagram independent of initialization.

Gradient descent near stability threshold shows sharpness oscillations.

problem Understanding sharpness and stability in non-Euclidean norms during gradient descent.
method Interpreted EoS through Directional Smoothness, defined generalized sharpness for arbitrary norms.
result Non-Euclidean GD exhibits sharpness oscillations around the stability threshold.

Momentum affects optimization differently at small vs large batch sizes near instability.

problem Understanding how momentum impacts optimization near the edge of stability.
method Demonstrated through batch-size dependent behavior of SGD with momentum.
result Momentum operates in two distinct regimes: amplifying stochastic fluctuations at small batch sizes and stabilizing at large batch sizes.

Weight decay stabilizes training dynamics by slowing progressive sharpening.

problem Understanding how weight decay affects training stability in deep learning models.
method Analyzing weight decay effects at the Edge of Stability, developing a mathematical framework.
result Weight decay dampens oscillations and stabilizes sharpness in CNNs, causing a phase transition in MLPs.

Study financial contagion and risk in sparse networks with directed edges.

problem Analyzing systemic risk in sparse financial networks with balance-sheet interactions.
method Linear fraction of institutions with zero out-degree, sender-truncated subgraph G_sh, adversarial and random systemic events, explicit fan-in accumulation bound.
result Maximal forward reachability in G_sh is O(log n) with high probability in the subcritical regime, and multi-hit defaults are negligible in the supercritical regime.

Polygonal meshes provide an efficient representation for 3D shapes. They explicitly capture both shape surface and topology, and leverage non-uniformity to represent large flat regions as well as sharp, intricate features. This non-uniformity and irregularity, however, inhibits mesh analysis efforts using neural networ…

2018-09-16abs ↗pdf ↗

GD converges faster to flatter minima than gradient flow in shallow networks.

problem Understanding the dynamics of gradient descent in shallow linear networks.
method Analyzing the convergence rate and solution of gradient descent in depth-2 linear neural networks.
result GD converges linearly to flatter minima than gradient flow, even with large step sizes.

A labeled oriented graph (LOG) is an oriented graph with a labeling function from the edge set into the vertex set. The complexity of a LOG is the minimal cardinality of an initial set SS of vertices such that every vertex can be reached successively from SS only using edges with labels in SS or already visited vert…

2014-12-23abs ↗pdf ↗

Study reconstructs hidden perfect matchings in random graphs with specific edge weights.

problem Reconstructing hidden perfect matchings in random weighted bipartite graphs.
method Analyzes the maximum likelihood estimator for matching reconstruction under different probability distributions of edge weights.
result Sharp threshold and infinite-order phase transition in reconstruction error for different probability distributions.

We introduce a scalable measure of curvature for analyzing training dynamics of large language models.

problem Analyzing the training dynamics of large language models due to high computational cost of measuring Hessian sharpness.
method We introduce critical sharpness and relative critical sharpness as computationally efficient measures capturing Hessian sharpness phenomena.
result We provide the first demonstration of sharpness phenomena at scale up to 7B parameters.

New theory explains how chaotic training improves neural network generalization.

problem Understanding how chaotic training improves neural network generalization.
method Representing stochastic optimizers as random dynamical systems and introducing a new dimension concept.
result Generalization in chaotic training depends on the complete Hessian spectrum and partial determinants.

DIF extends NF with stochastic discrete latent variables for better density estimation.

problem Improving density estimation with discontinuities and fine details.
method Discretely indexed flows as an extension of Normalizing Flows with stochastic latent variables.
result DIF inherit good computational behavior of NF and can capture distributions with discontinuities.

The study characterizes 3-pseudomanifolds with up to two singularities.

problem Characterizing face-number-related invariants of normal 3-pseudomanifolds with up to two singularities.
method Proves properties of normal 3-pseudomanifolds using specific operations and upper bounds.
result Proves that normal 3-pseudomanifolds with up to two singularities are constructed from boundary complexes of 4-simplices.

A new ODE model explains gradient descent dynamics near edge of stability.

problem Understanding gradient-based training over non-convex landscapes.
method Rod Flow, a new ODE approximation of GD dynamics.
result Rod Flow accurately predicts critical sharpness threshold and self-stabilization in quartic potentials.

This paper resolves the all-or-nothing phase transition in graph matching.

problem Recovering vertex correspondence between edge-correlated random graphs.
method Analysis of mutual information, truncated second-moment computation, and maximum likelihood estimator.
result Sharp thresholds for correct matching in both dense and sparse graphs.

Stable commutator length scl_G(g) of an element g in a group G is an invariant for group elements sensitive to the geometry and dynamics of G. For any group G acting on a tree, we prove a sharp bound scl_G(g)>=1/2 for any g acting without fixed points, provided that the stabilizer of each edge is relatively torsion-fre…

2019-10-30abs ↗pdf ↗

We investigate slicings of combinatorial manifolds as properly embedded co-dimension 1 submanifolds. A focus is given to dimension 3 where slicings are normal surfaces. In the case of 2-neighborly 3-manifolds and quadrangulated slicings, a lower bound on the number of quadrilaterals of normal surfaces depending on the …

2010-04-06abs ↗pdf ↗

Detecting edge correlation between two graphs sharpens a threshold based on densest subgraph.

problem Detecting edge correlation between two Erdős-Rényi graphs.
method Formulated as a hypothesis testing problem, connecting to densest subgraph detection.
result Sharp information-theoretic threshold established for edge correlation detection.

Quantum stochastic walks optimize portfolios by leveraging financial networks, improving Sharpe ratios and reducing turnover.

problem Optimizing portfolios in noisy financial markets with superior risk-adjusted returns.
method Embed assets in a weighted graph, using quantum stochastic walks to derive optimal portfolio weights from the stationary distribution.
result Quantum stochastic walks can lift Sharpe ratios by up to 27% and reduce turnover from 480% to 2-90%.

Method solves optimisation problems on non-Riemannian surfaces with bilateral curvature bounds.

problem Optimisation problems on non-Riemannian surfaces with sharp edges.
method Forward-backward splitting in Alexandrov spaces with bilateral curvature bounds.
result Convergence of the forward-backward method in Alexandrov spaces with bilateral curvature bounds.

This paper tackles exact recovery of clusters in a stochastic Ising model on a SBM graph.

problem Recovering clusters in a stochastic Ising model on a SBM graph.
method Proposes a Stochastic Ising Block Model (SIBM) and establishes a sharp threshold for exact recovery.
result Sharp threshold mm^\ast for exact recovery of clusters in SIBM, with O(n)O(n) time complexity for mmm \ge m^\ast.

Background: Three-dimensional, whole heart, balanced steady state free precession (WH-bSSFP) sequences provide delineation of intra-cardiac and vascular anatomy. However, they have long acquisition times. Here, we propose significant speed ups using a deep learning single volume super resolution reconstruction, to reco…

2019-12-22abs ↗pdf ↗