SAM improves neural network generalization by penalizing sharpness, clarifying its exact notion and mechanism.
problem Improving deep neural network generalization for various settings.
method Sharpness-Aware Minimization (SAM) technique that penalizes a notion of sharpness of the model.
result SAM regularizes the third notion of sharpness, most likely preferred for practical performance.
Sharp bounds on diameter and eigenvalues for amply regular graphs.
problem Finding bounds for amply regular graphs' diameter and eigenvalues.
method New ideas relating discrete Ricci curvature to local matching properties, including a novel construction of a regular bipartite graph.
result Sharp diameter and eigenvalue bounds for amply regular graphs.
New insights into network generalization show learning rate affects both norm and sharpness.
problem Understanding the generalization of overparameterized networks.
method Empirical analysis and theoretical proof of the trade-off between norm and sharpness.
result Learning rate influences both norm and sharpness, neither alone minimizes generalization error.
A new tradeoff between regularization and sharpness improves model performance in overparameterized settings.
problem Improving model performance in overparameterized settings with minimum-norm interpolators.
method Proposes a regularization-sharpness tradeoff for overparameterized linear regression with an ℓ^p penalty.
result Empirical validation shows the tradeoff terms can distinguish performant linear interpolators.
Sharp regularity for Pfaff system leads to isometric immersions in arbitrary dimensions.
problem Existence and regularity of isometric immersions in arbitrary dimensions.
method Proving W 1 , 2 W^{1,2} W 1 , 2 -regularity for Pfaff system with antisymmetric L 2 L^2 L 2 -coefficient matrix. result Equivalence between W 2 , 2 W^{2,2} W 2 , 2 -isometric immersions and weak solubility of Gauss--Codazzi--Ricci equations. Study on convergence rates for optimal transport with regularization.
problem Convergence analysis of divergence-regularized optimal transport.
method Novel methodology using quantization and martingale couplings.
result Sharp rates for various divergences and transport costs.
The study examines MCMC methods for arbitrary objectives and finds likelihood sharpness impacts performance and regularization.
problem Limitations of MCMC methods for arbitrary objective functions.
method Two-block MCMC framework with Metropolis-Hastings and Gibbs sampling, exploring likelihood curvature and sharpness.
result Likelihood sharpness governs in-sample performance and regularization inferred by training data.
New algorithm avoids spurious sharpness minimization for NLP models.
problem SAM fails in NLP, leading to performance degradation.
method Developed Functional-SAM, which modifies logit statistics instead of function geometry.
result Functional-SAM and combined methods outperform AdamW and SAM in NLP tasks.
SAM improves deep learning tasks by promoting balancedness, reducing outlier impact.
problem Improving generalization in deep learning tasks, especially with scale-invariant problems.
method Introduces balancedness as a new concept to depict global behaviors of SAM, focusing on the difference between squared norms of two variables.
result SAM promotes balancedness and is data-responsive, outperforming SGD in outlier scenarios.
Sharp results link DLN gradient flow to basis pursuit optimization and GHA phase transitions.
problem Understanding implicit regularization in Diagonal Linear Networks.
method Sharp convergence bounds and characterization of ℓ 1 \ell_1 ℓ 1 minimizers. result Gradient flow of DLNs with tiny initialization approximates minimizers of basis pursuit optimization problem.
Enhances deep learning by boosting generalization and convergence.
problem Improving generalization and convergence in deep learning models.
method Implicit Regularization Enhancement (IRE) framework that decouples flat and sharp directions.
result IRE consistently improves generalization performance across various deep learning tasks and models.
Deep linear networks minimize sharpness, avoiding large eigenvalues.
problem Understanding optimization dynamics in deep linear networks for regression.
method Analyzing sharpness (largest eigenvalue of Hessian) of minimizers and gradient flow solutions.
result Gradient flow implicitly regularizes towards flat minima, with sharpness bounded by a constant.
Sharp Hölder regularity found for complex Frobenius theorem coordinates.
problem Finding optimal Hölder-Zygmund regularity for complex Frobenius theorem coordinates.
method Analyzing necessary and sufficient conditions for coordinate charts achieving the theorem's structure.
result The optimal Hölder-Zygmund regularity for coordinate charts is shown to be α α α . Sharp proof of sub-Riemannian length-minimizing curves being at least C 2 C^2 C 2
problem Smoothness of sub-Riemannian length-minimizing curves
method Study of a class of sub-Riemannian structures, proving C 2 C^2 C 2 regularity result Theorem 1.1 in [6] is sharp
Non-convex regularizers usually improve the performance of sparse estimation in practice. To prove this fact, we study the conditions of sparse estimations for the sharp concave regularizers which are a general family of non-convex regularizers including many existing regularizers. For the global solutions of the regul…
Sharp threshold found for metric uniqueness in Riemannian Calderón-type problems.
problem Determining metrics uniquely from Dirichlet-to-Neumann maps in Riemannian Schrödinger problems.
method Adaptation of Lassas-Uhlmann reconstruction theorem and novel Gevrey space techniques.
result Analytic metrics uniquely determine the metric up to boundary-preserving diffeomorphisms, but non-analytic metrics are not uniquely determined.
The paper examines how ESG constraints affect portfolio optimization in large datasets.
problem Investment optimization with ESG constraints in large portfolios.
method Asymptotic analysis of out-of-sample Sharpe ratio, regularization matrix estimation, and adaptive portfolio selection.
result The proposed adaptive ESG-constrained portfolio yields a high out-of-sample Sharpe ratio while meeting ESG requirements.
Sharp Lipschitz bounds for flow-matching and diffusion models with optimal sampling rates.
problem Establishing optimal Lipschitz regularity for flow-matching and diffusion models.
method Sharp Lipschitz regularity theory for flow-matching vector fields and diffusion-model scores.
result Achieves optimal sampling rate of d / N \sqrt{d}/N d / N for Euler-type samplers in dimension d d d . Sharp inequality for p p p -harmonic maps with new optimal constant.
problem Deriving the sharp vectorial Kato inequality for p p p -harmonic mappings. method Analyzing the inequality for p p p -harmonic mappings and comparing with scalar valued cases. result Established the optimal constant for p p p -harmonic maps and enhanced the range of p p p values for regularity. Sharp 3D Alexandrov inequality applied to volume-preserving flows.
problem Volume-preserving geometric flows in 3D space.
method Sharp quantitative Alexandrov inequality for C 2 C^2 C 2 -regular sets. result Established a 3D sharp quantitative version of the Alexandrov inequality.
On simple geodesic disks of constant curvature, we derive new functional relations for the geodesic X-ray transform, involving a certain class of elliptic differential operators whose ellipticity degenerates normally at the boundary. We then use these relations to derive sharp mapping properties for the X-ray transform…
Sharp analysis improves RLHF sample complexity with KL-regularization.
problem Improving RLHF sample complexity with KL-regularization.
method Sharp analysis of KL-regularized contextual bandits and RLHF.
result Achieved an O(1/ε) sample complexity when ε is sufficiently small.
We prove the two theorems of the title, settling two long standing questions in the local theory of singular minimal hypersurfaces. The sharpness of either result is with respect to its hypothesis on the size of the allowable singular sets. The proofs of both theorems rely heavily on the author's recent regularity and …
Sharp generalization of boundary regularity for area minimizing currents with arbitrary multiplicity.
problem Boundary regularity of area minimizing currents with multiplicity.
method Sharp generalization of Allard's boundary regularity theorem to higher multiplicity settings.
result The set of density Q / 2 Q/2 Q /2 singular boundary points of T T T is H m − 3 \mathcal{H}^{m-3} H m − 3 -rectifiable. Classifies regularity for Lagrangian mean curvature type equations.
problem Classifying regularity for Lagrangian mean curvature type equations.
method Generalized constant rank theorem for Legendre transform, constructed convex solutions, and showed regularity conditions.
result Optimal regularity conditions for Lagrangian mean curvature type equations.
Researchers extend regularity of p p p -harmonic maps into spheres for a new range of p p p .
problem Establishing regularity of p p p -harmonic maps for a broader range of p p p . method Combining Morrey's methods with Hardt and Lin's Extension Theorem, and proving a sharp Kato inequality.
result Regularity for p ∈ [ 2.961 , 3 ] p \in [2.961, 3] p ∈ [ 2.961 , 3 ] and p ∈ [ 2 , p 0 ] p \in [2, p_0] p ∈ [ 2 , p 0 ] with p 0 ≈ 2.366 p_0 \approx 2.366 p 0 ≈ 2.366 . According to the classical Plante-Thurston Theorem, all nilpotent groups of C 2 C^2 C 2 -diffeomorphisms of the closed interval are Abelian. Using techniques coming from the works of Denjoy and Pixton, Farb and Franks constructed a faithful action by C 1 C^1 C 1 -diffeomorphisms of [ 0 , 1 ] [0,1] [ 0 , 1 ] for every finitely-generated, torsion-free,…
A Jacobi structure J J J on a line bundle L → M L\to M L → M is weakly regular if the sharp map J ♯ : J 1 L → D L J^\sharp : J^1 L \to DL J ♯ : J 1 L → D L has constant rank. A generalized contact bundle with regular Jacobi structure possess a transverse complex structure. Paralleling the work of Bailey in generalized complex geometry, we find condition on a pair …
Sharp curvature estimates lead to optimal C 1 , 1 C^{1,1} C 1 , 1 regularity for L p L_p L p Minkowski problems.
problem Optimal regularity for solutions to L p L_p L p Minkowski problems. method Anisotropic Gauss curvature flows and curvature estimates.
result Sharp C 1 , 1 C^{1,1} C 1 , 1 regularity for solutions to L p L_p L p Minkowski problems. In this paper, we give a new sharp generalization bound of lp-MKL which is a generalized framework of multiple kernel learning (MKL) and imposes lp-mixed-norm regularization instead of l1-mixed-norm regularization. We utilize localization techniques to obtain the sharp learning rate. The bound is characterized by the d…
Regularity properties of intrinsic objects for a large class of Stein Manifolds, namely of Monge-Ampère exhaustions and Kobayashi distance, is interpreted in terms of modular data. The results lead to a construction of an infinite dimensional family of convex domains with squared Kobayashi distance of prescribed regula…
In this paper, we investigate a multivariate multi-response (MVMR) linear regression problem, which contains multiple linear regression models with differently distributed design matrices, and different regression and output vectors. The goal is to recover the support union of all regression vectors using l 1 / l 2 l_1/l_2 l 1 / l 2 -reg…
The paper studies the loss landscape of regularized deep matrix factorization, revealing unique and sharp minimizers.
problem Understanding the loss landscape and minimizers of regularized deep matrix factorization problems.
method Theoretical analysis of ℓ 2 \ell^2 ℓ 2 -regularized deep matrix factorization/deep linear network training problems with squared-error loss. result The unique end-to-end minimizer exists for all target matrices except for a set of Lebesgue measure zero.
WARPd method solves inverse problems with approximate sharpness conditions.
problem Reconstruction of signals from undersampled and noisy measurements.
method First-order method based on primal-dual iterations with restart-reweight scheme.
result WARPd achieves stable linear convergence under generic approximate sharpness condition.
Sharp log-Sobolev inequalities proved for C D ( 0 , N ) {\sf CD}(0,N) CD ( 0 , N ) spaces.
problem Proving log-Sobolev inequalities in noncompact metric measure spaces.
method Sharp isoperimetric inequality, symmetrisation, scaling argument, Hamilton-Jacobi inequality, Sobolev regularity.
result Sharp log-Sobolev inequalities established in C D ( 0 , N ) {\sf CD}(0,N) CD ( 0 , N ) spaces. SAM improves generalization in overparameterized models, but its behavior in tensorized models is less understood.
problem Understanding the implicit regularization of SAM in tensorized models.
method Scale-invariance analysis and gradient flow analysis to derive Norm Deviation as a measure of core norm imbalance, and propose Deviation-Aware Scaling (DAS).
result DAS achieves competitive or improved performance over SAM, while offering reduced computational overhead.
Sharp Sobolev theory for scalar elliptic equations on minimal regular manifolds.
problem Well-posedness and regularity for scalar elliptic equations on manifolds of minimal regularity.
method Localization and flat domain techniques combined with Calderón–Zygmund theory and Fredholm alternative.
result Sharp L p L^p L p -based Sobolev regularity for scalar elliptic problems on manifolds of minimal regularity. Sharp Gaussian isoperimetry proven along Ricci flow.
problem Proving sharp Gaussian isoperimetric inequality for Ricci flow.
method Using monotonicity formula to prove inequality.
result Exact Gaussian enlargement theorem and concentration estimates.
Weight decay stabilizes training dynamics by slowing progressive sharpening.
problem Understanding how weight decay affects training stability in deep learning models.
method Analyzing weight decay effects at the Edge of Stability, developing a mathematical framework.
result Weight decay dampens oscillations and stabilizes sharpness in CNNs, causing a phase transition in MLPs.
We derive a selection of energy estimates for a generalisation of a critical equation on the unit disc in R 2 \mathbb{R}^2 R 2 introduced by Rivière. Applications include sharp regularity results and compactness theorems which generalise a large amount of previous geometric PDE theory, including some of the theory of harmoni…
GD converges faster to flatter minima than gradient flow in shallow networks.
problem Understanding the dynamics of gradient descent in shallow linear networks.
method Analyzing the convergence rate and solution of gradient descent in depth-2 linear neural networks.
result GD converges linearly to flatter minima than gradient flow, even with large step sizes.
The paper explores sharp isoperimetric properties on non-compact spaces with Ricci bounds.
problem Sharp isoperimetric properties on non-compact spaces with Ricci bounds.
method Sharp isoperimetric comparison theorems and asymptotic isoperimetric properties.
result Almost regularity theorems and enhanced functional inequalities.
We establish sharp regularity and Fredholm theorems for the \bar{\partial}_b-Neumann problem on domains satisfying some non-generic geometric conditions. We use these domains to construct explicit examples of bad behaviour of the Kohn Laplacian: it is not always hypoelliptic up to the boundary, its partial inverse is n…
Paper proposes a method to estimate scientific parameters in hybrid models without relying on model architecture.
problem Estimating unknown parameters in hybrid models combining machine learning and scientific models.
method Sharpness-aware minimization adapted for hybrid modeling, focusing on model simplicity.
result Demonstrates effectiveness of SAM-based hybrid model learning for scientific parameter estimation.
We prove a sharp version of the Hopf boundary point lemma for Black-Scholes type equations. We also investigate the existence and the regularity of the spatial derivative of the solutions at the spatial boundary.
Paper analyzes sample complexity for offline f f f -divergence-regularized contextual bandits.
problem Lack of tight analyses for sample complexity in offline reinforcement learning.
method Novel pessimism-based analysis for reverse KL divergence, establishing i l d e O ( ε − 1 ) ilde{O}(ε^{-1}) i l d e O ( ε − 1 ) sample complexity. result Achieves i l d e O ( ε − 1 ) ilde{O}(ε^{-1}) i l d e O ( ε − 1 ) sample complexity for reverse KL divergence, surpassing existing bounds. Sharp bounds for spanning tree entropy in planar lattices.
problem Estimating spanning tree entropy in planar lattice graphs.
method Using hyperbolic geometry and polyhedra volumes.
result Proved bounds are easy to compute and provide excellent estimates.
The paper extends submanifold reach to less regular C 1 , α C^{1,α} C 1 , α classes.
problem Extending submanifold reach to less regular classes.
method Using μ μ μ -reach and C 1 , α C^{1,α} C 1 , α regularity to quantify reach. result Intermediate regularities C 1 , α C^{1,α} C 1 , α induce quantitative results on reach.