Sharpness minimization algorithms don't solely improve generalization.
problem Why do overparameterized neural networks generalize?
method Theoretical and empirical investigation of two-layer ReLU networks.
result Sharpness minimization algorithms do not always lead to better generalization.
ASAM improves deep neural network generalization by adapting sharpness to scale.
problem Fixed-radius sharpness measure is sensitive to parameter scaling, weakening its connection to generalization.
method Introduces adaptive sharpness, a scale-invariant measure, and proposes ASAM for deep learning.
result ASAM significantly improves model generalization performance across various datasets.
DGSAM improves domain generalization by minimizing individual sharpness.
problem Improving domain generalization models that perform well on unseen target domains.
method Shifts DG paradigm toward minimizing individual sharpness across source domains.
result DGSAM reduces performance variance across domains with less computational overhead.
Kronheimer and Mrowka introduced a new knot invariant, called s♯, which is a gauge theoretic analogue of Rasmussen's s invariant. In this article, we compute Kronheimer and Mrowka's invariant for some classes of knots, including algebraic knots and the connected sums of quasi-positive knots with non-trivial r…
SAM improves neural network generalization by penalizing sharpness, clarifying its exact notion and mechanism.
problem Improving deep neural network generalization for various settings.
method Sharpness-Aware Minimization (SAM) technique that penalizes a notion of sharpness of the model.
result SAM regularizes the third notion of sharpness, most likely preferred for practical performance.
SALR improves deep learning generalization by dynamically adjusting learning rates.
problem Improving generalization in deep learning models.
method Sharpness-aware learning rate scheduling based on local loss function sharpness.
result SALR drives solutions to flatter regions, improving generalization and convergence.
SharpBalance improves deep ensemble performance by balancing sharpness and diversity.
problem Improving deep ensemble performance in both in-distribution and out-of-distribution scenarios.
method Introducing SharpBalance, a novel training approach that balances sharpness and diversity within ensembles.
result SharpBalance effectively improves the sharpness-diversity trade-off and ensemble performance in ID and OOD scenarios.
New insights into network generalization show learning rate affects both norm and sharpness.
problem Understanding the generalization of overparameterized networks.
method Empirical analysis and theoretical proof of the trade-off between norm and sharpness.
result Learning rate influences both norm and sharpness, neither alone minimizes generalization error.
LSAM optimizes deep learning training with improved efficiency.
problem Inefficiency in distributed large-batch training with Sharpness-Aware Minimization (SAM).
method Integrates SAM's adversarial steps with an asynchronous distributed sampling strategy.
result Higher final accuracy compared to data-parallel SAM.
The Sharpe ratio is a way to compare the excess returns (over the risk free asset) of portfolios for each unit of volatility that is generated by a portfolio. In this paper we introduce a robust Sharpe ratio portfolio under the assumption that the risk free asset is unknown. We propose a robust portfolio that maximizes…
Sharp-MAML improves MAML by reducing saddle points in few-shot learning.
problem Challenges in optimizing MAML due to complex loss landscape.
method Sharpness-aware minimization applied to MAML.
result Sharp-MAML and its variant outperform plain MAML on few-shot learning tasks.
We prove sharp bounds for the growth rate of eigenfunctions of the Ornstein-Uhlenbeck operator and its natural generalizations. The bounds are sharp even up to lower order terms and have important applications to geometric flows.
Sharp spectral gap estimates on manifolds with integral curvature bounds.
problem Proving spectral gap estimates on manifolds with integral curvature bounds.
method Generalizing previous results to include integral curvature bounds.
result Confirms a conjecture about spectral gap estimates on manifolds with integral curvature bounds.
Truncated SGD with heavy-tailed noise eliminates sharp local minima.
problem Avoiding sharp local minima in deep learning models.
method Truncated SGD with heavy-tailed gradient noise.
result Truncated SGD can eliminate sharp local minima entirely from its training trajectory.
In Deep Learning, Stochastic Gradient Descent (SGD) is usually selected as a training method because of its efficiency; however, recently, a problem in SGD gains research interest: sharp minima in Deep Neural Networks (DNNs) have poor generalization; especially, large-batch SGD tends to converge to sharp minima. It bec…
Sharp spectral gap estimates for higher-order operators on hyperbolic spaces.
problem Estimating spectral gaps for higher-order operators on Cartan-Hadamard manifolds.
method Symmetrization-free proofs based on general functional inequalities.
result Solves a sharp asymptotic problem from Cheng and Yang and answers a question from Kristály.
Generalized Blaschke rolling theorem for curved spaces.
problem Extending classical theorem to curved spaces.
method Generalization to Riemannian manifolds with bounded curvature.
result Sharp results in arbitrary dimensions, new even in constant curvature spaces.
Paper generalizes complex Brunn-Minkowski theory and proves new extension theorems.
problem Complex Brunn-Minkowski theory and extension theorems.
method Hilbert bundle approach to complex Brunn-Minkowski theory.
result Generalizes Guan's sharp strong openness theorem and sharp Ohsawa-Takegoshi extension theorem.
We discuss - in what is intended to be a pedagogical fashion - generalized "mean-to-risk" ratios for portfolio optimization. The Sharpe ratio is only one example of such generalized "mean-to-risk" ratios. Another example is what we term the Fano ratio (which, unlike the Sharpe ratio, is independent of the time horizon)…
Sharp estimate for flow in any dimension.
problem Interior gradient estimate for graphical mean curvature flow.
method Proving sharp interior gradient estimate for area decreasing graphical mean curvature flow in arbitrary codimension.
result Generalized result in arbitrary codimension.
Sharp estimate shown to be rigid on curved surfaces.
problem Sharp Bezout estimate on nonnegatively curved Riemann surfaces.
method General three circle theorem applied.
result Rigidity of the sharp Bezout estimate.
Enhances deep learning by boosting generalization and convergence.
problem Improving generalization and convergence in deep learning models.
method Implicit Regularization Enhancement (IRE) framework that decouples flat and sharp directions.
result IRE consistently improves generalization performance across various deep learning tasks and models.
Sharp spectral theorem splits certain non-compact manifolds.
problem Proving spectral splitting for non-compact manifolds with specific curvature conditions.
method Sharp spectral analysis and geometric splitting theorem.
result Non-compact manifolds split as RimesN under given curvature constraints. Sharp criterion for Chern-Gauss-Bonnet integral using Q curvature.
problem Quantifying the Chern-Gauss-Bonnet integral using Q curvature.
method New approach involving singular integral estimation.
result Derivation of asymptotic formula for Q curvature equation.
Sharpe ratio is widely used in asset management to compare and benchmark funds and asset managers. It computes the ratio of the excess return over the strategy standard deviation. However, the elements to compute the Sharpe ratio, namely, the expected returns and the volatilities are unknown numbers and need to be esti…
Gradient descent near stability threshold exhibits sharpness oscillations.
problem Understanding sharpness behavior near stability threshold in non-Euclidean norms.
method Interpreted EoS through Directional Smoothness and generalized sharpness under arbitrary norms.
result Non-Euclidean GD with generalized sharpness shows sharpness oscillations near 2/η. The study analyzes sharpness dynamics in neural networks, revealing mechanisms and conditions.
problem Understanding sharpness in neural network training.
method Fixed point analysis and edge of stability analysis in a simplified 2-layer linear network.
result Reveals mechanisms behind sharpness trends, conditions for edge of stability, and a period-doubling route to chaos.
Sharp boundaries for detecting dense subhypergraphs established.
problem Detecting dense subhypergraphs in random hypergraphs.
method Established sharp detection boundaries for known and unknown edge probabilities.
result Sharp detectable regions differ significantly from graph counterparts.
Gradient descent near stability threshold shows sharpness oscillations.
problem Understanding sharpness and stability in non-Euclidean norms during gradient descent.
method Interpreted EoS through Directional Smoothness, defined generalized sharpness for arbitrary norms.
result Non-Euclidean GD exhibits sharpness oscillations around the stability threshold.
Sharp bounds on mean curvature and geodesic lengths in convex hypersurfaces.
problem Finding sharp bounds on total mean curvature of convex hypersurfaces.
method Sharp lower bounds for mean width and Birkhoff invariant, characterizing spheres.
result Generalization of Álvarez Paiva's result to convex hypersurfaces.
Sharp curvature estimates for mean curvature flow in spheres.
problem Understanding the behavior of surfaces evolving under mean curvature flow in spheres.
method Proving asymptotically sharp curvature pinching estimates and using them to derive derivative and convexity estimates.
result Partial classification of singularity models and new rigidity results for ancient solutions.
Sharp results link DLN gradient flow to basis pursuit optimization and GHA phase transitions.
problem Understanding implicit regularization in Diagonal Linear Networks.
method Sharp convergence bounds and characterization of ℓ1 minimizers. result Gradient flow of DLNs with tiny initialization approximates minimizers of basis pursuit optimization problem.
Sharp inequalities in unit ball with constraints on moments.
problem Establishing Sobolev trace inequalities with constraints.
method Constructing smooth test functions for higher order moments.
result Almost optimal Sobolev trace inequalities for 2nd and 4th orders.
Improves model generalization by minimizing loss sharpness.
problem Overparameterized models often fail to generalize well despite low training loss.
method Sharpness-Aware Minimization (SAM) minimizes both loss value and sharpness.
result SAM improves model generalization across various datasets and models.
Sharp bounds on scalar curvature spectrum and rigidity theorems.
problem Understanding scalar curvature bounds and rigidity on manifolds.
method Sharp upper bounds for the bottom spectrum of the Beltrami Laplacian, scalar curvature rigidity theorem.
result Sharp upper bound for the bottom spectrum of the Beltrami Laplacian and scalar curvature rigidity theorem.
Sharp comparison for sub-Gaussian random variables in convex order.
problem Comparing sub-Gaussian random variables in convex order.
method Proving dominance using moment generating functions and convex functions.
result Sharp comparison established between specific sub-Gaussian random variables.
Sharp inequality on Siegel domain involving weighted norms and sub-Laplacian.
problem Establishing a Sobolev trace inequality on a specific domain.
method Using weighted norms and fractional powers of sub-Laplacian on Heisenberg group.
result Sharp Sobolev trace inequality on Siegel domain involving weighted norms.
Averaged SGD optimizes a smoothed objective, leading to better generalization.
problem Improving generalization performance in machine learning models.
method Analyzed the smoothed objective function of SGD and proved that averaged SGD can optimize this smoothed function efficiently.
result Averaged SGD can efficiently optimize a smoothed objective, leading to better generalization.
Monge SAM improves deep learning by making sharpness-aware minimization invariant to reparametrizations.
problem Non-invariance of sharpness-aware minimization (SAM) to reparametrizations.
method Introduces Monge SAM, a reparametrization-invariant version of SAM using a Riemannian metric.
result Monge SAM enhances robustness and generalization compared to previous methods.
Sharp inequality found on three-balls for fourth order Sobolev traces.
problem Fourth order Sobolev trace inequality on three-balls.
method Established through equivalence to a third order Sobolev inequality on two-spheres.
result Sharp fourth order Sobolev trace inequality on three-balls.
This work analyzes statistical properties of SAM, showing it outperforms GD.
problem Improving deep neural network generalization through flatter solutions.
method Directly studies statistical performance of Sharpness-Aware Minimization (SAM).
result SAM has smaller prediction error than Gradient Descent (GD) under certain conditions.
Sharp characterization of Willmore invariant in higher dimensions.
problem Understanding the Willmore invariant in various dimensions.
method Characterization using conformal fundamental forms and tensors.
result Sharp sufficient condition for vanishing Willmore invariant in even dimensions.
In this paper, we extend the sharp lower bounds of spectal gap, due to Chen- Wang [10, 11], Bakry-Qian [6] and Andrews-Clutterbuck [5], from smooth Riemaniannian manifolds to general metric measure spaces with Riemannian curvature-dimension condition RCD*(K;N).
Sharp heat kernel estimates on manifolds lead to solutions of the Parabolic Anderson model.
problem Well-posedness and intermittency of solutions to the Parabolic Anderson model on Riemannian manifolds.
method Sharp global heat kernel bounds and geodesic comparison geometry.
result Upper and lower moment bounds for solutions of the Parabolic Anderson model on general compact Riemannian manifolds.
Sharp inequalities for star bodies in 2D space.
problem Understanding star bodies in 2D space.
method Sharp inequalities for star bodies in R2. result New inequalities and proofs for star bodies.
Sharp estimates for mean curvature flow of graphs are shown and examples are given to illustrate why these are sharp. The estimates improves earlier (non-sharp) estimates of Klaus Ecker and Gerhard Huisken.
DASH improves ensemble generalizability by encouraging diverse, flat loss landscapes.
problem Improving generalization and robustness of deep ensembles.
method DASH promotes diversity and flatness in deep ensembles by encouraging base learners to move towards low-loss regions of minimal sharpness.
result DASH improves ensemble generalizability, as demonstrated by extensive empirical evidence.
Sharpe ratio (sometimes also referred to as information ratio) is widely used in asset management to compare and benchmark funds and asset managers. It computes the ratio of the (excess) net return over the strategy standard deviation. However, the elements to compute the Sharpe ratio, namely, the expected returns and …