The study analyzes sharpness dynamics in neural networks, revealing mechanisms and conditions.
problem Understanding sharpness in neural network training.
method Fixed point analysis and edge of stability analysis in a simplified 2-layer linear network.
result Reveals mechanisms behind sharpness trends, conditions for edge of stability, and a period-doubling route to chaos.
We introduce a scalable measure of curvature for analyzing training dynamics of large language models.
problem Analyzing the training dynamics of large language models due to high computational cost of measuring Hessian sharpness.
method We introduce critical sharpness and relative critical sharpness as computationally efficient measures capturing Hessian sharpness phenomena.
result We provide the first demonstration of sharpness phenomena at scale up to 7B parameters.
A simple function shows how neural nets can converge despite high sharpness.
problem Understanding why neural nets converge with high sharpness.
method Constructed a minimal example function and analyzed its training dynamics rigorously.
result Final converging point has sharpness close to 2/η. SALR improves deep learning generalization by dynamically adjusting learning rates.
problem Improving generalization in deep learning models.
method Sharpness-aware learning rate scheduling based on local loss function sharpness.
result SALR drives solutions to flatter regions, improving generalization and convergence.
Develops a dynamical method to prove the sharp Berezin-Li-Yau inequality.
problem Proving the sharp Berezin-Li-Yau inequality for convex domains.
method Volume-preserving mean curvature flow and a new monotonicity principle.
result Shows the sharp Berezin-Li-Yau bound for every smooth convex domain.
New couplings improve understanding of molecular dynamics convergence.
problem Understanding convergence of Andersen dynamics in high dimensions.
method Presented couplings to obtain sharp convergence bounds in the Wasserstein sense.
result Sharp convergence bounds in the Wasserstein sense without global convexity.
Noise in linear networks minimizes sharpness and leads to shrinkage-thresholding.
problem Minimizing sharpness in diagonal linear networks.
method Stochastic sharpness-aware minimization (SAM) with isotropic noise.
result Noise forces shrinkage-thresholding of true parameters.
In this paper we investigate the expected terminal utility maximization approach for a dynamic stochastic portfolio optimization problem. We solve it numerically by solving an evolutionary Hamilton-Jacobi-Bellman equation which is transformed by means of the Riccati transformation. We examine the dependence of the resu…
Deep RL optimizes dynamic portfolio weights in China's stock market.
problem Traditional portfolio optimization methods struggle with dynamic asset weight adjustments.
method Developed a deep reinforcement learning framework with novel reward functions and random sampling.
result Model outperforms traditional methods in portfolio optimization and risk mitigation.
VAE struggles with distribution class sharpness, which can be learned dynamically.
problem VAE struggles with distribution class sharpness, leading to blurry reconstructions.
method Established that VAE fails if distribution class sharpness does not match data scale. Suggested learning distribution sharpness dynamically.
result Dynamic adjustment of distribution sharpness improves VAE performance, escaping local optima.
Truncated SGD with heavy-tailed noise eliminates sharp local minima.
problem Avoiding sharp local minima in deep learning models.
method Truncated SGD with heavy-tailed gradient noise.
result Truncated SGD can eliminate sharp local minima entirely from its training trajectory.
Weight decay stabilizes training dynamics by slowing progressive sharpening.
problem Understanding how weight decay affects training stability in deep learning models.
method Analyzing weight decay effects at the Edge of Stability, developing a mathematical framework.
result Weight decay dampens oscillations and stabilizes sharpness in CNNs, causing a phase transition in MLPs.
Study reveals sharp characterisation of local minima in neural network loss landscapes.
problem Characterizing local minima in high-dimensional two-layer ReLU neural networks.
method Exact low-dimensional representation of local minima using summary statistics and link with one-pass SGD dynamics.
result Local minima in overparameterized neural networks form discrete families with varying stability and reachability.
Optimal option portfolios under Sharpe Ratio maximization with skew-elliptical t-distributed returns
problem Optimal option portfolios under Sharpe Ratio maximization
method Formulation for explicit portfolio weights
result Different optimal portfolios for Sharpe Ratio and return-to-Value-at-Risk (VaR) ratio
Sharp changes in time series representing market dynamics are studied by means of the self--similar analysis suggested earlier by the authors. These sharp changes are market booms and crashes. Such crises phenomena in markets are analogous to critical phenomena in physics. A simple classification of the market crisis p…
SAM selects flatter minima late in training, improving generalization.
problem Improving neural network generalization under various settings.
method Sharpness-Aware Minimization (SAM) applied late in training.
result SAM efficiently selects flatter minima late in training, improving generalization.
Dynamic trading strategies, in the spirit of trend-following or mean-reversion, represent an only partly understood but lucrative and pervasive area of modern finance. Assuming Gaussian returns and Gaussian dynamic weights or signals, (e.g., linear filters of past returns, such as simple moving averages, exponential we…
New insights into network generalization show learning rate affects both norm and sharpness.
problem Understanding the generalization of overparameterized networks.
method Empirical analysis and theoretical proof of the trade-off between norm and sharpness.
result Learning rate influences both norm and sharpness, neither alone minimizes generalization error.
Study community detection in multi-view data with various types of information.
problem Community detection in multi-view data with different types of information.
method Unified theoretical framework, mutual information analysis, sharp thresholds, iterative algorithms.
result Sharp thresholds for community recovery in various multi-view settings.
Improved error estimate for SGLD sampling algorithm.
problem Establishing a precise error bound for SGLD.
method Sharp uniform-in-time error estimate for SGLD under mild assumptions.
result Uniform-in-time O(η2) bound for KL-divergence between SGLD and Langevin diffusion. The paper characterizes optimal dynamic portfolios for a modified mean-variance utility.
problem Optimal dynamic portfolio choice for a modified mean-variance utility.
method Complete characterization under minimal assumptions, no restrictions on asset return moments.
result Maximal MMV utility is linked to the monotone Sharpe ratio, with global squared MSR as the nominal yield.
Study of SGD with state-dependent noise, improving escape from local minima.
problem Understanding and improving the dynamics of SGD in non-convex optimization.
method Formal study on SGD with state-dependent noise, proposing power-law dynamic with state-dependent diffusion.
result Power-law dynamic can escape from sharp minima exponentially faster than flat minima.
Enhances deep learning by boosting generalization and convergence.
problem Improving generalization and convergence in deep learning models.
method Implicit Regularization Enhancement (IRE) framework that decouples flat and sharp directions.
result IRE consistently improves generalization performance across various deep learning tasks and models.
It is well established that in a market with inclusion of a risk-free asset the single-period mean-variance efficient frontier is a straight line tangent to the risky region, a fact that is the very foundation of the classical CAPM. In this paper, it is shown that in a continuous-time market where the risky prices are …
The article uses dynamic factor allocation to improve portfolio performance by integrating regime-switching signals.
problem Improving portfolio performance through dynamic factor allocation.
method The authors apply the sparse jump model (SJM) to identify bull and bear market regimes for individual factors, then fine-tune hyperparameters using a hypothetical single-factor long-short strategy. These regime inferences are incorporated into the Black-Litterman framework to dynamically adjust allocations among indices.
result The constructed multi-factor portfolio significantly improves the information ratio (IR) relative to the market, raising it from 0.05 to approximately 0.4.
Understanding the behavior of stochastic gradient descent (SGD) in the context of deep neural networks has raised lots of concerns recently. Along this line, we study a general form of gradient based optimization dynamics with unbiased noise, which unifies SGD and standard Langevin dynamics. Through investigating this …
Deep linear networks minimize sharpness, avoiding large eigenvalues.
problem Understanding optimization dynamics in deep linear networks for regression.
method Analyzing sharpness (largest eigenvalue of Hessian) of minimizers and gradient flow solutions.
result Gradient flow implicitly regularizes towards flat minima, with sharpness bounded by a constant.
Model explains herding and volatility in urban housing prices.
problem Understanding non-linear price dynamics in urban housing markets.
method Agent-based model with rational households and trend-following behavior.
result Model accurately predicts price variability and herding behavior.
Modeling financial markets with sandpile model to understand price volatility and arbitrage constraints.
problem Understanding price volatility and arbitrage constraints in financial markets.
method Uses a sandpile model to represent information and price changes, linking size of price volatility to the scaling law of avalanches.
result Identifies a structural tension between non-arbitrage condition and price adjustments consistent with a constant Sharpe ratio.
SGD favors flat minima exponentially more than sharp minima in deep learning.
problem Understanding how SGD selects flat minima in deep learning.
method Developed a density diffusion theory (DDT) to analyze minima selection.
result SGD exponentially favors flat minima over sharp minima due to Hessian-dependent noise.
Proposes ICC method for dynamic portfolio optimization.
problem Non-stationarity in market conditions makes traditional portfolio optimization ineffective.
method Inverse Covariance Clustering (ICC) to identify market states and integrate into dynamic optimization.
result ICC-PO generates portfolios with higher Sharpe Ratios and greater robustness.
A machine learning approach for dynamic stock recommendation outperforms traditional strategies.
problem Lack of time for analysts to check all S&P 500 stocks and the need for a reliable stock selection strategy.
method Selecting representative stock indicators, using five machine learning methods, and choosing the model with the lowest Mean Square Error to rank stocks.
result The proposed scheme outperforms the long-only strategy on the S&P 500 index in terms of Sharpe ratio and cumulative returns.
The paper explores proper actions and their relation to representation theory, with new quantitative methods.
problem Understanding proper actions and their connection to representation theory.
method Geometric criteria, sharpness measure, and dynamical volume estimates.
result New quantitative methods have established temperedness criteria for unitary representations.
SAM improves generalization in overparameterized models, but its behavior in tensorized models is less understood.
problem Understanding the implicit regularization of SAM in tensorized models.
method Scale-invariance analysis and gradient flow analysis to derive Norm Deviation as a measure of core norm imbalance, and propose Deviation-Aware Scaling (DAS).
result DAS achieves competitive or improved performance over SAM, while offering reduced computational overhead.
We consider the sharp interface limit of the Allen-Cahn equation with Dirichlet or dynamic boundary conditions and give a varifold characterization of its limit which is formally a mean curvature flow with Dirichlet or dynamic boundary conditions. In order to show the existence of the limit, we apply the phase field me…
Develops Heuristic Portfolio Optimization (HPO) as an information-restricted projection of Markowitz/tangency solution
problem Practitioners allocate capital with forecast-light rules like equal weight, inverse volatility, risk parity, HRP, and RA-HRP
method Implies-return principle and fixed-tree cluster-Sharpe recursion
result Formalizes HPO maps, proves defect equals squared inefficiency, and identifies nodewise alphas as policy-gradient coordinates
Paper introduces Market-adaptive Ratio for better portfolio management.
problem Traditional risk-adjusted ratios fail to account for bull and bear markets.
method Integrates ρ parameter and uses reinforcement learning to adjust portfolio allocations dynamically. result Market-adaptive Ratio outperforms traditional ratios in bull and bear markets.
We conduct a post hoc analysis of solar flare predictions made by a Long Short Term Memory (LSTM) model employing data in the form of Space-weather HMI Active Region Patches (SHARP) parameters calculated from data in proximity to the magnetic polarity inversion line where the flares originate. We train the the LSTM mod…
GD converges faster to flatter minima than gradient flow in shallow networks.
problem Understanding the dynamics of gradient descent in shallow linear networks.
method Analyzing the convergence rate and solution of gradient descent in depth-2 linear neural networks.
result GD converges linearly to flatter minima than gradient flow, even with large step sizes.
Study on curve diffusion flows with scale-critical curvature term.
problem Analyzing stability of curve diffusion flows with scale-critical curvature.
method Introduced and studied a one-parameter family of curve diffusion flows with a scale-critical cubic curvature term. Analyzed dynamical stability of homothetic circles using variational methods.
result Established that any small perturbation of an ω-fold circle monotonically approaches the unit ω-circle after rescaling, translation, and reparametrisation. We propose a non-parametric link prediction algorithm for a sequence of graph snapshots over time. The model predicts links based on the features of its endpoints, as well as those of the local neighborhood around the endpoints. This allows for different types of neighborhoods in a graph, each with its own dynamics (e.…
SAM optimizes deep networks by oscillating between sides of the minimum.
problem Improving performance of deep networks.
method Gradient-based optimization method that oscillates between sides of the minimum.
result SAM effectively performs gradient descent on the spectral norm of the Hessian, encouraging drift towards wider minima.
We give a sharp lower bound for the number of geometrically distinct contractible periodic orbits of dynamically convex Reeb flows on prequantizations of symplectic manifolds that are not aspherical. Several consequences of this result are obtained, like a new proof that every bumpy Finsler metric on Sn carries at l…
This paper is devoted to rigidity of smooth bundles which are equipped with fiberwise geometric or dynamical structure. We show that the fiberwise associated sphere bundle to a bundle whose leaves are equipped with (continuously varying) metrics of negative curvature is a topologically trivial bundle when either the ba…
Semi-static trading strategies make frequent appearances in mathematical finance, where dynamic trading in a liquid asset is combined with static buy-and-hold positions in options on that asset. We show that the space of outcomes of such strategies can have very poor closure properties when all European options for a f…
Complex dynamical systems driven by the unravelling of information can be modelled effectively by treating the underlying flow of information as the model input. Complicated dynamical behaviour of the system is then derived as an output. Such an information-based approach is in sharp contrast to the conventional mathem…
The paper shows how to uniquely determine a connection up to gauge.
problem Determining a connection up to gauge from its holonomies.
method Using hyperbolic dynamical systems and transport operators.
result The primitive trace map is locally injective near generic points.
New theory explains how chaotic training improves neural network generalization.
problem Understanding how chaotic training improves neural network generalization.
method Representing stochastic optimizers as random dynamical systems and introducing a new dimension concept.
result Generalization in chaotic training depends on the complete Hessian spectrum and partial determinants.