Lion optimizer performs well in training AI models with memory efficiency.
problem Lion optimizer's theoretical basis is unclear.
method Continuous-time and discrete-time analysis of Lion updates with a new Lyapunov function.
result Lion is a novel and principled approach for constrained optimization.
Distributed Lion optimizes large model training by reducing communication costs.
problem Training large AI models efficiently with reduced communication costs.
method Adapted Lion optimizer for distributed training, using binary or lower-precision vectors for communication.
result Distributed Lion achieves comparable performance to standard optimizers but with significantly reduced communication bandwidth.
Enhanced Lion optimizer CLion improves generalization with lower error.
problem Lion optimizer's generalization analysis is lacking.
method Algorithmic stability analysis and cautious sign function use.
result CLion has a lower generalization error of O(N1). Unified view of Lion and Muon as Stochastic Frank-Wolfe methods.
problem Optimization of constrained problems in deep learning.
method Interpreting Lion and Muon as Stochastic Frank-Wolfe methods and extending the approach to heavy-tailed noise.
result Convergence guarantees and KKT point convergence for Lion and Muon.
New optimizers improve stock market forecasting accuracy.
problem Forecasting S&P 500 Index returns with MambaStock model.
method Evaluation of various optimizers (Adam, RMSProp, Lion, Roaree).
result Roaree optimizers combine faster training with reduced oscillations.
This work uses Lasry-Lions envelopes to solve nonconvex optimization problems.
problem Nonconvex and nonsmooth terms in optimization problems.
method Develops a homotopy approach using Lasry-Lions envelopes to approximate and solve the original problem.
result The method can solve composite minimization problems and is more effective than classical alternatives in certain domains.
Muon optimizer improves deep learning with spectral norm constraints.
problem Improving optimization algorithms in deep learning.
method Theoretical analysis of Muon optimizer within the Lion-K family. result Muon implicitly solves an optimization problem enforcing spectral norm constraints.
LION generates high-quality 3D shapes using hierarchical latent diffusion models.
problem Creating high-quality 3D shapes for digital artists.
method Hierarchical Latent Point Diffusion Model (LION) with a global shape latent and point-structured latent space.
result LION achieves state-of-the-art generation performance on ShapeNet benchmarks.
A technique identifies memoryless algorithms approximating memory-dependent optimization methods.
problem Understanding how memory in optimization algorithms affects loss and generalization.
method Introducing a general technique to replace past iterates with the current one and adding a correction term.
result Lion does not have the same implicit anti-regularization as AdamW, explaining its better generalization performance.
Establishes a Lorentzian Lasry-Lions regularization theorem for functions on globally hyperbolic spacetimes.
problem Optimal transport with C1,1 regularizing pairs method Local semiconcavity and future-directed timelike superdifferentials
result Derives C1,1 regularizing pairs for optimal transport under general assumptions Cautious Weight Decay modifies weight decay for better optimization.
problem Improving optimization in deep learning models.
method Applies weight decay selectively based on parameter sign alignment.
result Consistently improves model performance across various tasks and scales.
We build a model using Gaussian processes to infer a spatio-temporal vector field from observed agent trajectories. Significant landmarks or influence points in agent surroundings are jointly derived through vector calculus operations that indicate presence of sources and sinks. We evaluate these influence points by us…
Inspired by recent work of P.-L. Lions on conditional optimal control, we introduce a problem of optimal stopping under bounded rationality: the objective is the expected payoff at the time of stopping, conditioned on another event. For instance, an agent may care only about states where she is still alive at the time …
Is AdamW effective under heavy-tailed noise?
problem Stochastic gradient noise in LLM pretraining is typically heavy-tailed.
method Formulate as an open problem, prove a positive weighted-metric benchmark, and give a corridor lower-bound mechanism.
result No rigorous convergence theory for AdamW established in heavy-tailed regime.
New optimization method combines gradient clipping and non-Euclidean smoothness.
problem Improving optimization in non-Euclidean spaces for machine learning.
method Hybrid of steepest descent and conditional gradient, incorporating weight decay.
result Achieves optimal convergence rate and demonstrates effectiveness in deep learning.
This paper improves SAM by reformulating it as a bilevel optimization problem.
problem Improving Sharpness-Aware Minimization (SAM) for better performance.
method Reformulate SAM as a bilevel optimization problem using a 0-1 loss surrogate.
result BiSAM consistently results in improved performance compared to SAM and its variants.
We consider two functions on Sp(g,R) with values in the cyclic group of order four {1,-1,i,-i}. One was defined by Lion and Vergne. The other is -i raised to the power given by an integer valued function defined by Masbaum and the author (initially on the mapping class group of a surface). We identify these functions w…
The Prescriptive Canvas improves business outcomes by directly prescribing actions based on predictions.
problem Sub-optimal performance in business projects due to a two-step approach of prediction and decision-making.
method The Prescriptive Canvas methodology for framing and communicating actions directly based on predictions.
result Improves framing and communication across stakeholders for successful business impact.
FedLion improves Federated Learning by speeding up convergence and reducing communication costs.
problem Slow convergence and high communication costs in Federated Learning.
method Integrates Lion's adaptive approach into Federated Learning framework, using signed gradients.
result FedLion outperforms existing adaptive algorithms in convergence rate and communication efficiency.
MARS optimizes large model training by reducing variance, outperforming AdamW.
problem Training large models efficiently and scalably.
method Unified optimization framework MARS combining preconditioned gradient updates and variance reduction.
result MARS outperforms AdamW in training GPT-2 models.
Mean field game theory studies the behavior of a large number of interacting individuals in a game theoretic setting and has received a lot of attention in the past decade (Lasry and Lions, Japanese journal of mathematics, 2007). In this work, we derive mean field game partial differential equation systems from determi…
SAMPa speeds up SAM by parallelizing its computations.
problem Improving neural network generalization through SAM.
method Parallelizing the two gradient computations in SAM.
result Achieves a twofold speedup of SAM.
In this paper, we consider the generalized lambda constant and the existence of ground states of the generalized Perelman's W-functional from a variational formulation. One result is concerned with the estimation of the generalized λ constant. The other results are about the existence of ground states of generalized …
Paper establishes NE existence and efficient algorithms for weakly monotone GMFGs.
problem Existence and efficient learning of Nash Equilibrium in λ-regularized GMFGs. method Establishes existence of NE for any λ-regularized GMFGs. Proposes efficient algorithms for weakly monotone GMFGs. result Efficient algorithms for weakly monotone GMFGs with provable convergence.
This article combines various methods of analysis to draw a comprehensive picture of penalty approximations to the value, hedge ratio, and optimal exercise strategy of American options. While convergence of the penalised solution for sufficiently smooth obstacles is well established in the literature, sharp rates of co…
We show how Lasry-Lions's result on regularization of functions defined on Rn or on Hilbert spaces by sup-inf convolutions with squares of distances can be extended to (finite or infinite dimensional) Riemannian manifolds M of bounded sectional curvature. More specifically, among other things we show that…
The paper examines stability of Sobolev inequalities on manifolds with Ricci curvature bounds.
problem Stability of Sobolev inequalities on Riemannian manifolds with Ricci curvature lower bounds.
method Generalized Lions' concentration compactness and rigidity results of Sobolev inequalities on singular spaces.
result Almost extremal functions are close to extremal functions on the round sphere and Euclidean Sobolev inequality.
Lions and Musiela (2007) give sufficient conditions to verify when a stochastic exponential of a continuous local martingale is a martingale or a uniformly integrable martingale. Blei and Engelbert (2009) and Mijatović and Urusov (2012c) give necessary and sufficient conditions in the case of perfect correlation (ρ=1).…
Central bank strategy to maintain currency exchange rate within limits.
problem Maintaining a currency exchange rate within a target zone despite adverse economic trends.
method Modeling the problem with a continuous-time market impact model and solving it as a stochastic control problem.
result Optimal strategy minimizes accumulated inventory of foreign currency.
We decompose, within an ARCH framework, the daily volatility of stocks into overnight and intra-day contributions. We find, as perhaps expected, that the overnight and intra-day returns behave completely differently. For example, while past intra-day returns affect equally the future intra-day and overnight volatilitie…
We prove some old and new isoperimetric inequalities with the best constant using the ABP method applied to an appropriate linear Neumann problem. More precisely, we obtain a new family of sharp isoperimetric inequalities with weights (also called densities) in open convex cones of Rn. Our result applies to…
In the first part of the paper we investigate some geometric features of Moser-Trudinger inequalities on complete non-compact Riemannian manifolds. By exploring rearrangement arguments, isoperimetric estimates, and gluing local uniform estimates via Gromov's covering lemma, we provide a Coulhon, Saloff-Coste and Varopo…
Charmer improves character-level adversarial attacks for language models.
problem Efficiency and effectiveness of character-level adversarial attacks for language models.
method Query-based adversarial attack method that maintains semantic similarity.
result Charmer achieves high attack success rate and similar adversarial examples.
We characterise the link of derivatives in measure, which are introduced in [AKR,Card,ORS] respectively by different means, for functions on the space M of finite measures over a Riemannian manifold M. For a reasonable class of functions f, the extrinsic derivative DEf coincides with the linear functio…
Proves existence of eigenvalue and eigenfunction for complex Monge-Ampère operator.
problem Eigenvalue problem for complex Monge-Ampère operator on bounded domains.
method Follows P.L. Lions' strategy for real case, proves new existence theorem for complex degenerate equations, uses a priori estimates and variational approach.
result Existence of first eigenvalue and eigenfunction with specified properties.
Unified framework for constructing kernels for transport equations and Koopman eigenfunctions.
problem Constructing kernels for transport equations and Koopman eigenfunctions.
method Three methods: variational principle, Green's function, and resolvent operator.
result Kernels constructed via these methods are identical under mild assumptions.
The IMH suggests market price fluctuations are driven by order flow, not fundamental values.
problem Reconciling IMH with microstructure literature on market dynamics.
method Reviewed empirical facts and applied Latent Liquidity Theory to predict price impact multiplier.
result The multiplier M is of order unity, consistent with IMH, and depends on stock volatility and daily traded market cap fraction. In his lectures at College de France, P.L. Lions introduced the concept of Master equation, see [5] for Mean Field Games. It is introduced in a heuristic fashion, from the system of partial differential equations, associated to a Nash equilibrium for a large, but finite, number of players. The method, also explained in…
Efficient regularization mitigates catastrophic overfitting in single-step adversarial training.
problem Catastrophic overfitting in single-step adversarial training.
method ELLE regularization term to enforce local linearity of the loss function.
result Our regularization term effectively mitigates catastrophic overfitting without the drawbacks of previous methods.
Unified kernel framework extends to stochastic systems, improving numerical stability.
problem Extending kernel methods to stochastic dynamical systems with diffusion.
method Unified kernel framework, Feynman-Kac path-integral representations, collocation-based computational framework.
result Kernel equivalence under uniform ellipticity assumptions and improved numerical stability with moderate diffusion.
The past several years have seen both an explosion in the use of Convolutional Neural Networks (CNNs) and the design of accelerators to make CNN inference practical. In the architecture community, the lion share of effort has targeted CNN inference for image recognition. The closely related problem of video recognition…
Bayesian optimization reduces computational effort in aircraft design optimization.
problem High computational cost in industrial aircraft design optimization.
method Constrained Bayesian optimization (Super Efficient Global Optimization with Mixture of Experts)
result Significant computational efficiency improvements over existing Isight optimizers.
Bayesian optimization method tackles combinatorial spaces, scalable for large data.
problem Optimization over combinatorial categorical spaces in natural sciences.
method Combines variational optimization and continuous relaxations for gradient-based optimization.
result Method performs comparably to state-of-the-art methods while scaling well.
New algorithm solves complex stopping problems with robust optimization.
problem Solving complex stochastic optimal stopping problems.
method Simulation-based robust optimization with exact reformulation as a zero-one bilinear program.
result Developed polynomial-time heuristics and algorithms for practical solution.
L2O uses ML to optimize traditional optimization techniques.
problem Real-world optimization problems with shared structures.
method Exploiting shared structures to enhance optimization techniques.
result Better or faster solutions through machine learning integration.
When hyperparameter optimization of a machine learning algorithm is repeated for multiple datasets it is possible to transfer knowledge to an optimization run on a new dataset. We develop a new hyperparameter-free ensemble model for Bayesian optimization that is a generalization of two existing transfer learning extens…
Meta algorithm solves multivariate optimization using univariate optimizers.
problem Multivariate global optimization problems.
method Meta algorithm combining univariate global optimizers.
result Meta algorithm provides robust regret guarantees.
A novel neural network approach for optimization problems.
problem Constrained optimization problems.
method Neural Optimization Machine (NOM) using a specially designed NN architecture and training procedure.
result Solves optimization problems efficiently, especially in high-dimensional spaces.