New tools understand and control dynamics in n-player differentiable games.
problem Understanding and controlling the behavior of gradient-based methods in games.
method Developed new tools to understand and control the dynamics in n-player differentiable games, decomposing the game Jacobian into symmetric and antisymmetric components.
result Motivated Symplectic Gradient Adjustment (SGA) algorithm for finding stable fixed points in differentiable games.
New flows introduced for symplectic geometry.
problem No specific problem stated; focuses on new flows.
method Introduces several geometric flows on symplectic manifolds.
result Examples include the Hitchin gradient flow and dual Ricci flow.
This study shows that certain cohomology groups of symplectic manifolds are always even-dimensional.
problem Understanding the cohomology structure of symplectic manifolds.
method Constructing and deforming a skew-adjoint operator to prove the vanishing property.
result The even dimensionality of even-degree cohomology groups in (4n+2)-dimensional symplectic manifolds.
Investigates metric degeneracies on symplectic leaves using a generalized gradient flow.
problem Degeneracies in metrics on symplectic leaves of Poisson manifolds.
method Introduces the generalized double bracket (GDB) vector field to generalize gradient dynamics.
result Identifies admissible regions where the double bracket metric remains non-degenerate on symplectic leaves, enabling GDB as a gradient flow.
Introduces a Morse complex on symplectic manifolds using gradient flows and proves its cohomology is independent of metrics and Morse functions.
problem Cohomology of symplectic manifolds under different metrics and Morse functions.
method Symplectic Morse complex with gradient flows and Witten deformation.
result Cohomology of the complex is isomorphic to Tsai, Tseng, and Yau's cohomology and independent of metrics and Morse functions.
Automatically tunes learning rate and momentum for SGD methods.
problem Manual hyperparameter tuning is costly and lacks theoretical justification.
method Uses statistics of gradient estimator to automatically adjust learning rate and momentum.
result Matches performance of best manual settings for CNN training.
LOSSGRAD automatically adjusts learning rates in neural networks.
problem Finding optimal learning rates in gradient descent.
method LOSSGRAD uses quadratic approximation to find locally optimal step-size.
result LOSSGRAD achieves comparable results to other methods while being insensitive to initial learning rate.
Much of the focus in machine learning research is placed in creating new architectures and optimization methods, but the overall loss function is seldom questioned. This paper interprets machine learning from a multi-objective optimization perspective, showing the limitations of the default linear combination of loss f…
This study proves the local existence of a symplectic gradient flow on a flat torus.
problem Proving the local existence of a symplectic gradient flow on a flat torus.
method Using a moment map and a DeTurck trick to make the flow strictly parabolic and showing local existence and regularity.
result The group of symplectomorphisms of the real four-dimensional torus is locally contractible.
Proves unique symplectic Lefschetz fibration from Morse functions.
problem Mapping Morse functions to symplectic Lefschetz fibrations.
method Homotopically unique complex-valued symplectic Lefschetz fibration on cotangent bundles.
result Existence and uniqueness of symplectic Lefschetz fibrations.
New method preserves convergence rates in gradient-based optimization.
problem How to discretize gradient-based optimization systems while preserving stability and convergence rates.
method Geometric framework for dissipative symplectic integration.
result Dissipative symplectic integrators preserve rates of convergence up to a controlled error.
Introspects convolutional speech recognition models using Gradient-adjusted Neuron Activation Profiles.
problem Lack of interpretability in deep learning ASR models.
method Gradient-adjusted Neuron Activation Profiles (GradNAPs) for feature and representation visualization.
result Gains insight into how data is processed in convolutional ASR models.
Algorithm finds causal effects from observational data using auxiliary variables.
problem Estimating causal effects from observational data with confounders.
method Gradient-based optimization using auxiliary variables.
result Algorithm outperforms alternatives in estimating true causal effect.
Bayesian optimisation for dynamically adjusting learning rates in machine learning models.
problem Dynamic adjustment of learning rates schedules in machine learning models.
method Probabilistic model based on latent Gaussian processes and auto-/regressive formulation.
result Flexibly adjusts learning rates schedules to abrupt changes of behaviours.
When training a machine learning model with observational data, it is often encountered that some values are systemically missing. Learning from the incomplete data in which the missingness depends on some covariates may lead to biased estimation of parameters and even harm the fairness of decision outcome. This paper …
Constructs a Morse-Bott function on symplectic Grassmannians.
problem Defines a function on symplectic Grassmannians.
method Uses a compatible linear complex structure to construct a quadratic Morse-Bott function.
result Critical loci consist of subspaces splitting into isotropic and complex parts.
Derives new optimization methods using variational integrators.
problem Optimization methods in machine learning.
method Variational integrators and principles of Hamilton and Lagrange-d'Alembert.
result Derives two families of optimization methods, including Nesterov's accelerated gradient method.
Stochastic Gradient Descent with a constant learning rate (constant SGD) simulates a Markov chain with a stationary distribution. With this perspective, we derive several new results. (1) We show that constant SGD can be used as an approximate Bayesian posterior inference algorithm. Specifically, we show how to adjust …
SympNets identify Hamiltonian systems from data using linear, activation, and gradient modules.
problem Identifying Hamiltonian systems from data.
method Composition of linear, activation, and gradient modules; universal approximation theorems.
result SympNets can approximate arbitrary symplectic maps and generalize well to various Hamiltonian systems.
In this article we construct Lagrangian torus fibrations for general quintic \cy hypersurfaces near the large complex limit and their mirror manifolds using gradient flow method. Then we prove the Strominger-Yau-Zaslow mirror conjecture for this class of \cy manifolds in symplectic category.
Modern statistical inference tasks often require iterative optimization methods to compute the solution. Convergence analysis from an optimization viewpoint only informs us how well the solution is approximated numerically but overlooks the sampling nature of the data. In contrast, recognizing the randomness in the dat…
Study on local convergence of min-max algorithms to differential equilibria on Riemannian manifolds.
problem Solving zero-sum differential games on Riemannian manifolds.
method Analysis of two simultaneous min-max algorithms, τ-GDA and τ-SGA, to differential Stackelberg and Nash equilibria, with conditions for linear convergence and asymptotic approximation. result Established sufficient conditions for linear convergence of τ-GDA and demonstrated faster convergence of τ-SGA in some cases. We develop the method of stochastic modified equations (SME), in which stochastic gradient algorithms are approximated in the weak sense by continuous-time stochastic differential equations. We exploit the continuous formulation together with optimal control theory to derive novel adaptive hyper-parameter adjustment po…
A new approach optimizes weights in DLP for better risk-adjusted performance.
problem Optimizing time-varying weights in Double Linear Policy (DLP) for better risk-adjusted performance.
method Stochastic Model Predictive Control (SMPC) framework to maximize risk-adjusted returns while enforcing constraints.
result Empirical results show improved risk-adjusted performance and drawdown control.
Two new proofs of Gromov's non-squeezing theorem using curve reparametrization and gradient bounds.
problem Gromov's non-squeezing theorem in symplectic geometry.
method Reparametrization of pseudo-holomorphic curves and application of mean value inequality or Gromov-Schwarz lemma.
result Uniform bounds on the gradient of pseudo-holomorphic curves leading to compactness of moduli space.
Analysis of momentum methods on quadratic models, showing SGD's superiority.
problem Analysis of stochastic gradient algorithms with momentum on quadratic models.
method Inspired by random matrix theory, exact characterization of loss values.
result Stochastic heavy-ball momentum does not improve over SGD in the strongly convex setting.
On a Poisson manifold endowed with a Riemannian metric we will construct a vector field that generalizes the double bracket vector field defined on semi-simple Lie algebras. On a regular symplectic leaf we will construct a generalization of the normal metric such that the above vector field restricted to the symplectic…
Novel multisymplectic framework for pseudo-Fueter curves in Hamiltonian field theory.
problem Generalizing Floer theory to multisymplectic geometry.
method Introducing pseudo-Fueter curves in a compatible almost hyperkähler structure.
result Gradient lines of multisymplectic action functional are pseudo-Fueter curves.
Stochastic Gradient TreeBoost is often found in many winning solutions in public data science challenges. Unfortunately, the best performance requires extensive parameter tuning and can be prone to overfitting. We propose PaloBoost, a Stochastic Gradient TreeBoost model that uses novel regularization techniques to guar…
This is one in a series of papers devoted to the foundations of Symplectic Field Theory sketched in [Y Eliashberg, A Givental and H Hofer, Introduction to Symplectic Field Theory, Geom. Funct. Anal. Special Volume, Part II (2000) 560--673]. We prove compactness results for moduli spaces of holomorphic curves arising in…
New method averages SGD iterates to achieve adjustable regularization.
problem Overfitting in machine learning models.
method Averaging SGD iterates for regularized solutions.
result Obtain regularized solutions without tuning parameters.
Floer theory constructs filtrations on quantum cohomology for symplectic manifolds.
problem Quantum cohomology of symplectic manifolds with C∗-actions. method Floer theory applied to C∗-actions on symplectic manifolds. result Constructs a family of filtrations on quantum cohomology for Conical Symplectic Resolutions.
Study of symplectic Stiefel and Grassmann manifolds with geodesics and applications.
problem Understanding symplectic bases and subspaces for data processing.
method Lie group approach to derive geodesics and retractions for pseudo-Riemannian and Riemannian metrics.
result Efficient formulas for geodesics and retractions on symplectic manifolds.
A new decentralized Bayesian learning method using Metropolis-adjusted Hamiltonian Monte Carlo.
problem Decentralized Bayesian learning with uncertainty quantification.
method Metropolis-adjusted Hamiltonian Monte Carlo in a decentralized federated learning setting.
result Theoretical guarantees and numerical effectiveness of the method on non-convex problems.
AutoClip automatically adjusts gradient clipping for better audio separation.
problem Improving generalization in audio source separation networks.
method Adaptive gradient clipping based on historical gradient norms.
result Improves generalization performance in audio source separation networks.
Proposes a new Langevin flow approach for VAEs.
problem Difficulty in constructing low variance ELBO for VAEs with large datasets.
method Integrates Langevin dynamic with quasi-symplectic integrator to improve posterior estimation.
result Shows theoretical and practical effectiveness compared to gradient flow-based methods.
Optimal preconditioning improves Langevin sampling efficiency.
problem Improving sampling efficiency in high-dimensional target distributions.
method Optimal preconditioning using Fisher information, applied to MALA.
result Adaptive MCMC scheme significantly outperforms other methods.
A novel neural network training method reduces gradient variance for faster and better reinforcement learning.
problem Improving convergence and generalization in deep reinforcement learning.
method Gradient Monitoring (GM) approach to dynamically adjust the learning process based on feedback.
result The proposed methods, especially AM-WGM, significantly enhance model performance and generalization.
We study first-order optimization methods obtained by discretizing ordinary differential equations (ODEs) corresponding to Nesterov's accelerated gradient methods (NAGs) and Polyak's heavy-ball method. We consider three discretization schemes: an explicit Euler scheme, an implicit Euler scheme, and a symplectic scheme.…
DCGD improves training of PINNs by adjusting gradients to avoid negative inner products.
problem Pathological behaviors in PINNs training, especially gradient imbalance.
method Dual Cone Gradient Descent (DCGD) framework to adjust gradient direction.
result DCGD outperforms other optimization algorithms in various evaluation metrics.
Adaptive HMC improves sampling efficiency by optimizing mass matrix.
problem Inefficient HMC performance due to mass matrix choice.
method Gradient-based adaptation of mass matrix to maximize proposal entropy.
result Adaptation method outperforms HMC variants by optimizing mass matrix.
Stochastic gradient descent updates parameters with summation gradient computed from a random data batch. This summation will lead to unbalanced training process if the data we obtained is unbalanced. To address this issue, this paper takes the error variance and error mean both into consideration. The adaptively adjus…
New algorithm speeds up sampling from complex distributions.
problem Efficiently sampling from non-log-concave distributions.
method Stochastic Proximal Samplers (SPS) based on SGLD and MALA.
result SPS-SGLD and SPS-MALA achieve faster sampling with reduced gradient complexity.
New algorithm preserves symplectic structure for faster optimization.
problem Optimization methods in machine learning.
method Structure-preserving discretizations of dissipative Hamiltonian systems.
result Proposes a new algorithm that generalizes Nesterov and heavy ball methods.
Optimizes functions on Lie groups using generalized eigenvalue problems.
problem Optimization on Lie groups with specific applications to eigenvalue problems.
method Generalizes NAG principle to Lie groups, resulting in continuous Lie-NAG dynamics converging to local optima.
result Discretized Lie-NAG dynamics yield structure-preserving optimization algorithms with faithful energy behavior.
PAGE is a simple gradient estimator for nonconvex optimization problems.
problem Nonconvex optimization problems in machine learning.
method PAGE is a probabilistic gradient estimator that uses vanilla SGD with probability and a small adjustment with probability 1-p.
result PAGE achieves optimal convergence rates for nonconvex finite-sum and online problems.
Algorithms for bandit convex optimization and online learning often rely on constructing noisy gradient estimates, which are then used in appropriately adjusted first-order algorithms, replacing actual gradients. Depending on the properties of the function to be optimized and the nature of ``noise'' in the bandit feedb…
Improved particle filters for estimating model parameters using differentiable resampling.
problem Inability to differentiate sampling and resampling steps in particle filters.
method Extended reparameterisation trick to include stochastic input, enabling differentiation. Used p-MCMC and NUTS for parameter estimation.
result NUTS improves mixing of Markov chain and produces more accurate results in less time.