MuonEq improves training of matrix-valued parameters by rebalancing momentum before orthogonalization.
problem Training matrix-valued parameters with orthogonalized-update optimizers like Muon.
method MuonEq introduces three lightweight pre-orthogonalization equilibration schemes: two-sided row/column normalization (RC), row normalization (R), and column normalization (C).
result Row/column normalization acts as a zeroth-order surrogate for whitening and improves the geometry seen by orthogonalization.
Proves finite step termination of Kähler-Einstein metric singularity formation.
problem Singularity formation of Kähler-Einstein metrics.
method Finite step termination of bubble trees for singularity formation.
result Finite step termination of Kähler-Einstein metric singularity formation proved in non-collapsing situation.
In this article we use the "escape from subvarieties lemma" introduced by Eskin--Mozes--Oh to prove finite step rigidity results for the Jordan-Lyapunov projection spectra of Hitchin representations and the Margulis-Smilga invariant spectra of some special Margulis-Smilga spacetimes. In the process, we also prove a sim…
The paper analyzes convergence properties of NGA and PAMe for L1-norm PCA.
problem Finite-step convergence of L1-norm PCA algorithms. method Conditional subgradient and alternating maximization interpretations of NGA, and PAMe with extrapolation.
result Iterative points of modified NGA and PAMe remain constant after finitely many steps under certain conditions.
New insights into how large learning rates affect transformer training dynamics.
problem Understanding how large learning rates impact the training of transformer models.
method Analyzing a simplified linear transformer model with a two-factor product map.
result Large learning rates can lead to various training outcomes including cycles, chaos, or divergence.
Study on deep matrix factorization with Bures-Wasserstein loss, focusing on critical points and convergence.
problem Analyzing critical points and convergence of generative deep linear networks trained with Bures-Wasserstein loss.
method Characterization of critical points and minimizers of Bures-Wasserstein distance, analysis of Hessian at low-rank matrices, convergence results for gradient flow and descent.
result Established convergence results for gradient flow and finite step size gradient descent under certain assumptions.
Optimal self-distillation improves generative models' velocity risk and mode recovery.
problem Improving generative models' velocity risk and mode recovery.
method Proved optimal self-distillation for rectified flow via linear probing, derived mixing coefficient, and provided validation tuning.
result Optimal self-distillation improves velocity risk and mode recovery.
We develop a primal dual active set with continuation algorithm for solving the \ell^0-regularized least-squares problem that frequently arises in compressed sensing. The algorithm couples the the primal dual active set method with a continuation strategy on the regularization parameter. At each inner iteration, it fir…
We propose a semismooth Newton algorithm for pathwise optimization (SNAP) for the LASSO and Enet in sparse, high-dimensional linear regression. SNAP is derived from a suitable formulation of the KKT conditions based on Newton derivatives. It solves the semismooth KKT equations efficiently by actively and continuously s…
This paper proposes a multi-grid method for learning energy-based generative ConvNet models of images. For each grid, we learn an energy-based probabilistic model where the energy function is defined by a bottom-up convolutional neural network (ConvNet or CNN). Learning such a model requires generating synthesized exam…
Fewer data weight updates lead to faster convergence in machine learning models.
problem Improving robustness of machine learning models through data mixing.
method Analyzing convergence behavior of data mixing with a finite number of inner steps.
result The optimal number of inner steps scales with the budget and type of gradients used.
SAGD uses Langevin algorithm for efficient gradient descent.
problem Efficiently approximating gradients in complex models.
method Langevin algorithm for biased but asymptotically accurate gradients.
result Theoretical convergence guarantee for SAGD.
When the data are stored in a distributed manner, direct application of traditional statistical inference procedures is often prohibitive due to communication cost and privacy concerns. This paper develops and investigates two Communication-Efficient Accurate Statistical Estimators (CEASE), implemented through iterativ…
In this article, we sketch an algorithm that extends the Q-learning algorithms to the continuous action space domain. Our method is based on the discretization of the action space. Despite the commonly used discretization methods, our method does not increase the discretized problem dimensionality exponentially. We wil…
Adam's hyperparameters implicitly regularize solutions, penalizing or impeding loss gradients' norms.
problem Implicit regularization in Adam's hyperparameters and training stage.
method Backward error analysis and ODE approximations to study Adam's behavior.
result Adam's implicit regularization depends on hyperparameters and training stage, involving different norms.
A new method learns latent space normalizing flow for approximate inference in generator models.
problem Approximate inference in generator models with complex posterior distributions.
method Jointly learns latent space normalizing flow and generator model using MCMC-based maximum likelihood.
result The short-run Langevin flow approximates the posterior and aligns with the normalizing flow prior.
This research accelerates sampling methods using Nesterov's Acceleration.
problem Improving sampling efficiency in MCMC methods.
method Developed a Hessian-Free High-Resolution ODE reformulation of NAG-SC, injected noise, and discretized the diffusion process.
result Quantified acceleration beyond underdamped Langevin in W2 distance for log-strongly-concave targets. Paper proposes CoopFlow, a two-flow generator for energy-based models.
problem Training energy-based models with Langevin flow and normalizing flow.
method CoopFlow trains an energy-based model using a normalizing flow initialization and a short-run Langevin flow revision.
result CoopFlow converges to a moment matching estimator and synthesizes realistic images.
We study geometric properties of the Lagrangian self-shrinking tori in R4. When the area is bounded above uniformly, we prove that the entropy for the Lagrangian self-shrinking tori can only take finitely many values; this is done by deriving a Łojasiewicz-Simon type gradient inequality for the branched conf…
This paper studies the fundamental problem of learning deep generative models that consist of multiple layers of latent variables organized in top-down architectures. Such models have high expressivity and allow for learning hierarchical representations. Learning such a generative model requires inferring the latent va…
We study realizable continual linear regression under random task orderings, a common setting for developing continual learning theory. In this setup, the worst-case expected loss after k learning iterations admits a lower bound of Ω(1/k). However, prior work using an unregularized scheme has only established an up…
Large GD stepsizes improve margins and speed up training for non-homogeneous networks.
problem Training efficiency and margin improvement in non-homogeneous two-layer networks.
method Investigation of two distinct phases in GD training, showing margin growth and empirical risk decrease.
result Large GD stepsizes lead to faster convergence and improved margins in non-homogeneous networks.
This paper studies the cooperative training of two generative models for image modeling and synthesis. Both models are parametrized by convolutional neural networks (ConvNets). The first model is a deep energy-based model, whose energy function is defined by a bottom-up ConvNet, which maps the observed image to the ene…
This paper proposes a method to train energy-based models using variational auto-encoders for efficient sampling.
problem Training energy-based models by maximum likelihood is challenging due to intractable partition functions and difficult sampling from the model distribution.
method The authors propose using a variational auto-encoder to initialize finite-step MCMC sampling, specifically Langevin dynamics, to train the energy-based model.
result The proposed method enables training energy-based models using maximum likelihood, generating samples comparable to GANs and EBMs.
Synthetic tabular data synthesis models balance utility and risk.
problem Generating synthetic tabular data for regulated domains.
method Latent flow models with various learning targets, paths, and sampling methods.
result Velocity and posterior matching objectives yield higher utility, while score and noise matching achieve lower risk.
Develops formal moduli theory for splitting complex supermanifolds.
problem Tackles the splitting problem of complex supermanifolds.
method Constructs a filtered dg Lie algebra to control splittings and transfers the theory to a minimal filtered L∞-model. result Recover classical obstruction classes as leading terms of Maurer-Cartan representatives and proves the existence of higher obstructions.
Study quantization effects on high-dimensional linear regression learning.
problem Understanding quantization's impact on learning high-dimensional linear regression models.
method Analyzes stochastic gradient descent for high-dimensional linear regression under various quantization targets.
result Establishes precise bounds on excess risk for different quantization schemes.
Analyzes symmetries in neural networks to predict learning dynamics.
problem Understanding the dynamics of neural network parameters during training.
method Unified theoretical framework based on symmetries and conservation laws.
result Symmetries impose geometric constraints on gradients and Hessians, leading to conservation laws.
Deep networks retain initial bias after training, affecting generalization.
problem Understanding how much initial bias in neural networks survives training.
method Introduced initialization memory to measure initial bias's survival.
result SGD can preserve initial bias, while Adam-family methods erase it.
Demanding sparsity in estimated models has become a routine practice in statistics. In many situations, we wish to require that the sparsity patterns attained honor certain problem-specific constraints. Hierarchical sparse modeling (HSM) refers to situations in which these constraints specify that one set of parameters…
Formalizes identifying information to answer key questions about machine learning from uncertain and novel observations.
problem Understanding and quantifying information from uncertain and novel observations in machine learning.
method Formalizes identifying information, defines hypothesis identification and sample complexity, and proves sample complexity properties for various data-generating processes.
result Proves the information theoretic characteristics of hypothesis identification and sample complexity, and shows how to compute identifying information and novel information.
Proves weak convergence equals mean convergence in GGC.
problem Proving convergence in GGC distributions.
method Using generalized gamma convolution (GGC) and expected utility maximization.
result Weak convergence implies mean convergence in GGC.
This is an intuitive survey of extrinsic and intrinsic notions of convergence of manifolds complete with pictures of key examples and a discussion of the properties associated with each notion. We begin with a description of three extrinsic notions which have been applied to study sequences of submanifolds in Euclidean…
The abstract discusses convergence properties of Lipschitz functions and sets defined by equations.
problem Convergence of Lipschitz functions and sets defined by equations.
method Painlevé-Kuratowski convergence applied to Lipschitz functions and sets defined by equations.
result Generalizations and reverses of classical theorems on convergence of functions and sets.
Study shows intrinsic timed Hausdorff convergence leads to Gromov-Hausdorff and big bang convergence.
problem Distance between Lorentzian manifolds.
method Intrinsic timed Hausdorff convergence.
result Intrinsic timed Hausdorff convergence implies Gromov-Hausdorff and big bang convergence.
Studied SGD convergence under weak conditions.
problem Convergence of SGD in nonconvex optimization.
method Analyzed biased nonconvex SGD under mild conditions.
result Provided convergence rates and complexities.
Study on convergence rate of Q-curvature flow in 6 dimensions.
problem Analyzing the convergence rate of Q-curvature flow in 6 dimensions. method Provided an example of a slowly converging Q6-curvature flow in dimension 6. result The Q-curvature flow in 6 dimensions does not always converge exponentially, unlike in 2 dimensions. The objective of this paper is to introduce the notion of generalized almost statistical (briefly, GAS) convergence of bounded real sequences, which generalizes the notion of almost convergence as well as statistical convergence of bounded real sequences. As a special kind of Banach limit functional, we also introduce …
Establishes geometric convergence of iterative optimization algorithms.
problem Analyzes convergence of iterative optimization algorithms under general assumptions.
method General framework for iterative optimization algorithms, proving asymptotic geometric convergence and providing convergence rates.
result Asymptotic geometric convergence of iterative optimization algorithms with exact rate.
Uniform counting formulas for orthogeodesics in Kleinian groups converge.
problem Counting orthogeodesics in Kleinian groups converging to a limit.
method Spectral gap of the limit manifold and geodesic flow mixing property.
result Asymptotically uniform counting formulas for orthogeodesics.
The paper explores null distance convergence for warped product spacetimes.
problem Defining convergence for sequences of spacetimes as metric spaces.
method Using the null distance to define convergence of spacetimes.
result Optimal convergence theorem for warped product spacetimes.
New quasi-Newton method guarantees global superlinear convergence.
problem Global convergence and superlinear convergence of quasi-Newton methods.
method Hybrid proximal extragradient method with online learning for Hessian approximation.
result First globally convergent quasi-Newton method with explicit superlinear convergence rate.
We investigate finite-time decoupled convergence in nonlinear two-time-scale stochastic approximation.
problem Achieving decoupled convergence in nonlinear two-time-scale stochastic approximation.
method Nested local linearity assumption, suitable step size selection, convergence analysis of matrix cross term, fourth-order moment convergence rates.
result Finite-time decoupled convergence rates can be achieved in nonlinear two-time-scale stochastic approximation with proper step size selection.
The article introduces a new convergence concept for Lorentzian spaces and applies it to generalized cones.
problem Stability of curvature bounds in generalized Lorentzian cones.
method Introduces ℓ-convergence for Lorentzian pre-length spaces, applies it to generalized cones, and proves stability of curvature bounds. result Sharp timelike curvature and curvature-dimension bounds for generalized cones are established.
AdaBoost's classifier and margins converge to a known value.
problem Convergence properties of AdaBoost algorithm.
method Formal proofs of convergence properties of AdaBoost's classifier and margins.
result AdaBoost's classifier and margins converge to a known value.
Study shows gap between uniform convergence and test error in random feature models.
problem Understanding the gap between uniform convergence and test error in random feature models.
method Analytical expressions for uniform convergence over norm balls, interpolators, and minimum norm interpolator risk derived and proved.
result Uniform convergence over interpolators still gives a non-trivial bound of test error even when classical uniform convergence is vacuous.
The Sinkhorn-Knopp derivatives converge with linear rate.
problem Optimal transport problem with entropic regularization.
method Iterative proportional fitting procedure.
result Derivatives converge with linear rate.
The paper examines convergence of distances in Lipschitz structures on manifolds.
problem Convergence of distances in Lipschitz vector fields and norms on manifolds.
method Analysis of convergence of distances associated to converging structures of Lipschitz vector fields and norms.
result Under mild controllability assumption, distances converge locally uniformly to the limit Carnot-Carathéodory distance.