PDA method optimizes neural networks with global convergence rate analysis.
problem Quantitative convergence rate for neural network optimization in mean field regime.
method Particle dual averaging (PDA) method, combining Langevin algorithm and outer loop optimization.
result Established quantitative global convergence for two-layer mean field neural networks.
Paper uses averaging from many particle filters to approximate posterior predictive distributions.
problem Approximating posterior predictive distributions efficiently and accurately.
method Particle swarm filter algorithm that averages many particle filter approximations.
result Law of large numbers and central limit theorem support the method's effectiveness.
Using holographic renormalization coupled with the Caffarelli/Silvestre\cite{caffarelli} extension theorem, we calculate the precise form of the boundary operator dual to a bulk scalar field rather than just its average value. We show that even in the presence of interactions in the bulk, the boundary operator dual to …
PAPAL algorithm finds mixed Nash equilibria in continuous games.
problem Finding mixed Nash equilibria in non-convex, non-concave games.
method Particle-based Primal-Dual Algorithm (PAPAL) for weakly entropy-regularized min-max optimization.
result PAPAL offers non-asymptotic convergence guarantees for ε-mixed Nash equilibrium. Dual training method for EBMs with overparametrized neural networks.
problem Training EBMs with non-convex energies is challenging.
method Derive variational principles and dual GDA algorithm for feature-learning regime.
result Dual GDA algorithm performs best with similar time scales for features and particles.
Single particle reconstruction (SPR) from cryo-electron microscopy (EM) is a technique in which the 3D structure of a molecule needs to be determined from its contrast transfer function (CTF) affected, noisy 2D projection images taken at unknown viewing directions. One of the main challenges in cryo-EM is the typically…
Effective action for Kerr-Newman black hole found in twistor theory.
problem Effective action for Kerr-Newman black hole.
method Twistor particle theory to achieve all-orders worldline effective action.
result Exact hidden symmetries identified in self-dual backgrounds.
Dual averaging-type methods are widely used in industrial machine learning applications due to their ability to promoting solution structure (e.g., sparsity) efficiently. In this paper, we propose a novel accelerated dual-averaging primal-dual algorithm for minimizing a composite convex function. We also derive a stoch…
The paper analyzes rates for a modified gradient descent method using Stein variational gradients.
problem Improving the accuracy of gradient descent methods for complex target distributions.
method Derives finite-particle rates for regularized Stein variational gradient descent (R-SVGD).
result Establishes explicit non-asymptotic bounds for time-averaged empirical measures.
New algorithm for federated learning with non-smooth regularizers.
problem Federated Learning with non-smooth composite optimization problems.
method Proposed Federated Dual Averaging (FedDualAvg) algorithm to overcome convergence issues.
result FedDualAvg outperforms other algorithms in federated composite optimization.
A new algorithm computes Wasserstein barycenters without entropic regularization.
problem Computing Wasserstein barycenters efficiently and accurately.
method Free-support algorithm based on particle flow and Riemannian geometry.
result The algorithm avoids entropic regularization and is computationally tractable.
DSPI connects natural policy gradient to policy iteration, proving global convergence.
problem Optimizing policies in reinforcement learning.
method DSPI framework, combining smoothed policy iteration and natural policy gradient.
result DSPI achieves geometric convergence and optimal complexity for policy optimization.
Bayesian inference for inverse problems using mean-shift interacting particles
problem Bayesian inference for inverse problems
method Amortized mean-shift interacting particles
result Improves accuracy of Bayesian inference by reducing the number of samples needed
Study on convergence of Langevin dynamics for zero-sum games in probability distributions.
problem Analyzing convergence of Langevin dynamics for zero-sum games in probability distributions.
method Proved exponential and biased convergence guarantees for mean-field and finite-particle min-max Langevin dynamics.
result Explicit iteration complexity for finite-particle algorithms to approximate equilibrium distributions.
Bayesian nonparametric models for data with heterogeneous particles.
problem Deconvolving data with heterogeneous particles, like voter tallies in elections.
method Nonparametric deconvolution models (NDMs) using two tiers of Dirichlet processes.
result NDMs can recover how factor distributions vary locally for each observation.
Model tracks structural changes in Brownian particle configurations on a sphere.
problem Tracking structural changes in Brownian particle configurations on a sphere.
method Introduces Frustrated Distance Matrix (FDM) model for dynamic distance matrices on S^2.
result Preserves static BBS template with dynamics as redistributed spectral mass.
New algorithm for solving minimax problems over distributions converges to Nash equilibrium.
problem Solving minimax problems over probability distributions.
method Symmetric Mean-field Langevin Dynamics (MFL-AG and MFL-ABR) with weighted averaging and best response dynamics.
result Converges to mixed Nash equilibrium with average-iterate and last-iterate convergence.
WSINDy identifies reduced Hamiltonian systems from particle interactions.
problem Coarse-graining Hamiltonian dynamics with approximate symmetries.
method WSINDy algorithm applied to Hamiltonian systems with timescale separation.
result WSINDy successfully identifies reduced Hamiltonian systems from noisy data.
Despite the success of single-agent reinforcement learning, multi-agent reinforcement learning (MARL) remains challenging due to complex interactions between agents. Motivated by decentralized applications such as sensor networks, swarm robotics, and power grids, we study policy evaluation in MARL, where agents with jo…
We consider a composite convex minimization problem associated with regularized empirical risk minimization, which often arises in machine learning. We propose two new stochastic gradient methods that are based on stochastic dual averaging method with variance reduction. Our methods generate a sparser solution than the…
In this study, we present a multi-class graphical Bayesian predictive classifier that incorporates the uncertainty in the model selection into the standard Bayesian formalism. For each class, the dependence structure underlying the observed features is represented by a set of decomposable Gaussian graphical models. Emp…
New samplers minimize KL divergence for constrained and non-Euclidean geometries.
problem Efficient sampling from constrained and non-Euclidean distributions.
method Stein Variational Mirror Descent and Mirrored Stein Variational Gradient Descent.
result New samplers converge more rapidly and accurately than prior methods.
New method tackles composite optimization with error feedback.
problem Challenges in distributed machine learning training and message compression.
method Combines Dual Averaging with EControl for composite optimization.
result First strong convergence analysis for composite optimization with error feedback.
Paper improves particle variational inference by optimizing generalization error bound.
problem Improving the diversity of models in particle variational inference to enhance generalization.
method Develops a new second-order Jensen inequality with a repulsion term based on the loss function, leading to a tighter generalization error bound.
result The proposed PVI optimizes the generalization error bound directly, improving performance compared to existing methods.
This dissertation advances the theoretical foundation of local optimization methods in Federated Learning.
problem Theoretical understanding of local optimization methods in Federated Learning is lacking.
method The dissertation proposes and analyzes new methods to improve convergence rates and communication efficiency in Federated Learning.
result Sharp bounds and convergence rates for FedAvg are established, and new methods like FedAc and Federated Dual Averaging are proposed.
Improved convergence rates for Stein Variational Gradient Descent in finite-particle settings.
problem Improving convergence rates for Stein Variational Gradient Descent in finite-particle settings.
method Analyzing the time derivative of relative entropy and splitting it into dominant and smaller parts.
result Finite-particle convergence rates of order 1/\sqrt{N} for Kernelized Stein Discrepancy and Wasserstein-2 metrics.
Paper proposes a new framework for robust multi-modal data fusion under uncertainty.
problem Unexpected modality failures in nonlinear non-Gaussian dynamic processes.
method Dynamic model averaging (DMA) based particle filter (PF) algorithm.
result The proposed solution outperforms state-of-the-art methods in experiments.
We generalise the description of the dynamics of the order book of financial markets in terms of a Brownian particle embedded in a fluid of incoming, exiting and annihilating particles by presenting a model of the velocity on each side (buy and sell) independently. The improved model builds on the time-averaged number …
New algorithm uses PSO to optimize DNN training parameters in distributed systems.
problem Reducing synchronization frequency in DNN training leads to poor convergence.
method Integrates PSO into distributed training to automatically compute new parameters.
result Proposed algorithm outperforms synchronous methods in distributed DNN training.
The chart of the nuclides is limited by particle drip lines beyond which nuclear stability to proton or neutron emission is lost. Predicting the range of particle-bound isotopes poses an appreciable challenge for nuclear theory as it involves extreme extrapolations of nuclear masses beyond the regions where experimenta…
MDA optimizer performs similarly to SGD+M in CV and Adam in NLP.
problem Performance degradation due to choosing the wrong optimizer.
method Modernized Dual Averaging (MDA) optimizer, inspired by dual averaging.
result MDA performs as well as SGD+M in CV and as Adam in NLP.
In quantum physics, the operators associated with the position and the momentum of a particle are unbounded operators and C∗-algebraic quantisation does therefore not deal with such operators. In the present article, I propose a quantisation of the Lie-Poisson structure of the dual of a Lie algebroid which deals wit…
New methods learn sampling distributions for particle filters without supervision.
problem Designing accurate sampling distributions for nonlinear dynamical systems.
method Proposed four unsupervised learning methods for multivariate Gaussian and nonparametric distributions.
result Learned sampling distributions outperform designed ones in accuracy.
In decentralized networks (of sensors, connected objects, etc.), there is an important need for efficient algorithms to optimize a global cost function, for instance to learn a global model from the local data collected by each computing unit. In this paper, we address the problem of decentralized minimization of pairw…
Dual Bayesian Affine Estimators for Wiener-type state-space models
problem Estimating parameters in Wiener-type state-space models
method Fixed-point architecture combining two affine estimators
result Dual basis-parameter estimator achieves comparable parameter MSE to purely affine estimator
Paper develops Byzantine-resilient algorithms for decentralized learning.
problem Vulnerability of distributed learning to Byzantine attacks.
method Dual approach for decentralized optimization.
result Convergence guarantees and experimental validation of the proposed algorithm.
OMD and DA perform similarly in static settings but OMD is inferior under dynamic learning rates.
problem Proving and understanding the performance difference between OMD and DA under dynamic learning rates.
method Introducing stabilization to OMD and modifying its convergence analysis.
result OMD with stabilization and DA have the same performance guarantees under dynamic learning rates.
We define an almost--cosymplectic--contact structure which generalizes cosymplectic and contact structures of an odd dimensional manifold. Analogously, we define an almost--coPoisson--Jacobi structure which generalizes a Jacobi structure. Moreover, we study relations between these structures and analyse the associated …
Discrete normal surfaces are normal surfaces whose intersection with each tetrahedron of a triangulation has at most one component. They are also natural Poincaré duals to 1-cocycles with $\ZZ/2\ZZ$-coefficients. For a fixed cohomology class in a simplicial poset the average Euler characteristic of the associated discr…
Novel analysis of EFP for finite-sum problems in neural networks.
problem Optimization of two-layer neural networks in the mean-field regime.
method Primal-dual analysis of entropic fictitious play (EFP) for finite-sum problems.
result Established global convergence guarantees for EFP dynamics.
We consider Lie(G)-valued G-invariant connections on bundles over spaces G/H, RxG/H and R^2xG/H, where G/H is a compact nearly Kaehler six-dimensional homogeneous space, and the manifolds RxG/H and R^2xG/H carry G_2- and Spin(7)-structures, respectively. By making a G-invariant ansatz, Yang-Mills theory with torsion on…
This paper introduces metrics for welfare analysis in dynamic models. We develop estimation and inference for these parameters even in the presence of a high-dimensional state space. Examples of welfare metrics include average welfare, average marginal welfare effects, and welfare decompositions into direct and indirec…
This paper presents a fast Bayesian filtering technique for state estimation.
problem Bottleneck in Bayesian inference for state estimation from noisy sensor data.
method Processor-native uncertainty tracking for uncertainty propagation and inference.
result Deterministic approximate filtering with up to 805x speedup and competitive accuracy.
We propose a voted dual averaging method for online classification problems with explicit regularization. This method employs the update rule of the regularized dual averaging (RDA) method, but only on the subsequence of training examples where a classification error is made. We derive a bound on the number of mistakes…
Study electric-magnetic duality in M-theory compactifications.
problem Restoring supersymmetry and understanding dualities in compactified M-theory.
method Dualizing M-theory on G2 manifolds with F-theory and studying D3-branes. result Demonstrates correspondence between D3-branes and shrinking surfaces/curves, revealing light particles with electric and magnetic charges.
Inference-Time Scaling can be extended to domains prone to systematic failure using intrinsic statistics.
problem Scaling inference time in domains prone to systematic failure
method Intrinsic Selection (iS), Intrinsic Particle Filtering (iPF), and Particle Distillation (dPF)
result Intrinsic Selection improves engineering design selection by 20% and pass@1 by 6.1 points on average.
Develops a gradient flow for Muon optimizer, a method for optimization.
problem Optimization of complex systems with matrix-valued parameters.
method Gradient flow on probability measures induced by regularized Muon optimizer.
result Derives continuous-time limits and proves Hamiltonian dissipation.
Study on Santaló point for convex bodies in normed spaces.
problem Exploring Santaló point for convex bodies in normed spaces.
method Existence and uniqueness proof for C1 norms, dual Santaló point for smooth curved unit balls. result Existence and uniqueness of Santaló point for convex bodies in normed spaces.