Study on the smoothness of solutions to a specific type of stochastic differential equation.
problem Regularity of solutions to mean-field G-SDEs. method Analysis of first and second order Fréchet differentiability in the random initial condition.
result Established the Fréchet differentiability of the solution and specified the corresponding equations.
Deep learning relies on good initialization schemes and hyperparameter choices prior to training a neural network. Random weight initializations induce random network ensembles, which give rise to the trainability, training speed, and sometimes also generalization ability of an instance. In addition, such ensembles pro…
Elman-type RNNs converge to globally optimal solutions in the mean-field regime.
problem Optimizing feature learning in wide RNNs.
method Analysis of gradient descent dynamics and mean-field limits.
result Fixed points of infinite-width dynamics are globally optimal.
This work studies clustering in transformer models, proving exponential convergence to a single token state.
problem Understanding the long-term behavior of tokens in transformer models.
method Investigates mean-field transformer models under specific conditions to prove exponential convergence to a single state.
result Transformer models synchronize exponentially fast to a single token state with explicit rates.
New proof links initial class bias to DNN trainability, challenging traditional understanding.
problem Understanding the initial class bias in DNNs and its impact on trainability.
method Theoretical proof linking initial class bias to mean field theories of DNNs.
result Efficient learning is connected to a network's prejudice towards a specific class, contradicting traditional understanding.
Global convergence of multilayer neural networks proven for any depth.
problem Global convergence of multilayer neural networks in the mean field regime.
method Mean field limit framework, neuronal embedding, bidirectional diversity condition.
result Global convergence for multilayer networks of any depths, including correlated initializations.
Study on market entry timing in stock liquidation with trading constraints.
problem Optimal timing of market entry and exit in portfolio liquidation with trading restrictions.
method Mean-field game approach to model N-player and mean-field games of optimal portfolio liquidation. result Existence of unique equilibrium in both mean-field and N-player games. We develop a mathematically rigorous framework for multilayer neural networks in the mean field regime. As the network's widths increase, the network's learning trajectory is shown to be well captured by a meaningful and dynamically nonlinear limit (the \textit{mean field} limit), which is characterized by a system of …
Study on price formation among investors with exponential utility and liabilities.
problem Equilibrium price formation among investors with heterogeneous risk-averseness and liabilities.
method Mean-field game theory and mean-field backward stochastic differential equations (BSDE).
result Existence of equilibrium risk-premium process and market clearing in the large population limit.
Gradient descent converges to minimum Bayes risk for two-layer ReLU networks in mean field regime.
problem Training two-layer ReLU networks using gradient descent in the mean field regime.
method Describes a condition for convergence to minimum Bayes risk, extending previous results to ReLU-activated networks.
result The condition for convergence does not depend on initialization and concerns weak convergence of network realization.
This work shows linear convergence for two-layer neural networks in mean-field regime.
problem Optimizing two-layer neural networks in the mean-field regime.
method Mean-field analysis and continuous-time noisy gradient descent.
result Establishes linear convergence rate for two-layer neural networks.
Paper proposes a mean-field gradient descent for zero-sum games, proving convergence to Nash equilibrium.
problem Finding mixed Nash equilibria in zero-sum games with multiple players.
method Mean-field gradient descent dynamics with time-averaging, incorporating exponentially discounted gradients.
result Exponential convergence rate to mixed Nash equilibrium with respect to total variation metric.
Paper analyzes convergence rates of mean-field SVGD method.
problem Establishing quantitative rates of convergence for mean-field SVGD.
method Quantitative analysis of mean-field SVGD dynamics on torus.
result Explicit polynomial convergence rates in L2-norm for Riesz-type kernels.
We study the mean field games equations, consisting of the coupled Kolmogorov-Fokker-Planck and Hamilton-Jacobi-Bellman equations. The equations are complemented by initial and terminal conditions. It is shown that with some specific choice of data, this problem can be reduced to solving a quadratically nonlinear syste…
Training recurrent neural networks (RNNs) on long sequence tasks is plagued with difficulties arising from the exponential explosion or vanishing of signals as they propagate forward or backward through the network. Many techniques have been proposed to ameliorate these issues, including various algorithmic and archite…
The study analyzes deep linear networks from random initialization, capturing dynamics and hyperparameter effects.
problem Understanding training dynamics in deep linear networks from random initialization.
method Theoretical analysis of gradient descent dynamics in deep linear networks with random initialization and large data.
result Captures the 'wider is better' effect and hyperparameter transfer effects, contrasting with neural-tangent parameterization.
This work analyzes a two-stage algorithm for single index models, showing precise asymptotics of gradient descent.
problem Learning single index models with non-convex optimization.
method Spectral initialization followed by gradient descent, with detailed analysis of dynamics and asymptotics.
result Gradient descent converges to long-time fixed points in the large system limit, representing mean field behavior.
Reducing the precision of weights and activation functions in neural network training, with minimal impact on performance, is essential for the deployment of these models in resource-constrained environments. We apply mean-field techniques to networks with quantized activations in order to evaluate the degree to which …
We develop a mean-field theory for multi-component ICA in high dimensions.
problem Understanding multi-component ICA in high-dimensional settings.
method Asymptotically exact mean-field theory for multi-component online ICA.
result Explicit learnability boundaries and competition conditions linking step size, data moments, and initialization.
Softmax policy gradient achieves global optimality in wide neural networks with entropy regularization.
problem Optimizing softmax policies with neural networks in the mean-field regime.
method Modeling neural networks as Wasserstein gradient flows and proving global optimality of fixed points.
result Global optimality of softmax policy gradient in wide single hidden layer neural networks with entropy regularization.
A framework solves parametric families of MFGs efficiently.
problem Efficiently solving MFG systems with varying initial distributions and terminal costs.
method Operator learning framework for parametric families of MFGs.
result Accurate approximation for cybersecurity and quadratic MFGs.
Study shows policy gradient convergence for entropy-regularized MDPs with neural nets in mean-field regime.
problem Global convergence of policy gradient for entropy-regularized MDPs with neural network approximation.
method Softmax policy with neural network approximation in mean-field regime, gradient flow in 2-Wasserstein metric, exponential convergence under sufficient regularization.
result Gradient flow converges exponentially fast to the unique stationary solution under sufficient regularization.
Study Transformer layers under cross-entropy training using mean field control.
problem Understanding the behavior of Transformer layers in cross-entropy training.
method Continuous-depth mean field control analysis, treating depth as time and layer parameters as controls.
result Derivation of a Pontryagin condition for the limiting population problem, involving the softmax residual.
Study on learning rates in neural networks of varying depth.
problem Dependence of maximal update learning rate on network depth.
method Analysis of random fully connected ReLU networks with mean-field weight initialization.
result Maximal update learning rate scales like L−3/2 with network depth. Develops asset pricing models with mean field game theory for heterogeneous agents.
problem Tackles equilibrium asset pricing in incomplete markets with heterogeneous agents.
method Uses mean field game theory and mean field backward stochastic differential equations (BSDEs).
result Derives equilibrium risk premium and shows market clearing in the large population limit.
This paper goes beyond the optimal trading Mean Field Game model introduced by Pierre Cardaliaguet and Charles-Albert Lehalle in [Cardaliaguet, P. and Lehalle, C.-A., Mean field game of controls and an application to trade crowding, Mathematics and Financial Economics (2018)]. It starts by extending it to portfolios of…
We analyze single-layer neural networks with the Xavier initialization in the asymptotic regime of large numbers of hidden units and large numbers of stochastic gradient descent training steps. The evolution of the neural network during training can be viewed as a stochastic system and, using techniques from stochastic…
Wide neural networks learn features under μP, identifying weights and decomposing support.
problem Feature learning in wide neural networks under μP. method Proving mean-field limit, characterizing identifiability, sparse-dictionary decomposition, and feature-learning-error decomposition.
result The triple (w∗,Dorb∗,S∗) identifies the natural learning cell of the architecture-data pair (σ,ρ). In recent years, state-of-the-art methods in computer vision have utilized increasingly deep convolutional neural network architectures (CNNs), with some of the most successful models employing hundreds or even thousands of layers. A variety of pathologies such as vanishing/exploding gradients make training such deep n…
MF-PID uses interacting samples to efficiently transport probability mass.
problem Efficiently transporting probability mass in generative models.
method Introducing Mean-Field Path-Integral Diffusion (MF-PID) where samples become interacting agents.
result MF-PID achieves 19-24% reductions in control energy for demand-response control of energy systems.
New framework for understanding infinite-width neural networks.
problem Understanding the infinite-width limit behavior of neural networks.
method General framework to study limit behavior of neural models based on hyperparameter scaling.
result Derives scaling for existing mean-field and neural tangent kernel limits and introduces new dynamically stable limits.
We propose a mean field game model to study the question of how centralization of reward and computational power occur in Bitcoin-like cryptocurrencies. Miners compete against each other for mining rewards by increasing their computational power. This leads to a novel mean field game of jump intensity control, which we…
Study on fluctuations in neural network kernels and predictions, focusing on finite width effects.
problem Characterizing fluctuations in finite width neural networks.
method Dynamical mean field theory analysis of wide but finite feature learning neural networks.
result Fluctuations in kernels and predictions are dynamically coupled, leading to reduced variance in feature learning regimes.
Gradient descent dynamics in wide neural networks are analyzed using a dynamical CLT.
problem Understanding the fluctuations in wide shallow neural networks trained via gradient descent.
method Dynamical Central Limit Theorem (CLT) applied to neural network dynamics.
result Asymptotic fluctuations remain bounded in mean square throughout training.
This paper analyzes MFVBI for GMM using statistical mechanics.
problem Approximate fast computation of Gaussian Mixture Model.
method Statistical mechanics and MFVBI applied to GMM.
result Rigorous analysis and mathematical foundation for MFVBI applied to GMM.
Unified analysis of DLNs using DMFT reveals dynamics of loss convergence and generalization trade-offs.
problem Understanding the overall dynamics of diagonal linear networks (DLNs) in neural network training.
method Dynamical Mean-Field Theory (DMFT) applied to DLNs.
result Derives low-dimensional effective process capturing high-dimensional gradient flow dynamics.
We introduce and analyze a linear kinetic model that describes the evolution of the probability density of the number of firms in a society, in which the microscopic rate of change obeys to the so-called law of proportional effect proposed by Gibrat. Despite its apparent simplicity, the possible mean field limits of th…
The dynamics of DNNs during gradient descent is described by the so-called Neural Tangent Kernel (NTK). In this article, we show that the NTK allows one to gain precise insight into the Hessian of the cost of DNNs. When the NTK is fixed during training, we obtain a full characterization of the asymptotics of the spectr…
New framework for portfolio management using binomial markets and game theory.
problem Investment behavior in competitive and incomplete markets.
method Introduces PRFPP framework, constructs and analyzes for both finite and mean field games.
result Relative performance concerns do not always lead to more risky asset investment.
Two neural network methods solve the master equation for MFGs.
problem Approximating Nash equilibria in stochastic, finite-agent games.
method Backward induction and direct PDE tackling neural networks.
result Neural networks can approximate the master equation's solution.
Theory proposes neural networks can be initialized for optimal information transmission.
problem Optimizing neural networks for optimal information transmission and representation.
method Developed a corrected mean-field framework to study neural networks as information channels, proving mutual information maximization at dynamic isometry.
result Mutual information maximization is realized between inputs and propagated signals when neural networks are initialized at dynamic isometry.
Residual networks (ResNet) and weight normalization play an important role in various deep learning applications. However, parameter initialization strategies have not been studied previously for weight normalized networks and, in practice, initialization methods designed for un-normalized networks are used as a proxy.…
This work studies fluctuation in multilayer neural networks using mean field theory.
problem Understanding fluctuation in multilayer neural networks with mean field training.
method Developed a second-order mean field limit to capture fluctuation, demonstrating stability of gradient descent training.
result Gradient descent training in multilayer networks biases towards minimal fluctuation, even after convergence.
Recurrent neural networks have gained widespread use in modeling sequence data across various domains. While many successful recurrent architectures employ a notion of gating, the exact mechanism that enables such remarkable performance is not well understood. We develop a theory for signal propagation in recurrent net…
New method solves supercooled Stefan problem, proving minimal solutions are physical.
problem Evolution of solid-liquid boundary in substances below freezing point.
method Construct solutions through McKean-Vlasov equation, proving tightness and propagation of chaos.
result Minimal solutions of McKean-Vlasov equation are physical under integrable initial conditions.
We give a rigorous analysis of the statistical behavior of gradients in a randomly initialized fully connected network N with ReLU activations. Our results show that the empirical variance of the squares of the entries in the input-output Jacobian of N is exponential in a simple architecture-dependent constant beta, gi…
Paper adds Fisher Information to mean field optimization for faster convergence.
problem Mean field optimization in neural networks training.
method Developed energy-dissipation method and gradient flow on probability space.
result Marginal distributions converge exponentially to minimizer.
New method tackles incomplete data in RBM inverse Ising problems.
problem Computing data and model expectations in inverse Ising problems with missing observations.
method Combines mean-field approximation, persistent contrastive divergence, and spatial Monte Carlo integration.
result Effective and accurate tuning of model parameters compared to conventional methods.