Learning-rate schedules for large models match optimization theory closely, leading to better training.
problem Improving training of large models with optimal learning rates.
method Used a bound from non-smooth convex optimization theory to match learning-rate schedules with practical benefits.
result Extending the learning-rate schedule with optimal learning-rate and transferring it across schedules improves model training.
Theory extends optimal learning rates without realizability assumption.
problem Agnostic binary classification without realizability assumption.
method Identifies tetrachotomy of optimal rates and combinatorial structures.
result Optimal universal rates for binary classification in agnostic setting.
Refined theorem on linear perturbations with applications in singularity theory and optimization.
problem Linear perturbations and their implications in singularity theory and optimization.
method New perspective of Hausdorff measures for refined transversality theorem.
result Applications in singularity theory and optimization.
Develops pathwise analysis for log-optimal portfolios using rough paths theory.
problem Analyzing stability and approximation of log-optimal portfolios.
method Pathwise approach based on càdlàg rough paths theory.
result Establishes pathwise stability and error estimates for log-optimal portfolios.
Lecture notes on linear neural networks for deep learning optimization and generalization.
problem Understanding optimization and generalization in deep learning models.
method Mathematical tools and dynamical systems theory.
result Potential of mathematical tools to enhance understanding of deep learning.
BraidNet uses braid theory to optimize neural networks for image classification.
problem Image classification problems
method Procedural optimization of neural networks combining information theory and braid theory
result BraidNet outperforms other networks in learning speed and accuracy
The optimal approach is to theorize after examining data, not before.
problem Optimal sequencing of theory and empirical analysis for economic questions.
method Formalized a Bayesian model to trade off Darwinian and Statistical Learning.
result Post hoc theorizing is typically optimal in modern economics.
Attempts from different disciplines to provide a fundamental understanding of deep learning have advanced rapidly in recent years, yet a unified framework remains relatively limited. In this article, we provide one possible way to align existing branches of deep learning theory through the lens of dynamical system and …
We developed a strategic of optimal portfolio based on information theory and Tsallis statistics. The growth rate of a stock market is defined by using q-deformed functions and we find that the wealth after n days with the optimal portfolio is given by a q-exponential function. In this context, the asymptotic optim…
These notes constitute a sort of Crash Course in Optimal Transport Theory. The different features of the problem of Monge-Kantorovitch are treated, starting from convex duality issues. The main properties of space of probability measures endowed with the distances Wp induced by optimal transport are detailed. The ke…
Braid theory optimizes neural network structures.
problem Optimizing the architecture of neural networks.
method Using braid theory to describe and construct neural network structures.
result Braid-based networks outperform other architectures in classification tasks.
Optimal portfolio yields a digital option payoff.
problem Portfolio optimization under generalized dual theory of choice.
method Characterized optimal solution and derived it in closed form.
result Payoff is a digital option that yields in-the-money payoff in good market scenarios.
This work broadens optimal transport map estimation theory to stochastic settings.
problem Existing theory for optimal transport map estimation is restricted to deterministic maps under specific conditions.
method Introduces a novel metric for evaluating stochastic maps, develops computationally efficient estimators with robust guarantees.
result First general-purpose theory for map estimation compatible with real-world stochastic applications.
Combines option pricing and portfolio theory for optimal hedging.
problem Optimal hedging of European options in various price dynamics.
method Derives optimal holdings and unhedged risk for different price dynamics.
result Derives solutions for various price dynamics including binomial, diffusion, volatility, volatility-of-volatility, and jump diffusion.
Improved optimal regularity for harmonic almost complex structures.
problem Establishing optimal regularity for harmonic almost complex structures.
method Quantitative stratification method and rectifiability of singular strata.
result Optimal regularity theory for energy minimizing harmonic almost complex structures.
Most decision theories, including expected utility theory, rank dependent utility theory and cumulative prospect theory, assume that investors are only interested in the distribution of returns and not in the states of the economy in which income is received. Optimal payoffs have their lowest outcomes when the economy …
Optimizes electric field to control molecule states in Hartree-Fock theory.
problem Optimizing electric field to drive molecule from initial to target state.
method Trust region optimization with gradients from adjoint state method.
result Achieves desired target states with minimal control effort.
Optimizes portfolios with utility theory, diversification, and leverage.
problem Finding optimal portfolio allocation strategies.
method Utility theory, exponential and logarithmic utilities, compound probability distributions, maximum expected utility, generalized mean-variance.
result Enhanced portfolio allocation strategies with natural explanations.
Proposes a new portfolio theory that optimizes returns and risk.
problem Inefficient market hypothesis and risk premium in finance markets.
method Introduces triplet (R, H, σ) model for portfolio optimization.
result Developed a global optimal strategy for different investor styles.
Theory proposes neural networks can be initialized for optimal information transmission.
problem Optimizing neural networks for optimal information transmission and representation.
method Developed a corrected mean-field framework to study neural networks as information channels, proving mutual information maximization at dynamic isometry.
result Mutual information maximization is realized between inputs and propagated signals when neural networks are initialized at dynamic isometry.
Develops theory for conditional optimal transport in infinite-dimensional spaces.
problem Bayesian inference with functional parameters in infinite-dimensional spaces.
method Theory of constrained optimal transport for block-triangular maps.
result Regularity estimates on conditioning maps from prior to posterior.
We study optimal investment problems under the framework of cumulative prospect theory (CPT). A CPT investor makes investment decisions in a single-period financial market with transaction costs. The objective is to seek the optimal investment strategy that maximizes the prospect value of the investor's final wealth. W…
Study preferences over uncertain time payments, finds growth-optimality better than expected utility theory.
problem Understanding how people make decisions with uncertain timing of payments.
method Normative model of growth-optimality, revisiting experimental evidence on time lotteries.
result Growth-optimality better explains experimental data on time lotteries than expected discounted utility theory.
The paper analyzes privacy leakage in federated learning using linear algebra and optimization theory.
problem Privacy leakage in federated learning despite its promise for data privacy.
method Theoretical analysis from linear algebra and optimization theory perspectives.
result Derives sufficient conditions to prevent data reconstruction attacks and establishes an upper bound on privacy leakage.
GT-DDP optimizer trains residual networks using game theory.
problem Applying DDP to residual networks.
method Game-theoretic DDP optimizer for residual networks.
result Improved training convergence and variance reduction.
PAC-Bayesian theory applied to learning optimization algorithms with generalization guarantees.
problem Learning optimization algorithms with provable generalization guarantees and explicit trade-offs.
method PAC-Bayes theory applied to learning-to-optimize, reformulating the learning procedure into a one-dimensional minimization problem.
result Learned optimization algorithms outperform deterministic worst-case analysis algorithms, even in the limit case of guaranteed convergence.
Study of 2D Lorentzian anti-de Sitter plane using geometric control theory.
problem Understanding extremal trajectories and reachable set on anti-de Sitter plane.
method Geometric control theory and differential geometry.
result Construction of optimal synthesis and description of Lorentzian distance.
The paper develops a new theory to understand deep learning optimization.
problem Understanding the dynamics of optimization in deep learning, especially in the edge of stability regime.
method Developed a central flow differential equation to describe the time-averaged trajectory of oscillatory optimizers.
result Central flows can predict long-term optimization trajectories with high numerical accuracy.
Some optimization or equilibrium problems involving somehow the concept of optimal transport are presented in these notes, mainly devoted to applications to economic and game theory settings. A variant model of transport, taking into account traffic congestion effects is the first topic, and it shows various links with…
The pathwise coordinate optimization is one of the most important computational frameworks for high dimensional convex and nonconvex sparse learning problems. It differs from the classical coordinate optimization algorithms in three salient features: {\it warm start initialization}, {\it active set updating}, and {\it …
Study optimal stopping times for multi-dimensional processes with non-exponential discounting.
problem Optimal stopping in multi-dimensional processes with non-exponential discounting.
method Probabilistic potential theory to establish existence of optimal equilibria.
result Existence of optimal equilibria for multi-dimensional stopping problems.
Optimizes communication in federated learning using rate-distortion theory.
problem Reduces communication cost in federated learning while maintaining model accuracy.
method Applies rate-distortion theory to model updates, proposing distortion as a proxy for accuracy.
result Near-optimal communication reduction, outperforming other methods on a FL benchmark.
New approach to optimal income tax theory tackles inequity issues.
problem Optimal tax schedules often lead to minimal tax rates for higher earners, contradicting ethical practices.
method Developed a theorem for piecewise-linear environment and introduced a new utility function parameter.
result New approach leads to more equitable tax schedules, interpreting optimality criteria easily.
Modern statistical inference tasks often require iterative optimization methods to compute the solution. Convergence analysis from an optimization viewpoint only informs us how well the solution is approximated numerically but overlooks the sampling nature of the data. In contrast, recognizing the randomness in the dat…
Unified theory for semi-implicit variational inference, bridging approximation and optimization.
problem Developing a statistical theory for semi-implicit variational inference.
method Unified theory combining approximation and optimization analyses.
result Unified theory characterizes SIVI's ability to recover target distributions and governs asymptotic behavior.
Bayesian optimization surveys information-theoretic acquisition functions.
problem Optimizing noisy, expensive, non-convex functions with unknown gradients.
method Bayesian optimization using Gaussian process surrogate models and information-theoretic acquisition functions.
result Information-theoretic acquisition functions outperform others in real scenarios.
The paper generalizes optimization algorithms using category theory.
problem Optimizing functions in a category-theoretic setting.
method Using the Cartesian reverse derivative to generalize gradient descent and Newton's method.
result Properties of optimization algorithms are preserved in the generalized setting, including invariances and convergence.
Quantum field theory connects Riemannian geometry to quantum fluctuations.
problem Generating Riemannian structures from quantum fluctuations.
method QFT approach to Riemannian Geometry, focusing on Ricci curvature.
result Ricci curvature is crucial in generating Riemannian structures.
This paper optimizes sports betting strategies using neural networks and portfolio theory.
problem Optimizing betting strategies in sports gambling.
method Combining neural network models with portfolio optimization, integrating Von Neumann-Morgenstern Expected Utility Theory and the Kelly Criterion.
result Achieved 135.8% relative profit during the English Premier League season.
Active inference minimizes expected free energy for optimal behavior.
problem Understanding and optimizing behavior in complex systems.
method Combines Bayesian decision theory, optimal Bayesian design, and the free energy principle.
result Active inference emerges as a unified framework for information-seeking, utility maximization, and goal-directed behavior.
Introduces a new geometric method for optimal experimental design.
problem Restrictive invariance properties of traditional OED approaches based on probability densities.
method Mutual transport dependence (MTD) using optimal transport theory.
result Demonstrates high-quality designs and flexibility compared to standard methods.
Optimal Control Theory optimizes neural networks, improving robustness and efficiency.
problem Optimizing deep neural networks (DNNs) for better performance and efficiency.
method Integrating Optimal Control Theory with Backpropagation to develop a new optimizer.
result Optimal Control Theoretic Neural Optimizer (OCNOpt) improves upon existing methods in robustness and efficiency.
Develops a new method for optimizing portfolios in stochastic markets.
problem Optimizing functionally generated portfolios in stochastic portfolio theory.
method Optimizes over a family of rank-based portfolios parameterized by an exponentially concave function.
result Proves existence and uniqueness of the optimization problem and provides stability estimates.
A general study of symmetries in optimal control theory is given, starting from the presymplectic description of this kind of system. Then, Noether's theorem, as well as the corresponding reduction procedure (based on the application of the Marsden-Weinstein theorem adapted to the presymplectic case) are stated both in…
Study connects covariance cleaning theory to information theory for heavy-tailed distributions.
problem Optimizing covariance matrices for heavy-tailed distributions using information theory.
method Minimizing Frobenius norm and information loss between true and estimated covariance matrices.
result Asymptotic regime of large matrices minimizes information loss for Student's t distributions.
We study sparse approximate solutions to convex optimization problems. It is known that in many engineering applications researchers are interested in an approximate solution of an optimization problem as a linear combination of elements from a given system of elements. There is an increasing interest in building such …
Develops regularity theory for Beckmann's optimal transport problem.
problem Minimizing total squared flux in continuous transport from source to target.
method Unconstrained Lagrangian formulation, variational first order optimality conditions, Schauder estimates.
result Exact Hölder regularity of potential, flux, and flow generating on bounded, regular domains.
Radial Basis Functions Neural Networks (RBFNNs) are tools widely used in regression problems. One of their principal drawbacks is that the formulation corresponding to the training with the supervision of both the centers and the weights is a highly non-convex optimization problem, which leads to some fundamentally dif…