Learning-rate schedules for large models match optimization theory closely, leading to better training.
problem Improving training of large models with optimal learning rates.
method Used a bound from non-smooth convex optimization theory to match learning-rate schedules with practical benefits.
result Extending the learning-rate schedule with optimal learning-rate and transferring it across schedules improves model training.
PAGE optimizes nonconvex problems with optimal convergence rates.
problem Nonconvex optimization problems.
method PAGE algorithm for achieving optimal convergence rates.
result PAGE achieves optimal convergence rates for nonconvex optimization.
Paper improves learning rates for SGD and NAG.
problem Generalization performance of stochastic optimization algorithms.
method Establishes new learning rates for SGD and NAG.
result Improved guarantees in some settings or comparable rates under weaker assumptions.
The paper analyzes how learning rate affects SGD and provides insights into optimal rates.
problem Understanding the impact of learning rate on stochastic gradient descent.
method Developed a learning-rate-dependent stochastic differential equation (lr-dependent SDE) to analyze SGD.
result Established a linear rate of convergence for SGD and found the optimal linear rate by analyzing the spectrum of the Witten-Laplacian.
Theory extends optimal learning rates without realizability assumption.
problem Agnostic binary classification without realizability assumption.
method Identifies tetrachotomy of optimal rates and combinatorial structures.
result Optimal universal rates for binary classification in agnostic setting.
Study on optimal capital injection for insurance companies under negative interest rates.
problem Optimal capital injection behavior of insurance companies under negative interest rates.
method Modelled the surplus process as a Brownian motion with drift and interest rate changes as a Markov-switching process. Established an algorithm for finding the value function and optimal strategy.
result Optimal strategy involves holding a positive reserve when interest rate is negative, unlike when positive.
Paper reconciles minimax rates and optimal recovery rates for noisy observations.
problem Estimating a function from noisy observations.
method Develops NLA minimax rates for Besov classes in Lq-norms. result NLA minimax rates continuously depend on noise level and match optimal recovery rates as noise decreases.
Optimal classification rules control error rates in multiclass mixture models.
problem Classifying observations in multiclass mixture models while controlling error rates.
method Finding optimal classification rules by searching an optimal region in the observation space, using Maximum A Posteriori (MAP) rule and heuristic computation.
result The FDR-like optimal rule can be significantly less conservative than thresholded MAP rules.
Study optimal dividends in dual risk model with stochastic interest rate.
problem Optimal dividend strategy in dual risk model with stochastic interest rate.
method Geometric Brownian motion or exponential Lévy process for discounting factor.
result Closed form solutions can be obtained for optimal dividends.
Study proposes optimal risk-aware interest rates for crypto lending protocols.
problem Determining optimal interest rates for decentralized lending protocols to maximize profit and minimize risk.
method Agent-based model, Riccati-type ODEs for linear behaviors, Monte-Carlo estimator and deep learning for nonlinear behaviors.
result Calibrated model shows superior risk-adjusted performance compared to industry-standard interest rate models.
MARTHE optimizes learning rates online using hypergradient approximations.
problem Optimizing task-specific learning rates for better generalization.
method Online algorithm guided by hypergradient approximations, interpolating between RTHO and HD.
result Produces more stable learning rate schedules leading to better model generalization.
Optimizes binary rating systems for item ranking.
problem Designing efficient feedback systems for item ranking.
method Formalizes performance, provides algorithm, empirically designs and validates.
result Empirically designed and validated approximately optimal rating system.
A new game-theoretic approach optimizes complex rate metrics.
problem Optimizing non-decomposable performance metrics and rate constraints.
method Extending two-player game approaches to a three-player game, seeking equilibrium.
result Generalizes and improves upon existing algorithms for constrained optimization.
Study optimizes dividend payout strategies under fluctuating interest rates.
problem Maximizing dividends under stochastic interest rates with negative values.
method Analytical HJB approach and backward SDEs for analysis.
result Explicit optimal strategies found for both time-dependent and strategy-independent stopping times.
Unified algorithm solves convex optimization problems with optimal rates.
problem Solving nonsmooth constrained convex optimization problems.
method Unified randomized block-coordinate primal-dual algorithm.
result Achieves optimal convergence rates of O(n/k) and O(n2/k2). New algorithms improve SGD's efficiency in convex and nonconvex optimization.
problem Optimizing gradient size in stochastic optimization.
method Designing SGD3 for convex objectives and SGD5 for nonconvex objectives.
result Near-optimal rates for gradient size reduction in both convex and nonconvex settings.
Study examines how data augmentation impacts optimization in linear regression.
problem Understanding how data augmentation schedules affect optimization in linear regression.
method Analyzed the effect of augmentation on optimization in linear regression with MSE loss, using classical convex optimization and recent work on implicit bias.
result Proved that under certain joint schedules for learning rate and augmentation scheme, augmented gradient descent converges and characterized the resulting minimum.
Cyclical learning rate improves neural machine translation performance.
problem Optimizing learning rate for neural machine translation.
method Applied cyclical learning rate to transformer-based neural networks.
result Cyclical learning rate significantly impacts neural machine translation performance.
Optimal rates for shallow ReLU networks in nonparametric regression.
problem Approximating smooth and non-smooth functions with shallow ReLU networks.
method Analysis of shallow ReLUk neural networks, using variation norms and deep learning theory. result Optimal approximation rates for shallow ReLU networks in nonparametric regression.
A new method for optimizing non-decomposable metrics with constraints.
problem Optimizing complex machine learning objectives with thresholded constraints.
method Formulate rate-constrained optimization using the Implicit Function theorem and solve with gradient-based methods.
result Demonstrated effectiveness over existing methods on benchmark datasets.
New insights into optimizing Local SGD's outer optimizer for faster convergence.
problem Understanding the impact of outer optimizer and its hyperparameters in Local SGD.
method Analyzing convergence guarantees with new outer learning rates and momentum.
result Tuning the outer learning rate can improve convergence and handle inner learning rate ill-tuning.
This work analyzes the convergence rate of unrolling for optimizing quadratic objectives.
problem The challenge of accurately computing Jacobians through optimization.
method Non-asymptotic convergence-rate analysis of unrolled differentiation for gradient descent and Chebyshev method.
result There is a trade-off between fast asymptotic convergence and immediate but slower convergence due to the learning rate.
Optimizes portfolio growth rate for a behavioral investor considering terminal relative growth rate.
problem Optimizing a behavioral investor's portfolio growth rate under relative growth criterion.
method Martingale method, concavification, and quantile optimization techniques.
result Derives closed-form optimal growth rate and finds significant impact of benchmark growth rate.
Con-TS optimizes wireless link throughput with latency constraints.
problem Optimizing rate selection for wireless links with latency constraints.
method Proposes Con-TS, a constrained Thompson sampling algorithm for stochastic MAB problems.
result Con-TS achieves upper bounds on expected constraint violations and throughput loss.
Unified framework for ESG-inclusive portfolio optimization and pricing.
problem Incorporating ESG ratings into dynamic asset pricing theory.
method Introducing ESG-valued return as a linear transformation of financial and ESG scores, preserving traditional risk aversion with an ESG affinity parameter.
result Developed a more complex portfolio optimization problem in a space governed by reward, risk, and ESG score.
AutoGD automatically adjusts learning rates for gradient descent.
problem Optimizing learning rates for gradient descent methods.
method AutoGD automatically adjusts learning rates based on iteration.
result AutoGD can recover the optimal rate of GD for a broad class of functions.
Method calibrates local volatility and stochastic short rate models for equity-rate dynamics.
problem Joint calibration of local volatility and stochastic short rate models.
method Iterative approach using semimartingale optimal transport.
result Demonstrated performance on market data using European SPX options and cap interest rate options.
The paper studies convergence rates of Tsallis entropic regularization in optimal transport.
problem Optimal transport with regularization.
method Γ-convergence and quantization/shadow arguments.
result Derives convergence rate of Tsallis entropic regularization.
The paper improves SVM learning rates for anisotropic Gaussian kernels.
problem Nonparametric regression with anisotropic Gaussian kernels.
method Establishing almost optimal learning rates for functions in anisotropic Besov spaces.
result Optimal learning rates up to logarithmic factors, faster than Sobolev space-based rates.
VAV method optimizes learning rate for faster, stable SGD convergence.
problem Optimizing learning rate for efficient and stable machine learning models.
method Energy-based self-adaptive learning rate with auxiliary variable r. result VAV method achieves faster convergence and superior stability with larger learning rates.
Optimizes identifying the best arm with fixed samples.
problem Finding the arm with the highest mean in a fixed number of samples.
method Characterizes minimax optimal rates and introduces algorithms to achieve them.
result Characterizes and introduces algorithms for optimal best arm identification.
Study on convergence rates for optimal transport with regularization.
problem Convergence analysis of divergence-regularized optimal transport.
method Novel methodology using quantization and martingale couplings.
result Sharp rates for various divergences and transport costs.
SALR improves deep learning generalization by dynamically adjusting learning rates.
problem Improving generalization in deep learning models.
method Sharpness-aware learning rate scheduling based on local loss function sharpness.
result SALR drives solutions to flatter regions, improving generalization and convergence.
Optimizes portfolios using anticipated interest rate information.
problem Maximizing utility in financial models with future interest rate trends.
method Enlargement of filtrations, affine diffusion process, Markov chain modeling.
result Explicit formulas for expected logarithmic utility.
Optimal dividend strategy with constraints on drawdown and ratcheting rates.
problem Maximizing discounted utility of dividends until bankruptcy with drawdown and ratcheting constraints.
method Formulated as a stochastic control problem, solved via Hamilton-Jacobi-Bellman variational inequality.
result Optimal dividend rate ct∗ depends on current surplus and historical maximum of dividend rate, with specific rules for different surplus levels. Paper proposes a new coin betting method for training deep networks without learning rates.
problem Deep learning requires tuning many hyperparameters, especially learning rates.
method Reduces deep network training to a coin betting game, eliminating learning rates.
result Empirical and theoretical evidence shows the new method outperforms existing stochastic gradient algorithms.
Optimal bounds proven for ordinal embedding convergence rate.
problem Optimal bounds for ordinal embedding convergence rate in 1D.
method Utilized results from additive number theory and conducted computational experiments.
result Proved optimal bounds for convergence rate in 1D.
Paper studies convergence rates from surrogate risk minimizers to Bayes optimal classifier.
problem Analyzing the convergence rates of surrogate risk minimizers to the Bayes optimal classifier.
method Introducing consistency intensity to characterize surrogate loss functions and using it to derive convergence rates.
result Empirical surrogate risk minimizers converge faster to the Bayes optimal classifier under certain conditions.
We establish minimax optimal rates of convergence for estimation in a high dimensional additive model assuming that it is approximately sparse. Our results reveal an interesting phase transition behavior universal to this class of high dimensional problems. In the {\it sparse regime} when the components are sufficientl…
D-Adaptation automatically sets optimal learning rates without manual tuning.
problem Optimizing learning rates for efficient convergence in machine learning.
method D-Adaptation, which asymptotically achieves optimal learning rates without back-tracking or additional evaluations.
result D-Adaptation automatically matches hand-tuned learning rates across diverse problems.
A new method automatically and dynamically sets learning rates in deep learning.
problem Determining the appropriate learning rate in deep learning tasks is challenging and often subjective.
method Local Quadratic Approximation (LQA) to automatically and dynamically set learning rates.
result The proposed method leads to nearly optimal learning rates in a computationally efficient way.
Study confirms optimal minimax rate for nonlocal interaction kernel estimation.
problem Estimating nonlocal interaction kernels in interacting particle systems.
method Introduced tamed least squares estimator (tLSE) achieving optimal convergence rate.
result Optimal minimax rate of convergence confirmed for β≥1/4. Meta learning estimator for Bayes classification error rates.
problem Estimating the optimal Bayes classifier's performance.
method Weighted nearest neighbor (WNN) graph estimator based on HP divergence.
result The estimator achieves rate-optimal mean squared estimation error.
Paper establishes universal lower bounds and optimal rates for clustering sub-exponential mixture models.
problem Achieving optimal error rates in clustering sub-exponential mixture models.
method Establishes universal lower bounds and demonstrates iterative algorithms' optimality in sub-exponential mixture models.
result Iterative algorithms achieve the universal lower bound in sub-exponential mixture models.
Optimal rates found for learning with Nyström stochastic gradient methods.
problem Nonparametric regression learning with improved computational efficiency.
method Combination of stochastic gradient methods with Nyström subsampling, allowing multiple passes and mini-batches.
result Derivation of optimal learning rates considering various parameters.
Develops a parameter-free SGD algorithm with optimal convergence rate.
problem Optimizing parameters in stochastic convex optimization.
method A novel parameter-free algorithm for SGD with high-probability guarantees and adaptive properties.
result Achieves optimal convergence rate with only a double-logarithmic factor increase compared to known-parameter settings.
Analyzes optimal learning rate schedules in high-dimensional non-convex optimization problems.
problem Optimizing high-dimensional non-convex loss landscapes.
method Langevin optimization with learning rate decaying as \(η(t) = t^{-β}\).
result To speed up optimization without getting stuck in saddles, a decay rate \(β < 1\) is optimal, contrary to convex setups where \(β = 1\).
Study optimizes learning rates for conditional mean embedding estimates.
problem Consistency of kernel ridge regression for conditional mean embedding.
method Adaptive statistical learning rate derived for misspecified setting.
result Upper bound matches optimal O(logn/n) rates without assuming finite dimensionality.