Learning rate annealing helps even in convex problems, improving generalization.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Learning rate annealing improves robustness in stochastic optimization.
Stochastic gradient descent with a large initial learning rate is widely used for training modern neural net architectures. Although a small initial learning rate allows for faster training and better test performance initially, the large learning rate achieves better generalization soon after the learning rate is anne…
CR-AIS improves AIS efficiency by constant rate annealing.
Optimal learning rate schedules derived for various tasks.
Mathematical analysis shows annealing prevents mode collapse in Gaussian mixtures.
In our recent paper, we showed that in exponential family, contrastive divergence (CD) with fixed learning rate will give asymptotically consistent estimates \cite{wu2016convergence}. In this paper, we establish consistency and convergence rate of CD with annealed learning rate . Specifically, suppose CD- gener…
The study optimizes simulated annealing's cooling schedule for better performance.
We describe an adaptation of the simulated annealing algorithm to nonparametric clustering and related probabilistic models. This new algorithm learns nonparametric latent structure over a growing and constantly churning subsample of training data, where the portion of data subsampled can be interpreted as the inverse …
Study convergence of simulated annealing in continuous and discrete settings.
Stochastic Gradient Descent shows directional bias with moderate learning rates, impacting optimization outcomes.
A new Q-learning controller improves line follower robot control.
New method uses reinforcement learning to improve Simulated Annealing.
In this paper we propose a modified version of the simulated annealing algorithm for solving a stochastic global optimization problem. More precisely, we address the problem of finding a global minimizer of a function with noisy evaluations. We provide a rate of convergence and its optimized parametrization to ensure a…
Annealed Entropic Allocation improves ranking and selection by mitigating hard switching and improving finite-budget discrimination.
We find optimal learning rate schedules for a random feature model.
We solve a multi-period portfolio optimization problem using D-Wave Systems' quantum annealer. We derive a formulation of the problem, discuss several possible integer encoding schemes, and present numerical examples that show high success rates. The formulation incorporates transaction costs (including permanent and t…
aMCL uses annealing to improve hypothesis diversity in ambiguous tasks.
Riemannian stochastic gradient descent converges faster with increasing batch size.
Proposes a method to improve SLMC for multimodal distributions.
CWGD measures gradient diversity weighted by curvature, improving SGD convergence.
Quantum annealers aim at solving non-convex optimization problems by exploiting cooperative tunneling effects to escape local minima. The underlying idea consists in designing a classical energy function whose ground states are the sought optimal solutions of the original optimization problem and add a controllable qua…
Quantum Annealing Enhanced Reinforcement Learning for Accurate RUL Prediction
Proposes a new sampling policy for ranking and selection problems.
Reverse annealing boosts quantum matrix factorization performance.
The paper analyzes the InfoNCE loss under different temperature schedules using Langevin dynamics.
KL annealing helps VAEs avoid posterior collapse and overfitting.
We introduce a novel framework for adversarial training where the target distribution is annealed between the uniform distribution and the data distribution. We posited a conjecture that learning under continuous annealing in the nonparametric regime is stable irrespective of the divergence measures in the objective fu…
Quantum machine learns faster by reverse annealing on AQCs.
The term structure of interest rates or yield curve is a function relating the interest rate with its own term. Nonlinear regression models of Nelson-Siegel and Svensson were used to estimate the yield curve using a sample of historical data supplied by the National Stock Exchange of Costa Rica. The optimization proble…
NVA combines variational posteriors, annealing, and natural-gradient learning for multimodal optimization.
DAIS improves AIS for differentiable marginal likelihood estimation.
CoolMomentum combines momentum and Simulated Annealing for deep learning optimization.
Simulated annealing is a popular method for approaching the solution of a global optimization problem. Existing results on its performance apply to discrete combinatorial optimization where the optimization variables can assume only a finite set of possible values. We introduce a new general formulation of simulated an…
We empirically evaluate a stochastic annealing strategy for Bayesian posterior optimization with variational inference. Variational inference is a deterministic approach to approximate posterior inference in Bayesian models in which a typically non-convex objective function is locally optimized over the parameters of t…
D-Wave quantum annealers represent a novel computational architecture and have attracted significant interest, but have been used for few real-world computations. Machine learning has been identified as an area where quantum annealing may be useful. Here, we show that the D-Wave 2X can be effectively used as part of an…
Maximum likelihood estimation (MLE) is one of the most important methods in machine learning, and the expectation-maximization (EM) algorithm is often used to obtain maximum likelihood estimates. However, EM heavily depends on initial configurations and fails to find the global optimum. On the other hand, in the field …
Despite the advances in the representational capacity of approximate distributions for variational inference, the optimization process can still limit the density that is ultimately learned. We demonstrate the drawbacks of biasing the true posterior to be unimodal, and introduce Annealed Variational Objectives (AVO) in…
Quantum computing optimizes ESG portfolios efficiently.
WSqD extends learning rate schedules for large model training without fixed horizons.
Quantum annealer speeds up RBM training for image classification.
Bayesian networks are a class of popular graphical models that encode causal and conditional independence relations among variables by directed acyclic graphs (DAGs). We propose a novel structure learning method, annealing on regularized Cholesky score (ARCS), to search over topological sorts, or permutations of nodes,…
We present an algorithm for learning a latent variable generative model via generative adversarial learning where the canonical uniform noise input is replaced by samples from a graphical model. This graphical model is learned by a Boltzmann machine which learns low-dimensional feature representation of data extracted …
This study optimizes currency arbitrage using quantum computing methods.
Introduces q-paths for generalizing geometric annealing paths in machine learning.
The convergence rate and final performance of common deep learning models have significantly benefited from heuristics such as learning rate schedules, knowledge distillation, skip connections, and normalization layers. In the absence of theoretical underpinnings, controlled experiments aimed at explaining these strate…
Quantum machine learns to clean up blurry images.
SKT improves EKI for Bayesian inverse problems with non-Gaussian targets.