Adam is shown not being able to converge to the optimal solution in certain cases. Researchers recently propose several algorithms to avoid the issue of non-convergence of Adam, but their efficiency turns out to be unsatisfactory in practice. In this paper, we provide new insight into the non-convergence issue of Adam …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper studies a curious phenomenon in learning energy-based model (EBM) using MCMC. In each learning iteration, we generate synthesized examples by running a non-convergent, non-mixing, and non-persistent short-run MCMC toward the current model, always starting from the same initial distribution such as uniform no…
This paper studies the market phenomenon of non-convergence between futures and spot prices in the grains market. We postulate that the positive basis observed at maturity stems from the futures holder's timing options to exercise the shipping certificate delivery item and subsequently liquidate the physical grain. In …
Polyak-Ruppert CLT for SA-Adam with momentum and non-convergent adaptive preconditioning
Study shows non-convergence of short-maturity expansion in SABR model.
The paper constructs non-convergent solutions to Vafa-Witten equations with specific harmonic 2-form limits.
This paper extends stability analysis to non-convergent neural network training.
ADOPT optimizes Adam to converge with any β2 without bounded noise.
Improved sampling for Bayesian neural networks reduces vanishing acceptance rates and increases predictive accuracy.
In this paper, we show that Generative Adversarial Networks (GANs) suffer from catastrophic forgetting even when they are trained to approximate a single target distribution. We show that GAN training is a continual learning problem in which the sequence of changing model distributions is the sequence of tasks to the d…
SGD avoids critical points on weakly convex functions.
SGD methods fail to converge to global minimizers in deep neural networks with ReLU activation.
Survey of GANs challenges and solutions for better model design and optimization.
In this paper, we shall study the Dirichlet problem for the minimal surfaces equation. We prove some results about the boundary behaviour of a solution of this problem. We describe the behaviour of a non-converging sequence of solutions in term of lines of divergence in the domain. Using this second result, we build so…
Study convergence of Yamabe flow on singular spaces with positive constant.
We study the problem of dynamically trading a futures contract and its underlying asset under a stochastic basis model. The basis evolution is modeled by a stopped scaled Brownian bridge to account for non-convergence of the basis at maturity. The optimal trading strategies are determined from a utility maximization pr…
CR Yamabe flow fails to converge on small deformations of the standard CR three-sphere.
Markov chain Monte Carlo (MCMC) methods have not been broadly adopted in Bayesian neural networks (BNNs). This paper initially reviews the main challenges in sampling from the parameter posterior of a neural network via MCMC. Such challenges culminate to lack of convergence to the parameter posterior. Nevertheless, thi…
Generative Adversarial Network (GAN) is a current focal point of research. The body of knowledge is fragmented, leading to a trial-error method while selecting an appropriate GAN for a given scenario. We provide a comprehensive summary of the evolution of GANs starting from its inception addressing issues like mode col…
Stochastic Stein Discrepancies improve inference efficiency.
Persistently trained EBMs generate images and estimate complex densities.
Developed a new algorithm to improve dynamic treatment regimens.
First-order optimization algorithms have been proven prominent in deep learning. In particular, algorithms such as RMSProp and Adam are extremely popular. However, recent works have pointed out the lack of ``long-term memory" in Adam-like algorithms, which could hamper their performance and lead to divergence. In our s…
Despite the growing prominence of generative adversarial networks (GANs), optimization in GANs is still a poorly understood topic. In this paper, we analyze the "gradient descent" form of GAN optimization i.e., the natural setting where we simultaneously take small gradient steps in both generator and discriminator par…
Belief Propagation (BP) is one of the most popular methods for inference in probabilistic graphical models. BP is guaranteed to return the correct answer for tree structures, but can be incorrect or non-convergent for loopy graphical models. Recently, several new approximate inference algorithms based on cavity distrib…
SGD fails to converge for deep ReLU networks with limited random initializations.
Generative Adversarial Networks (GANs) have shown remarkable results in modeling complex distributions, but their evaluation remains an unsettled issue. Evaluations are essential for: (i) relative assessment of different models and (ii) monitoring the progress of a single model throughout training. The latter cannot be…
Generative adversarial networks (GANs) are notoriously difficult to train and the reasons underlying their (non-)convergence behaviors are still not completely understood. By first considering a simple yet representative GAN example, we mathematically analyze its local convergence behavior in a non-asymptotic way. Furt…
Approximate Markov chain Monte Carlo (MCMC) offers the promise of more rapid sampling at the cost of more biased inference. Since standard MCMC diagnostics fail to detect these biases, researchers have developed computable Stein discrepancy measures that provably determine the convergence of a sample to its target dist…
In recent years, Generative Adversarial Networks (GANs) have received significant attention from the research community. With a straightforward implementation and outstanding results, GANs have been used for numerous applications. Despite the success, GANs lack a proper theoretical explanation. These models suffer from…
This paper tackles sampling issues in latent space EBMs by introducing diffusion-based amortization.
This study compares parallel SMC and MCMC for Bayesian deep learning, showing SMC parallel is faster.
Paper studies stochastic optimization methods with momentum, proving convergence and avoiding traps.
Stochastic gradient descent approximates Gaussian process posteriors efficiently.
Training deep neural networks requires intricate initialization and careful selection of learning rates. The emergence of stochastic gradient optimization methods that use adaptive learning rates based on squared past gradients, e.g., AdaGrad, AdaDelta, and Adam, eases the job slightly. However, such methods have also …
In complex simulation environments, certain parameter space regions may result in non-convergent or unphysical outcomes. All parameters can therefore be labeled with a binary class describing whether or not they lead to valid results. In general, it can be very difficult to determine feasible parameter regions, especia…
Recently, there is a growing interest in the study of median-based algorithms for distributed non-convex optimization. Two prominent such algorithms include signSGD with majority vote, an effective approach for communication reduction via 1-bit compression on the local gradients, and medianSGD, an algorithm recently pr…
New NPG variants ensure parameter convergence in multi-agent learning.
Learning an efficient update rule from data that promotes rapid learning of new tasks from the same distribution remains an open problem in meta-learning. Typically, previous works have approached this issue either by attempting to train a neural network that directly produces updates or by attempting to learn better i…
Adaptive inference for -estimators in bandit data with model misspecification.
Koopman mode analysis applied to neural networks for training optimization.
Network Embedding is the task of learning continuous node representations for networks, which has been shown effective in a variety of tasks such as link prediction and node classification. Most of existing works aim to preserve different network structures and properties in low-dimensional embedding vectors, while neg…
A new method learns latent space normalizing flow for approximate inference in generator models.
Recent years have witnessed significant advances in reinforcement learning (RL), which has registered great success in solving various sequential decision-making problems in machine learning. Most of the successful RL applications, e.g., the games of Go and Poker, robotics, and autonomous driving, involve the participa…
The paper tackles temporal coverage bias in financial panel data, proposing a structuring framework to correct for incomplete histories.
Bayesian Optimization tackles hidden constraints in architecture optimization.
Linearized attention fails to converge to NTK limit even at large widths.
This paper clarifies Bitcoin's volatility and predictability across daily, weekly, and monthly scales.