New analysis improves sample complexity for vanilla policy gradient methods.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new method reduces variance in training early-stage rankers for large-scale search systems.
We present a novel method in the family of particle MCMC methods that we refer to as particle Gibbs with ancestor sampling (PG-AS). Similarly to the existing PG with backward simulation (PG-BS) procedure, we use backward sampling to (considerably) improve the mixing of the PG kernel. Instead of using separate forward a…
Deep reinforcement learning (DRL) on Markov decision processes (MDPs) with continuous action spaces is often approached by directly training parametric policies along the direction of estimated policy gradients (PGs). Previous research revealed that the performance of these PG algorithms depends heavily on the bias-var…
STORM-PG uses momentum for faster policy gradient updates.
A new method reduces variance in PG methods for RL, improving efficiency and convergence.
Humans are capable of building holistic representations for images at various levels, from local objects, to pairwise relations, to global structures. The interpretation of structures involves reasoning over repetition and symmetry of the objects in the image. In this paper, we present the Program-Guided Image Manipula…
CoPhy-PGNN tackles competing PG losses in neural networks for solving eigenvalue problems.
Consider the stochastic composition optimization problem where the objective is a composition of two expected-value functions. We propose a new stochastic first-order method, namely the accelerated stochastic compositional proximal gradient (ASC-PG) method, which updates based on queries to the sampling oracle using tw…
New PG methods tackle nonconvex optimization with auto-conditioned stepsizes.
New PG losses improve decision optimization in misspecified models.
Softmax PG methods can take extremely long to converge, even with exact gradients.
We address the problem of regret minimization in logistic contextual bandits, where a learner decides among sequential actions or arms given their respective contexts to maximize binary rewards. Using a fast inference procedure with Polya-Gamma distributed augmentation variables, we propose an improved version of Thomp…
Paper proposes FPG algorithm for unbiased off-policy PG estimation.
PC-PG balances exploration and exploitation in reinforcement learning.
This paper introduces a number of new intrinsically 3-linked graphs through five new constructions. We then prove that intrinsic 3-linkedness is not preserved by moves. We will see that the graph , which is obtained through a move on , is not intrinsically 3-linked.
Paper explores PG for MCR, finding suboptimal policies but providing bounds.
We show that on-policy policy gradient (PG) and its variance reduction variants can be derived by taking finite difference of function evaluations supplied by estimators from the importance sampling (IS) family for off-policy evaluation (OPE). Starting from the doubly robust (DR) estimator (Jiang & Li, 2016), we provid…
Infinite Hidden Markov Models (iHMM's) are an attractive, nonparametric generalization of the classical Hidden Markov Model which can automatically infer the number of hidden states in the system. However, due to the infinite-dimensional nature of transition dynamics performing inference in the iHMM is difficult. In th…
We consider the problem of minimizing a Lipschitz differentiable function over a class of sparse symmetric sets that has wide applications in engineering and science. For this problem, it is known that any accumulation point of the classical projected gradient (PG) method with a constant stepsize satisfies the $L…
Bayesian sOED uses PG reinforcement learning for efficient experiment design.
Let G be a connected Lie group, LG its loop group, and PG->G the principal LG-bundle defined by quasi-periodic paths in G. This paper is devoted to differential geometry of the Atiyah algebroid A=T(PG)/LG of this bundle. Given a symmetric bilinear form on the Lie algebra g and the corresponding central extension of Lg,…
ZDPG learns model-free policies without critics, improving on PG.
Monte Carlo Tree Search (MCTS) algorithms perform simulation-based search to improve policies online. During search, the simulation policy is adapted to explore the most promising lines of play. MCTS has been used by state-of-the-art programs for many problems, however a disadvantage to MCTS is that it estimates the va…
Variable metric proximal gradient (VM-PG) is a widely used class of convex optimization method. Lately, there has been a lot of research on the theoretical guarantees of VM-PG with different metric selections. However, most such metric selections are dependent on (an expensive) Hessian, or limited to scalar stepsizes l…
We study the convergence rate of stochastic optimization of exact (NP-hard) objectives, for which only biased estimates of the gradient are available. We motivate this problem in the context of learning the structure and parameters of Ising models. We first provide a convergence-rate analysis of deterministic errors fo…
New PG samplers improve inference in coupled state-space models.
Study on PG learning for LQ MFC problems with common noise, proving convergence and sample complexity.
New proof confirms surfaces can be divided into polygons.
Additive regression trees are flexible non-parametric models and popular off-the-shelf tools for real-world non-linear regression. In application domains, such as bioinformatics, where there is also demand for probabilistic predictions with measures of uncertainty, the Bayesian additive regression trees (BART) model, i…
A new framework improves reinforcement learning algorithms with policy guarantees.
New framework optimizes multi-asset portfolio choice for high dimensions.
ZOSPI improves RL policies with global value function exploitation.
Post-training optimizes model performance beyond base model limits.
PG-EVIKAL refines molecular property predictions using neighbor fusion and evidential neural networks.
Betten and Riesinger constructed Parallelisms of with automorphism group by applying the reducible -action to a rotational Betten spread. This was generalized by the present author so as to include oriented parallelisms (i.e., p…
Study policy gradient and actor-critic methods for continuous-time reinforcement learning.
Evolution Strategies (ES) are a powerful class of blackbox optimization techniques that recently became a competitive alternative to state-of-the-art policy gradient (PG) algorithms for reinforcement learning (RL). We propose a new method for improving accuracy of the ES algorithms, that as opposed to recent approaches…
Paper analyzes convergence of Adam-type RL algorithms under Markovian sampling.
We study continuous action reinforcement learning problems in which it is crucial that the agent interacts with the environment only through safe policies, i.e.,~policies that do not take the agent to undesirable situations. We formulate these problems as constrained Markov decision processes (CMDPs) and present safe p…
Improving the sample efficiency in reinforcement learning has been a long-standing research problem. In this work, we aim to reduce the sample complexity of existing policy gradient methods. We propose a novel policy gradient algorithm called SRVR-PG, which only requires episodes to find an -approxima…
Adaptive optimal transport priors improve few-shot learning robustness.
Optimal hedging strategies for exotic options using vanilla options.
We give a proof of the LMO conjecture which say that for any simply connectd simple Lie group , the LMO invariant of rational homology 3-spheres recovers the perturvative invariant . By Habiro-Le theorem, this implies that the LMO invariant is the universal quantum invariant of integral homology 3-spheres.
Unified approach to Merton's portfolio problem using Pontryagin's principles.
Power system studies require the topological structures of real-world power networks; however, such data is confidential due to important security concerns. Thus, power grid synthesis (PGS), i.e., creating realistic power grids that imitate actual power networks, has gained significant attention. In this letter, we cas…
Vanilla GANs are connected to Wasserstein distance for better understanding.
Recent studies have shown that proximal gradient (PG) method and accelerated gradient method (APG) with restarting can enjoy a linear convergence under a weaker condition than strong convexity, namely a quadratic growth condition (QGC). However, the faster convergence of restarting APG method relies on the potentially …