Meta-algorithm reduces exploration steps in changing CMPs.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We consider the exploration/exploitation problem in reinforcement learning. For exploitation, it is well known that the Bellman equation connects the value at any time-step to the expected value at subsequent time-steps. In this paper we consider a similar \textit{uncertainty} Bellman equation (UBE), which connects the…
In this paper, we introduce a system called GamePad that can be used to explore the application of machine learning methods to theorem proving in the Coq proof assistant. Interactive theorem provers such as Coq enable users to construct machine-checkable proofs in a step-by-step manner. Hence, they provide an opportuni…
Note on instabilities in super-time-stepping methods for Heston model.
Q-learning is a simple and powerful tool in solving dynamic problems where environments are unknown. It uses a balance of exploration and exploitation to find an optimal solution to the problem. In this paper, we propose using four basic emotions: joy, sadness, fear, and anger to influence a Qlearning agent. Simulation…
We study an exploration method for model-free RL that generalizes the counter-based exploration bonus methods and takes into account long term exploratory value of actions rather than a single step look-ahead. We propose a model-free RL method that modifies Delayed Q-learning and utilizes the long-term exploration bonu…
A 2-step nilpotent Lie algebra n is called nonsingular if ad(X): n --> [n,n] is onto for any X not in [n,n]. We explore nonsingular algebras in several directions, including the classification problem (isomorphism invariants), the existence of canonical inner products (nilsolitons) and their automorphism groups (maxima…
Despite the prevalence of collaborative filtering in recommendation systems, there has been little theoretical development on why and how well it works, especially in the "online" setting, where items are recommended to users over time. We address this theoretical gap by introducing a model for online recommendation sy…
ETGL-DDPG improves DDPG for sparse reward control with new exploration and replay techniques.
The paper explores meromorphic connections over Frobenius manifolds.
Under a nondegeneracy condition, we show that an equiregular sub-Riemannian manifold of step size admits a canonical, -rigid complement defined from the sub-Riemannian data that is preserved the by action of sub-Riemannian isometries. We explore how the existence of such a complement relates to results from the …
Enhanced Gaussian process models accelerate optimization and posterior approximation.
New algorithm learns optimal exploration parameters for contextual bandits.
We study the exploration problem in episodic MDPs with rich observations generated from a small number of latent states. Under certain identifiability assumptions, we demonstrate how to estimate a mapping from the observations to latent states inductively through a sequence of regression and clustering steps -- where p…
SPAQL improves RL by adaptively partitioning state-action space and learning a time-invariant policy.
New model predicts chemical reactions without human input.
This study explores complex structures on Lie algebras from graph perspectives.
Best arm identification (or, pure exploration) in multi-armed bandits is a fundamental problem in machine learning. In this paper we study the distributed version of this problem where we have multiple agents, and they want to learn the best arm collaboratively. We want to quantify the power of collaboration under limi…
This paper explores the problem of learning transforms for image compression via autoencoders. Usually, the rate-distortion performances of image compression are tuned by varying the quantization step size. In the case of autoen-coders, this in principle would require learning one transform per rate-distortion point at…
Randomized exploration in linear bandits achieves optimal regret bounds.
ANN with GA optimizes flexible disc design for lower mass and stress.
In the present paper we study the rigidity of 2-step Carnot groups, or equivalently, of graded 2-step nilpotent Lie algebras. We prove the alternative that depending on bi-dimensions of the algebra, the Lie algebra structure makes it either always of infinite type or generically rigid, and we specify the bi-dimensions …
Exploration is a difficult challenge in reinforcement learning and even recent state-of-the art curiosity-based methods rely on the simple epsilon-greedy strategy to generate novelty. We argue that pure random walks do not succeed to properly expand the exploration area in most environments and propose to replace singl…
MAME models a separate exploration policy for faster adaptation.
Variational inference lies at the core of many state-of-the-art algorithms. To improve the approximation of the posterior beyond parametric families, it was proposed to include MCMC steps into the variational lower bound. In this work we explore this idea using steps of the Hamiltonian Monte Carlo (HMC) algorithm, an e…
The paper explores dynamic ensembles for multi-step forecasting.
We take initial steps in studying PAC-MDP algorithms with limited adaptivity, that is, algorithms that change its exploration policy as infrequently as possible during regret minimization. This is motivated by the difficulty of running fully adaptive algorithms in real-world applications (such as medical domains), and …
BeBold improves exploration in sparse-reward tasks by regulating visitation counts.
RECODE uses clustering and embedding to track state visitation counts in RL.
In graph-based active learning, algorithms based on expected error minimization (EEM) have been popular and yield good empirical performance. The exact computation of EEM optimally balances exploration and exploitation. In practice, however, EEM-based algorithms employ various approximations due to the computational ha…
WITCHcraft improves PGD attacks with random step size, enhancing efficiency.
Greedy policy maximizes information in unknown linear systems.
Reinforcement learning (RL) has had many successes in both "deep" and "shallow" settings. In both cases, significant hyperparameter tuning is often required to achieve good performance. Furthermore, when nonlinear function approximation is used, non-stationarity in the state representation can lead to learning instabil…
Two new exploration methods for multi-agent systems improve team performance.
Improved exploration algorithm for unknown MDPs with reduced sample complexity.
Proposes r2SGLD for efficient constrained exploration in non-convex learning.
Adaptive importance sampling (AIS) uses past samples to update the \textit{sampling policy} at each stage . Each stage is formed with two steps : (i) to explore the space with points according to and (ii) to exploit the current amount of information to update the sampling policy. The very funda…
Paper tackles efficient exploration of unseen graph-structured environments.
New model-free algorithms learn representations for low-rank MDPs efficiently.
New algorithm identifies best policy in MDPs faster.
FLEX optimizes exploration for nonlinear systems with minimal data.
New method improves smoothness of robot learning.
Improved RL algorithm stabilizes unknown linear systems with polynomial regret.
The paper explores MAB strategies for very short horizons, introducing new methods and showing improved performance.
Although reinforcement learning has made great strides recently, a continuing limitation is that it requires an extremely high number of interactions with the environment. In this paper, we explore the effectiveness of reusing experience from the experience replay buffer in the Deep Q-Learning algorithm. We test the ef…
Stratify unifies and improves multi-step forecasting strategies.
New RL algorithm explains why deep learning works in stochastic environments.
The posteriors over neural network weights are high dimensional and multimodal. Each mode typically characterizes a meaningfully different representation of the data. We develop Cyclical Stochastic Gradient MCMC (SG-MCMC) to automatically explore such distributions. In particular, we propose a cyclical stepsize schedul…