New approach handles stochastic and partially-observable environments using discrete autoencoders and Monte Carlo tree search.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Analyzes non-Markovian environments in stochastic approximation.
New algorithm for reinforcement learning in uncertain environments with unknown thresholds.
In this note, we extend an evolutionary stochastic portfolio optimization framework to include probabilistic constraints. Both the stochastic programming-based modeling environment as well as the evolutionary optimization environment are ideally suited for an integration of various types of probabilistic constraints. W…
New method uses hindsight to make exploration robust in stochastic environments.
New issue found in value-based reinforcement learning for stochastic environments.
Adaptive linear bandit algorithm with best-of-three-worlds regret bounds.
Develops model selection for bandits balancing adversarial and stochastic guarantees.
We discuss a variant of Thompson sampling for nonparametric reinforcement learning in a countable classes of general stochastic environments. These environments can be non-Markov, non-ergodic, and partially observable. We show that Thompson sampling learns the environment class in the sense that (1) asymptotically its …
New algorithms reduce regret in both stochastic and deterministic environments.
Surrogate models speed up RL training in dynamic systems.
Revisits VIC method to correct intrinsic reward bias in stochastic environments.
UDRL fails to converge in stochastic environments with episodic resets.
Develops a strategy to minimize loss in both stochastic and adversarial environments for linear contextual bandits.
We derive an algorithm that achieves the optimal (within constants) pseudo-regret in both adversarial and stochastic multi-armed bandits without prior knowledge of the regime and time horizon. The algorithm is based on online mirror descent (OMD) with Tsallis entropy regularization with power and reduced-varian…
This article extends, in a stochastic environment, the Yagil (1987) model which establishes, in a deterministic dividend discount model, a range for the exchange ratio in a stock-for-stock merger agreement. Here, we generalize Yagil's work letting both pre- and post-merger dividends grow randomly over time. If Yagil fo…
Autonomous robots need to interact with unknown, unstructured and changing environments, constantly facing novel challenges. Therefore, continuous online adaptation for lifelong-learning and the need of sample-efficient mechanisms to adapt to changes in the environment, the constraints, the tasks, or the robot itself a…
This paper considers online convex optimization (OCO) with stochastic constraints, which generalizes Zinkevich's OCO over a known simple fixed set by introducing multiple stochastic functional constraints that are i.i.d. generated at each round and are disclosed to the decision maker only after the decision is made. Th…
We provide bounds on control learning error in stochastic systems.
The paper analyzes trade execution strategies for large traders in a stochastic market environment.
We present analytical investigations of a multiplicative stochastic process that models a simple investor dynamics in a random environment. The dynamics of the investor's budget, , depends on the stochasticity of the return on investment, , for which different model assumptions are discussed. The fat-tail d…
Trade-R1 bridges verifiable rewards to stochastic financial markets via process-level reasoning verification.
Diffusion models mimic human actions in sequential tasks.
A method for self-supervised representation learning in partially observable environments.
Proposes a new resampling method for off-policy evaluation in stochastic control.
We develop the first general semi-bandit algorithm that simultaneously achieves regret for stochastic environments and regret for adversarial environments without knowledge of the regime or the number of rounds . The leading problem-dependent constants of our bounds are …
Stochastic Q-learning tackles large action spaces with reduced computation.
We designed a grid world task to study human planning and re-planning behavior in an unknown stochastic environment. In our grid world, participants were asked to travel from a random starting point to a random goal position while maximizing their reward. Because they were not familiar with the environment, they needed…
Study robust control for systems with continuous states using adversarial perturbations.
Bayesian Federated Learning improves model reliability in dynamic environments.
New RL algorithm explains why deep learning works in stochastic environments.
We establish convergence to an invariant measure as time tends to infinity, for a large class of (possibly non-Markovian) stochastic volatility models. Our arguments are based on a novel coupling idea for Markov chains which also extends to Markov chains in random environments in an efficient way.
Improved ExO method achieves near-optimal bounds in both stochastic and adversarial settings.
AEC Games model represents software MARL environments better than POSGs.
Meta-learning agents excel at rapidly learning new tasks from open-ended task distributions; yet, they forget what they learn about each task as soon as the next begins. When tasks reoccur - as they do in natural environments - metalearning agents must explore again instead of immediately exploiting previously discover…
In classical contagion models, default systems are Markovian conditionally on the observation of their stochastic environment, with interacting intensities. This necessitates that the environment evolves autonomously and is not influenced by the history of the default events. We extend the classical literature and allo…
Rough stochastic volatility models have attracted a lot of attentions recently, in particular for the linear option pricing problem. In this paper, starting with power utilities, we propose to use a martingale distortion representation of the optimal value function for the nonlinear asset allocation problem in a (non-M…
Algorithm optimizes system design and control for better rewards.
We revisit the fundamental problem of prediction with expert advice, in a setting where the environment is benign and generates losses stochastically, but the feedback observed by the learner is subject to a moderate adversarial corruption. We prove that a variant of the classical Multiplicative Weights algorithm with …
This paper studies the portfolio optimization problem when the investor's utility is general and the return and volatility of the risky asset are fast mean-reverting, which are important to capture the fast-time scale in the modeling of stock price volatility. Motivated by the heuristic derivation in [J.-P. Fouque, R. …
NetHack Learning Environment (NLE) tests RL algorithms, offering scalable, complex, and challenging gameplay.
Inverse reinforcement learning has proved its ability to explain state-action trajectories of expert agents by recovering their underlying reward functions in increasingly challenging environments. Recent advances in adversarial learning have allowed extending inverse RL to applications with non-stationary environment …
Reproducibility in reinforcement learning is challenging: uncontrolled stochasticity from many sources, such as the learning algorithm, the learned policy, and the environment itself have led researchers to report the performance of learned agents using aggregate metrics of performance over multiple random seeds for a …
We introduce a stochastic contextual bandit model where at each time step the environment chooses a distribution over a context set and samples the context from this distribution. The learner observes only the context distribution while the exact context realization remains hidden. This allows for a broad range of appl…
We present MDP Playground, a testbed for Reinforcement Learning (RL) agents with dimensions of hardness that can be controlled independently to challenge agents in different ways and obtain varying degrees of hardness in toy and complex RL environments. We consider and allow control over a wide variety of dimensions, i…
Investment strategies in occupational pension plans are optimized for non-tradable income risk.
LBQL improves Q-learning by using lookahead bounds for better performance.
Optimally explores dynamical systems with varying properties using context inference.