New approach handles stochastic and partially-observable environments using discrete autoencoders and Monte Carlo tree search.
problem Challenges in planning for stochastic and partially-observable environments.
method Uses discrete autoencoders and a stochastic variant of Monte Carlo tree search.
result Significantly outperforms MuZero on stochastic chess and scales to DeepMind Lab.
Thompson Sampling learns unknown stochastic environments efficiently.
problem Learning unknown stochastic environments efficiently.
method Thompson Sampling applied to nonparametric reinforcement learning in general environments.
result Thompson Sampling asymptotically converges to optimal value and has sublinear regret.
Analyzes non-Markovian environments in stochastic approximation.
problem Understanding learning mechanisms in non-ergodic, non-Markovian settings.
method Analytic framework for transformer learning and continual learning.
result Proposes a new approach to transformer and continual learning.
New algorithm for reinforcement learning in uncertain environments with unknown thresholds.
problem Safety in reinforcement learning in unknown and uncertain environments.
method Growing-Window estimator sampling and Stochastic Pessimistic-Optimistic Thresholding (SPOT) algorithm.
result Achieves sublinear regret and constraint violation of i l d e O ( T ) ilde{\mathcal{O}}(\sqrt{T}) i l d e O ( T ) . Optimal reinsurance and investment strategies in a stochastic environment modelled by diffusion.
problem Maximizing terminal expected utility in an insurance risk model with reinsurance and financial investment.
method Diffusion approximation, SAHARA utility functions, explicit results for proportional and excess-of-loss reinsurance.
result Explicit results for optimal reinsurance and investment strategies.
Study shows how certain stochastic models reach a steady state over time.
problem Understanding long-term behavior of stochastic volatility models.
method Novel coupling technique for Markov chains, applicable to random environments.
result Convergence to an invariant measure for multidimensional fractional models.
Study on human planning and re-planning in unknown stochastic environments.
problem Understanding how humans adjust plans in unfamiliar environments.
method Grid world task, 12 different models, model-based reinforcement learning approach.
result Model-based reinforcement learning approach best explains human re-planning behavior.
Optimal semi-bandit algorithm for both stochastic and adversarial environments.
problem Optimal semi-bandit algorithm for both stochastic and adversarial environments.
method Developed a general semi-bandit algorithm that achieves O ( log T ) \mathcal{O}(\log T) O ( log T ) regret for stochastic and O ( T ) \mathcal{O}(\sqrt{T}) O ( T ) regret for adversarial environments without regime or T T T knowledge. result First algorithm to achieve optimal O ( log T ) \mathcal{O}(\log T) O ( log T ) and O ( T ) \mathcal{O}(\sqrt{T}) O ( T ) regret simultaneously for stochastic and adversarial environments. In this note, we extend an evolutionary stochastic portfolio optimization framework to include probabilistic constraints. Both the stochastic programming-based modeling environment as well as the evolutionary optimization environment are ideally suited for an integration of various types of probabilistic constraints. W…
Investigates Merton's portfolio problem in a rough stochastic environment with Volterra Heston model.
problem Optimizing investment strategies in a non-Markovian, non-semimartingale stochastic environment.
method Solves the portfolio optimization problem using the martingale optimality principle and auxiliary random process.
result Derives semi-closed form solutions for optimal strategies under power and exponential utilities.
New method uses hindsight to make exploration robust in stochastic environments.
problem Exploration in sparse-reward or reward-free environments, especially in stochastic settings.
method Learn representations of the future that capture unpredictable aspects, using them to predict and reward only the predictable parts of the world.
result Improves exploration in Atari games and Montezuma's Revenge, robust to stochasticity.
New issue found in value-based reinforcement learning for stochastic environments.
problem Value-based reinforcement learning struggles with stochastic state transitions.
method Demonstrated using a multiobjective Markov Decision Process (MOMDP).
result Approaches may converge to Pareto-dominated solutions instead of optimal ones.
Adaptive linear bandit algorithm with best-of-three-worlds regret bounds.
problem Adaptive to adversarial and stochastic environments with varying sub-optimality gaps and corruption.
method Combines SCRiBLe algorithm with scaled-up sampling and optimistic online learning.
result Achieves best-of-three-worlds regret bounds of O ( T log T ) O(\sqrt{T \log T}) O ( T log T ) for adversarial and O ( log T Δ min + C log T Δ min ) O(\frac{\log T}{Δ_{\min}} + \sqrt{\frac{C \log T}{Δ_{\min}}}) O ( Δ m i n l o g T + Δ m i n C l o g T ) for stochastic environments. Develops model selection for bandits balancing adversarial and stochastic guarantees.
problem Model selection in bandit scenarios with simultaneous adversarial and stochastic high-probability regret.
method Nested policy classes, balanced candidate regret bounds, mis-specification tests.
result Best of both world guarantees in linear bandits with simultaneous adversarial and stochastic environments.
A new bandit model with stochastic context distributions and UCB algorithm.
problem Learning optimal actions in a stochastic environment with hidden contexts.
method Stochastic contextual bandit model and UCB algorithm.
result Order-optimal high-probability bound on cumulative regret for linear and kernelized reward functions.
New algorithms reduce regret in both stochastic and deterministic environments.
problem Designing algorithms that perform well in both types of MDPs.
method Proposed new environment norms and algorithms with variance-dependent regret bounds.
result First algorithm with simultaneously optimal bounds for both stochastic and deterministic MDPs.
Paper proposes optimal portfolio strategy under rough stochastic volatility models.
problem Optimal asset allocation in a fractional stochastic environment with rough volatility.
method Martingale distortion representation and fixed zeroth order trading strategy.
result Asymptotic optimality of a fixed strategy for general utility functions.
Surrogate models speed up RL training in dynamic systems.
problem High computational cost of high-fidelity simulations.
method Developed and tested surrogate models for RL training.
result Surrogate models can significantly accelerate RL training.
Revisits VIC method to correct intrinsic reward bias in stochastic environments.
problem Intrinsic reward bias in VIC leading to suboptimal solutions.
method Proposes two methods based on transitional probability model and Gaussian mixture model to correct bias.
result Achieves maximal empowerment through corrected intrinsic reward.
UDRL fails to converge in stochastic environments with episodic resets.
problem UDRL's convergence in stochastic environments with resets is questioned.
method UDRL is a supervised learning approach that does not use value functions.
result UDRL diverges in a simple stochastic environment with resets.
Develops a strategy to minimize loss in both stochastic and adversarial environments for linear contextual bandits.
problem Linear contextual bandits with adversarial corruption.
method Proposes a novel strategy called Best-of-Both-Worlds (BoBW) RealFTRL, extending RealLinExp3 and FTRL.
result Regret upper bound of $O\left(\min\left\{\frac{(\log(T))^3}{Δ_{*}} + \sqrt{\frac{C(\log(T))^3}{Δ_{*}}},\ \ \sqrt{T}(\log(T))^2
ight\}
ight)$ , showing effectiveness in both stochastic and adversarial environments.
Optimal algorithm for both stochastic and adversarial bandits without prior knowledge.
problem Optimal algorithm for both stochastic and adversarial bandits.
method Online mirror descent with Tsallis entropy regularization and reduced-variance loss estimators.
result Achieves optimal pseudo-regret in both adversarial and stochastic bandits.
Paper proposes using expectation models for planning in stochastic environments.
problem Intractability of learning distribution and sample models in large state and action spaces.
method Proposes using approximate expectation models for MBRL, analyzes linear and non-linear parametrizations, and presents a policy evaluation algorithm.
result Planning with an expectation model is equivalent to planning with a distribution model under certain conditions.
Extends Yagil's model to include stochastic dividends in stock-for-stock mergers.
problem Determining exchange ratios in stock-for-stock mergers with uncertain dividend growth.
method Generalizes Yagil's deterministic model to a stochastic environment, considering both expected values and variance of dividends.
result Identifies a more complex bargaining region for exchange ratios, influenced by the mean and standard deviation of dividends' growth rate.
Paper analyzes learning dynamics in quasi-periodic environments, showing consistent solutions.
problem Challenges in stochastic gradient learning for complex environments.
method Uses energy balance equations derived from Caldirola-Kanai Hamiltonian to model learning.
result In quasi-periodic environments, learning yields consistent solutions for similar patterns.
Optimizes portfolios in fast mean-reverting markets, achieving asymptotic efficiency.
problem Optimizing portfolios in markets with fast mean-reverting returns and volatility.
method Proposes a zeroth order strategy and uses singular perturbation method for asymptotic optimality.
result Shows asymptotic optimality of the proposed strategy under specific assumptions.
Paper tackles online convex optimization with stochastic constraints.
problem Online convex optimization with stochastic constraints.
method Proposes a new algorithm achieving O ( T ) O(\sqrt{T}) O ( T ) expected regret and constraint violations and O ( T log ( T ) ) O(\sqrt{T}\log(T)) O ( T log ( T )) high probability regret and constraint violations. result Achieves optimal regret and constraint violation bounds.
Robots learn quickly from few interactions using mental replay and intrinsic motivation.
problem Continuous online adaptation for robots in changing environments.
method Bio-inspired stochastic recurrent neural network with learning signals and mental replay.
result Robots can adapt to novel environments in seconds from few interactions.
We provide bounds on control learning error in stochastic systems.
problem Learning optimal controls in stochastic environments with uncontrolled parts.
method Dynamic programming and mean-field interpretation of neural networks.
result Non-asymptotic bounds on generalization error for stable overparametrised settings.
The paper analyzes trade execution strategies for large traders in a stochastic market environment.
problem Analyzing trade execution strategies in a stochastic market with price impact.
method Formulated a Markov game model and used backward induction method of dynamic programming.
result Explicit closed-form execution strategy at Markov perfect equilibrium.
We present analytical investigations of a multiplicative stochastic process that models a simple investor dynamics in a random environment. The dynamics of the investor's budget, x ( t ) x(t) x ( t ) , depends on the stochasticity of the return on investment, r ( t ) r(t) r ( t ) , for which different model assumptions are discussed. The fat-tail d…
Trade-R1 bridges verifiable rewards to stochastic financial markets via process-level reasoning verification.
problem Extending RL to financial markets where rewards are verifiable but noisy.
method A verification method that transforms reasoning over financial documents into a structured RAG task, using a triangular consistency metric.
result DSR achieves superior cross-market generalization while maintaining reasoning consistency.
Diffusion models mimic human actions in sequential tasks.
problem Cloning human behavior in dynamic environments is challenging.
method Adapting diffusion models to handle stochastic, multimodal, and correlated actions.
result Diffusion models closely replicate human behavior in robotic and gaming tasks.
A method for self-supervised representation learning in partially observable environments.
problem Sparse rewards and stochasticity in partially observable environments.
method World model in latent space for estimating missing information.
result Significant improvement in exploration compared to prior work.
Proposes a new resampling method for off-policy evaluation in stochastic control.
problem Estimating policy performance from data generated under a different policy.
method K-nearest neighbor resampling procedure for off-policy evaluation.
result Statistical consistency results for the proposed method under weak conditions.
The paper analyzes optimal portfolio allocation under a fast mean-reverting fractional stochastic environment.
problem Optimal portfolio allocation under a fractional stochastic environment with long-range dependence.
method Analyzes the nonlinear optimal portfolio allocation problem using a stationary fractional Ornstein-Uhlenbeck process with fast mean-reverting.
result Establishes asymptotic optimality of zeroth order trading strategies and general utility functions within specific families of admissible strategies.
This paper optimizes portfolios in a fast-reverting volatility environment.
problem Optimizing portfolios in a fast-reverting volatility environment.
method Fractional Brownian motions with Hurst index H, modeling fast or slow regimes with small parameters.
result Only one deterministic term of order √ε appears in the first order correction for the fast-varying rough environment.
This paper proposes a self-supervised exploration method using disagreement of dynamics models.
problem Efficient exploration in stochastic environments with real robots.
method Train an ensemble of dynamics models and incentivize exploration to maximize disagreement.
result Sample-efficient exploration achieved without external rewards or reinforcement learning.
Stochastic Q-learning tackles large action spaces with reduced computation.
problem Effective decision-making in complex environments with large discrete action spaces.
method Stochastic value-based RL approaches that consider a sublinear number of actions in each iteration.
result Stochastic Q-learning achieves near-optimal returns with significantly reduced computation time.
Deep RL agents vary significantly in Atari environments.
problem Challenges in reproducibility due to stochasticity.
method Experiments with OpenAI Baselines agents.
result Variability in agent performance is significant and underreported.
Study robust control for systems with continuous states using adversarial perturbations.
problem Fragile policies in Markov control models under internal or external perturbations.
method Distributionally robust stochastic control with adaptive adversarial perturbations.
result Optimal robust policies for continuous state systems with uniform learning guarantees.
New algorithm learns good interventions faster using causal models.
problem Improving online learning of good interventions in stochastic environments.
method Combines multi-arm bandits and causal inference for novel bandit feedback.
result Proves a bound on simple regret strictly better than non-causal algorithms.
New RL algorithm explains why deep learning works in stochastic environments.
problem Why deep RL algorithms perform well in practice despite using random exploration.
method Introducing SQIRL, an iterative RL algorithm that separates exploration and learning.
result Effective horizon explains why deep RL works in stochastic environments.
Bayesian Federated Learning improves model reliability in dynamic environments.
problem Uncertainty quantification and robust adaptation in distributed learning.
method Proposes a continual BFL framework using SGLD for sequential updates and continual learning challenges.
result Continual Bayesian updates preserve knowledge and adapt to evolving data.
Improved algorithm reduces regret in corrupted expert advice setting.
problem Prediction with expert advice in the presence of adversarial corruption.
method Multiplicative Weights algorithm with decreasing step sizes.
result Achieves constant regret and optimal performance in various environments.
Improved ExO method achieves near-optimal bounds in both stochastic and adversarial settings.
problem Finding optimal exploration strategies in online decision-making with limited feedback.
method Exploration by Optimization with hybrid regularizers for locally observable games.
result Achieved nearly optimal bounds of O ( ∑ a e q a ∗ k 2 m 2 log T / Δ a ) O(\sum_{a
eq a^*} k^2 m^2 \log T / Δ_a) O ( ∑ a e q a ∗ k 2 m 2 log T / Δ a ) in stochastic and adversarial environments. AEC Games model represents software MARL environments better than POSGs.
problem POSGs are conceptually unsuitable for software MARL environments.
method Introduced AEC Games model as an equivalent to POSGs.
result AEC Games model is more representative of software MARL environments.
Algorithm optimizes system design and control for better rewards.
problem Optimizing system design and control for maximum rewards.
method Deep reinforcement learning combining policy gradient and model-based optimization.
result DEPS algorithm outperforms state-of-the-art methods in various environments.