Trading off exploration and exploitation in an unknown environment is key to maximising expected return during learning. A Bayes-optimal policy, which does so optimally, conditions its actions not only on the environment state but on the agent's uncertainty about the environment. Computing a Bayes-optimal policy is how…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Sequence models quantify uncertainty over latent concepts.
Model based predictions of future trajectories of a dynamical system often suffer from inaccuracies, forcing model based control algorithms to re-plan often, thus being computationally expensive, suboptimal and not reliable. In this work, we propose a model agnostic method for estimating the uncertainty of a model?s pr…
Risk-averse model uncertainty framework for safe reinforcement learning.
Improved RL algorithm for robustness against parameter mismatches.
Bayesian segmentation and uncertainty estimation improve 3D model accuracy for factory planning.
Learning a policy using only observational data is challenging because the distribution of states it induces at execution time may differ from the distribution observed during training. We propose to train a policy by unrolling a learned model of the environment dynamics over multiple time steps while explicitly penali…
Model-free reinforcement learning based methods such as Proximal Policy Optimization, or Q-learning typically require thousands of interactions with the environment to approximate the optimum controller which may not always be feasible in robotics due to safety and time consumption. Model-based methods such as PILCO or…
Reinforcement learning agents are faced with two types of uncertainty. Epistemic uncertainty stems from limited data and is useful for exploration, whereas aleatoric uncertainty arises from stochastic environments and must be accounted for in risk-sensitive applications. We highlight the challenges involved in simultan…
Algorithm integrates uncertainty for lifelong learning in dynamic environments.
SAMPLR optimizes for ground truth in aleatoric parameters to avoid curriculum-induced covariate shift.
We can overcome uncertainty with uncertainty. Using randomness in our choices and in what we control, and hence in the decision making process, could potentially offset the uncertainty inherent in the environment and yield better outcomes. The example we develop in greater detail is the news-vendor inventory management…
Improves robust transfer learning with side information.
New algorithm tracks deep RL value functions with uncertainty.
LiveTradeBench evaluates LLMs in live trading environments.
Bayesian Federated Learning improves model reliability in dynamic environments.
A new RL model ensures safe learning in uncertain environments.
Safe learning in uncertain systems with state measurements and optimization.
BCPO optimizes offline RL policies by converting uncertainty into conservative bounds.
The paper tackles mean-variance analysis in Bayesian optimization under uncertainty.
A reinforcement learning framework combining value function and tree search planner for strategic and tactical decisions.
Autonomous lane changing is a critical feature for advanced autonomous driving systems, that involves several challenges such as uncertainty in other driver's behaviors and the trade-off between safety and agility. In this work, we develop a novel simulation environment that emulates these challenges and train a deep r…
Efficient exploration remains a challenging problem in reinforcement learning, especially for those tasks where rewards from environments are sparse. A commonly used approach for exploring such environments is to introduce some "intrinsic" reward. In this work, we focus on model uncertainty estimation as an intrinsic r…
Unsupervised representation learning has succeeded with excellent results in many applications. It is an especially powerful tool to learn a good representation of environments with partial or noisy observations. In partially observable domains it is important for the representation to encode a belief state, a sufficie…
A simple uncertainty measure improves deep bandit performance.
Paper proposes a framework for reliable off-policy evaluation in reinforcement learning.
New framework for robust uncertainty quantification in strategic settings.
We integrate information-theoretic concepts into the design and analysis of optimistic algorithms and Thompson sampling. By making a connection between information-theoretic quantities and confidence bounds, we obtain results that relate the per-period performance of the agent with its information gain about the enviro…
SARL uses predicted asset movements to improve financial portfolio management.
Paper introduces probabilistic digital twins for optimal decision making under uncertainty.
Enhances PlaNet for better planning in uncertain environments.
New RL algorithm tackles online robust MDPs with uncertainty.
Bayesian framework improves uncertainty estimates under covariate shifts.
Ultrasonic guided waves are commonly used to localize structural damage in infrastructures such as buildings, airplanes, bridges. Damage localization can be viewed as an inverse problem. Physical model based techniques are popular for guided wave based damage localization. The performance of these techniques depend on …
Before deploying autonomous agents in the real world, we need to be confident they will perform safely in novel situations. Ideally, we would expose agents to a very wide range of situations during training, allowing them to learn about every possible danger, but this is often impractical. This paper investigates safet…
This work tackles robust RL in multi-agent settings, improving sample efficiency.
UTE improves reinforcement learning by measuring action uncertainty, enhancing policy learning efficiency.
Traditional model-based RL relies on hand-specified or learned models of transition dynamics of the environment. These methods are sample efficient and facilitate learning in the real world but fail to generalize to subtle variations in the underlying dynamics, e.g., due to differences in mass, friction, or actuators a…
Paper tackles robust reinforcement learning with minimal data.
New algorithms reduce dynamic regret in non-stationary RL environments.
Animals need to devise strategies to maximize returns while interacting with their environment based on incoming noisy sensory observations. Task-relevant states, such as the agent's location within an environment or the presence of a predator, are often not directly observable but must be inferred using available sens…
In the edge computing paradigm, mobile devices offload the computational tasks to an edge server by routing the required data over the wireless network. The full potential of edge computing becomes realized only if a smart device selects the most appropriate server in terms of the latency and energy consumption, among …
We consider an agent's uncertainty about its environment and the problem of generalizing this uncertainty across observations. Specifically, we focus on the problem of exploration in non-tabular reinforcement learning. Drawing inspiration from the intrinsic motivation literature, we use density models to measure uncert…
For mobile robots to operate autonomously in general environments, perception is required in the form of a dense metric map. For this purpose, we present the stochastic triangular mesh (STM) mapping technique: a 2.5-D representation of the surface of the environment using a continuous mesh of triangular surface element…
Selective planning with imperfect models reduces harmful effects of model inadequacy.
A Robust Markov Decision Process (RMDP) is a sequential decision making model that accounts for uncertainty in the parameters of dynamic systems. This uncertainty introduces difficulties in learning an optimal policy, especially for environments with large state spaces. We propose two algorithms, RTD-DQN and Deep-RoK, …
This research tackles balancing exploration and exploitation in deep RL for partially observable systems.
New research finds uncertainty estimation techniques fail to reliably detect abnormal medical cases.