The abstract explores connections between reinforcement learning, scaling, and diffusion.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Machine learning models are often used at test-time subject to constraints and trade-offs not present at training-time. For example, a computer vision model operating on an embedded device may need to perform real-time inference, or a translation model operating on a cell phone may wish to bound its average compute tim…
Improves reinforcement learning extrapolation in Gridworlds.
Unified framework for certifying LLM reliability without extra supervision.
New method uses MCTS only at test time for faster game learning.
For an autonomous agent to fulfill a wide range of user-specified goals at test time, it must be able to learn broadly applicable and general-purpose skill repertoires. Furthermore, to provide the requisite level of generality, these skills must handle raw sensory input such as images. In this paper, we propose an algo…
Transfer and adaptation to new unknown environmental dynamics is a key challenge for reinforcement learning (RL). An even greater challenge is performing near-optimally in a single attempt at test time, possibly without access to dense rewards, which is not addressed by current methods that require multiple experience …
In multi-task reinforcement learning there are two main challenges: at training time, the ability to learn different policies with a single model; at test time, inferring which of those policies applying without an external signal. In the case of continual reinforcement learning a third challenge arises: learning tasks…
Deep Sets improve reinforcement learning agent's object-centered navigation and generalization.
Bayesian meta-reinforcement learning improves over point estimates with Laplace approximation.
Max entropy exploration guides reinforcement learning agents to pursue achievable goals.
Adversarial attacks can manipulate deep trading policies, compromising their performance.
New method improves convergence of RL meta-learning.
Algorithm learns new tasks efficiently from past experience.
Reinforcement learning after next-token prediction aids in learning from diverse sequence lengths.
New method improves meta-reinforcement learning efficiency.
Simple policy search outperforms advanced learnable test-time augmentation techniques.
This work uses reinforcement learning to optimize task scheduling and execution in a dynamic multi-agent warehouse environment.
We propose CAVIA for meta-learning, a simple extension to MAML that is less prone to meta-overfitting, easier to parallelise, and more interpretable. CAVIA partitions the model parameters into two parts: context parameters that serve as additional input to the model and are adapted on individual tasks, and shared param…
A probabilistic framework for online test-time adaptation
RAD enhances RL algorithms with data augmentations.
New method provides scalable safety guarantees for RL agents.
AdMRL improves meta-reinforcement learning by minimizing worst-case sub-optimality gap.
The aim of multi-task reinforcement learning is two-fold: (1) efficiently learn by training against multiple tasks and (2) quickly adapt, using limited samples, to a variety of new tasks. In this work, the tasks correspond to reward functions for environments with the same (or similar) dynamical models. We propose to l…
Integrates skills and world models for efficient task solving and transfer.
Deep neural network (DNN) based approaches hold significant potential for reinforcement learning (RL) and have already shown remarkable gains over state-of-art methods in a number of applications. The effectiveness of DNN methods can be attributed to leveraging the abundance of supervised data to learn value functions,…
Although reinforcement learning methods can achieve impressive results in simulation, the real world presents two major challenges: generating samples is exceedingly expensive, and unexpected perturbations or unseen situations cause proficient but specialized policies to fail at test time. Given that it is impractical …
Exposure bias refers to the train-test discrepancy that seemingly arises when an autoregressive generative model uses only ground-truth contexts at training time but generated ones at test time. We separate the contributions of the model and the learning framework to clarify the debate on consequences and review propos…
Empirical study shows consistent meta-RL algorithms adapt to OOD tasks.
We compare the model-free reinforcement learning with the model-based approaches through the lens of the expressive power of neural networks for policies, -functions, and dynamics. We show, theoretically and empirically, that even for one-dimensional continuous state space, there are many MDPs whose optimal -func…
Memory is an important aspect of intelligence and plays a role in many deep reinforcement learning models. However, little progress has been made in understanding when specific memory systems help more than others and how well they generalize. The field also has yet to see a prevalent consistent and rigorous approach f…
Meta-SAGE improves deep RL scalability for CO tasks by adapting pre-trained models to larger-scale problems.
SAGE enhances reinforcement learning by injecting hints to prevent model stagnation.
This paper improves test-time adaptation for distribution shifts using confidence maximization and input transformation.
Implicit models can match or exceed explicit models with more test-time compute.
Many real-world problems can be reduced to combinatorial optimization on a graph, where the subset or ordering of vertices that maximize some objective function must be found. With such tasks often NP-hard and analytically intractable, reinforcement learning (RL) has shown promise as a framework with which efficient he…
We show that several popular few-shot learning benchmarks can be solved with varying degrees of success without using support set Labels at Test-time (LT). To this end, we introduce a new baseline called Centroid Networks, a modification of Prototypical Networks in which the support set labels are hidden from the metho…
New approach makes deep reinforcement learning robust without assuming adversary knowledge.
Paper improves reinforcement learning in multi-scene tasks.
Competition aims to develop sample-efficient reinforcement learning methods.
This work explores test-time scaling strategies for LLMs, improving sample efficiency and expressiveness.
We tackle the Multi-task Batch Reinforcement Learning problem. Given multiple datasets collected from different tasks, we train a multi-task policy to perform well in unseen tasks sampled from the same distribution. The task identities of the unseen tasks are not provided. To perform well, the policy must infer the tas…
We propose the use of Bayesian networks, which provide both a mean value and an uncertainty estimate as output, to enhance the safety of learned control policies under circumstances in which a test-time input differs significantly from the training set. Our algorithm combines reinforcement learning and end-to-end imita…
Ordinary stochastic neural networks mostly rely on the expected values of their weights to make predictions, whereas the induced noise is mostly used to capture the uncertainty, prevent overfitting and slightly boost the performance through test-time averaging. In this paper, we introduce variance layers, a different k…
This paper addresses detection of a reverse engineering (RE) attack targeting a deep neural network (DNN) image classifier; by querying, RE's aim is to discover the classifier's decision rule. RE can enable test-time evasion attacks, which require knowledge of the classifier. Recently, we proposed a quite effective app…
Adaptive compute allocation improves model performance by prioritizing harder queries.
A new approach for test-time adaptation detects and reacts to distribution shifts.
Machine learning classifiers are known to be vulnerable to inputs maliciously constructed by adversaries to force misclassification. Such adversarial examples have been extensively studied in the context of computer vision applications. In this work, we show adversarial attacks are also effective when targeting neural …