Improves exploration in reinforcement learning with diverse population.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
AIPS improves ranking policy evaluation by adapting to diverse user behavior.
New method recovers diverse policies from expert data using state-action pair weighting.
Quality-Diversity algorithms explore multiple high-performing solutions in a search space.
The paper improves QD policy ensembles using distribution ratio estimators.
This paper analyzes DRL strategies in finance, revealing unique trading patterns and performance differences.
We derive a class of macroscopic differential equations that describe collective adaptation, starting from a discrete-time stochastic microscopic model. The behavior of each agent is a dynamic balance between adaptation that locally achieves the best action and memory loss that leads to randomized behavior. We show tha…
A method for a single policy to solve various tasks across diverse agent morphologies.
Hydra distills ensemble models into a single model while preserving diversity and uncertainty.
this paper has been withdrawn
Interactive news recommendation has been launched and attracted much attention recently. In this scenario, user's behavior evolves from single click behavior to multiple behaviors including like, comment, share etc. However, most of the existing methods still use single click behavior as the unique criterion of judging…
Two hitherto disconnected threads of research, diverse exploration (DE) and maximum entropy RL have addressed a wide range of problems facing reinforcement learning algorithms via ostensibly distinct mechanisms. In this work, we identify a connection between these two approaches. First, a discriminator-based diversity …
Mobile phones can record individual's daily behavioral data as a time-series. In this paper, we present an effective time-series segmentation technique that extracts optimal time segments of individual's similar behavioral characteristics utilizing their mobile phone data. One of the determinants of an individual's beh…
This study investigates the potential effects of different Dynamic Message Signs (DMSs) on driver behavior using a full-scale high-fidelity driving simulator. Different DMSs are categorized by their content, structure, and type of messages. A random forest algorithm is used for three separate behavioral analyses; a rou…
This paper formulates the problem of building a context-aware predictive model based on user diverse behavioral activities with smartphones. In the area of machine learning and data science, a tree-like model as that of decision tree is considered as one of the most popular classification techniques, which can be used …
In recent years, reinforcement learning (RL) methods have been applied to model gameplay with great success, achieving super-human performance in various environments, such as Atari, Go, and Poker. However, those studies mostly focus on winning the game and have largely ignored the rich and complex human motivations, w…
Proposes method to discover diverse near-optimal policies in reinforcement learning.
Learning algorithms are enabling robots to solve increasingly challenging real-world tasks. These approaches often rely on demonstrations and reproduce the behavior shown. Unexpected changes in the environment may require using different behaviors to achieve the same effect, for instance to reach and grasp an object in…
Diverse projection ensembles improve distributional reinforcement learning.
New method uses SHapley Additive Explanations to identify anomaly detectors with complementary behaviors.
Most animals possess the ability to actuate a vast diversity of movements, ostensibly constrained only by morphology and physics. In practice, however, a frequent assumption in behavioral science is that most of an animal's activities can be described in terms of a small set of stereotyped motifs. Here we introduce a m…
A new method finds diverse near-optimal portfolios using quality-diversity.
Standard reinforcement learning methods aim to master one way of solving a task whereas there may exist multiple near-optimal policies. Being able to identify this collection of near-optimal policies can allow a domain expert to efficiently explore the space of reasonable solutions. Unfortunately, existing approaches t…
Synthesizes computational approaches to understand neural timescales.
DIVA generates diverse tasks for complex simulators, enabling adaptive agent training.
This paper improves model generalization by integrating diverse pretrained models.
DiCE uses diverse agents to explore and learn, avoiding local minima.
RE enhances DL by learning model behavior, enabling iterative self-improvement.
Model assesses credit risk using behavioral data from Experian and Bank of Italy.
This paper proposes a new method to connect language and physical actions in reinforcement learning.
Transformers learn to use induction heads or shortcuts based on data diversity.
RL enhances LLM planning but introduces spurious solutions and diversity collapse.
Study shows diverse data types improve SARS-COV-2 case surge predictions.
We study the problem of controllable generation of long-term sequential behaviors, where the goal is to calibrate to multiple behavior styles simultaneously. In contrast to the well-studied areas of controllable generation of images, text, and speech, there are two questions that pose significant challenges when genera…
The Hull-Strominger system for supersymmetric vacua of the heterotic string allows general unitary Hermitian connections with torsion and not just the Chern unitary connection. Solutions on unimodular Lie groups exploiting this flexibility were found by T. Fei and S.T. Yau. The Anomaly flow is a flow whose stationary p…
DreamerV3 learns diverse tasks with a single configuration.
SODA-RL learns diverse treatment options for hypotension from data.
A recommendation framework helps users choose healthcare interventions.
We study --both in theory and practice-- the use of momentum motions in classic iterative hard thresholding (IHT) methods. By simply modifying plain IHT, we investigate its convergence behavior on convex optimization criteria with non-convex constraints, under standard assumptions. In diverse scenaria, we observe that …
Valid certifies LLMs' domain adherence, bounding out-of-domain behavior.
GGP models multivariate time series with latent sub-sequences for diverse behaviors.
ActiLabel learns activity patterns across diverse sensor devices.
The study provides precise asymptotic theory for in-context learning by Transformers.
Efficient exploration remains a challenging research problem in reinforcement learning, especially when an environment contains large state spaces, deceptive local optima, or sparse rewards. To tackle this problem, we present a diversity-driven approach for exploration, which can be easily combined with both off- and o…
Current learning machines have successfully solved hard application problems, reaching high accuracy and displaying seemingly "intelligent" behavior. Here we apply recent techniques for explaining decisions of state-of-the-art learning machines and analyze various tasks from computer vision and arcade games. This showc…
Learning to rank is an important problem in machine learning and recommender systems. In a recommender system, a user is typically recommended a list of items. Since the user is unlikely to examine the entire recommended list, partial feedback arises naturally. At the same time, diverse recommendations are important be…
Inspired by the unsupervised learning or self-organization in the machine learning context, here we attempt to draw `learning curve' for the collective behavior of job-seeking `zero-intelligence' labors in successive job-hunting processes. Our labor market is supposed to be opened especially for university graduates in…
HealthSyn generates synthetic user behavior data for health interventions.