Agent learns policies from world models in simulated environments.
problem Training reinforcement learning agents in diverse environments.
method Generative neural network models environments, compact policies evolved from features.
result Achieves state-of-the-art performance across various environments.
ES-MAML uses Evolution Strategies for MAML, avoiding second derivative estimation.
problem Solving the MAML problem with efficient second derivative estimation.
method Applies Evolution Strategies to MAML, avoiding second derivative estimation.
result ES-MAML performs competitively and often better with fewer queries.
We present a new method of blackbox optimization via gradient approximation with the use of structured random orthogonal matrices, providing more accurate estimators than baselines and with provable theoretical guarantees. We show that this algorithm can be successfully applied to learn better quality compact policies …
Combines NES and PPO to enhance exploration in various environments.
problem Improving exploration in reinforcement learning environments.
method Parameter transfer and parameter space noise methods for combining NES and PPO.
result PPO benefits from both NES methods in discrete and continuous control tasks.
A novel meta-learning method using ES for efficient reinforcement learning.
problem Sample inefficiency in reinforcement learning.
method Evolution strategies (ES) for exploration in parameter space, deterministic policy gradients for adaptation.
result Demonstrates improved performance in high-dimensional control tasks compared to gradient-based methods.
ReSkill reconciles RL skill creation with policy optimization.
problem RL policies lack reusable strategies across tasks.
method Integrates skill creation into RL loop with three mechanisms.
result Consistently outperforms existing methods, especially on unseen tasks.
The paper optimizes pension policies with guarantees and sustainability constraints.
problem Designing optimal pension policies with guarantees and sustainability constraints.
method Dynamic utility model, stochastic domain, overlapping generations, time-consistent decision criterion.
result Optimal investment/pension policy computed for a general framework.
We present a broad agenda for meaningful banking regulation reform aiming the creation of evolutive competitive environment to maximize the effectiveness of international financial system through the introduction of fair competition process among the banks in free market capitalism. We assume that the international fin…
We study the optimal trading policies for a wind energy producer who aims to sell the future production in the open forward, spot, intraday and adjustment markets, and who has access to imperfect dynamically updated forecasts of the future production. We construct a stochastic model for the forecast evolution and deter…
New approach uses Wasserstein distances to score and optimize policy behaviors.
problem Comparing reinforcement learning policies and guiding policy optimization.
method Dual formulation of Wasserstein distances in latent behavioral space, learning score functions, smoothed WDs, stochastic gradient descent, on-policy algorithms.
result Demonstrated improved performance over existing methods in various environments.
IW-ES improves ES efficiency by using Importance Sampling.
problem Data inefficiency in Evolution Strategies.
method Importance Sampling to perform multiple updates per batch.
result IW-ES shows promising results in efficiency.
A new method simulates large, diverse populations of learning agents evolving in games.
problem Limited scalability and efficiency of Multi-Agent Reinforcement Learning.
method Parallelizable implementation of Policy Gradient and Opponent-Learning Awareness for evolutionary simulations.
result Simulated large, diverse populations of learning agents evolve under various strategies.
Paper proposes MCTSPO for better reinforcement learning policy optimization.
problem Local optima and saddle points in gradient-based methods and poor initialization in gradient-free methods.
method Monte-Carlo tree search combined with gradient-free optimization.
result Improved performance on reinforcement learning tasks with deceptive or sparse reward functions.
Current economic theories miss most of economic dynamics.
problem Accuracy of economic theories and policies depend on economic variables and processes.
method Identify and analyze overlooked economic variables and processes.
result Many economic variables and processes not accounted for in current theories.
This work uses a recommendation system to enhance exporting countries' fitness.
problem Predicting and optimizing the evolution of international trade networks.
method A recommendation system was used to identify overlooked products for countries.
result Countries can improve their national competitiveness by diversifying exported products.
ESPD improves learning efficiency in sparse reward reinforcement learning.
problem Sparse reward reinforcement learning challenges.
method Evolutionary Stochastic Policy Distillation (ESPD) based on drifted random walk insight.
result High learning efficiency demonstrated in MuJoCo robotics control suite experiments.
Optimizes antenna tilt for better QoS in cellular networks.
problem Hard to learn optimal antenna tilt policies in real networks due to risk and simulation gap.
method Uses off-policy Contextual Multi-Armed-Bandit (CMAB) techniques to learn from existing data.
result Trained policies show consistent improvements over existing logging policies.
ES optimization improved by structured control variates.
problem Improving accuracy of Evolution Strategies in RL.
method RL-specific variance reduction through structured control variates.
result Structured control variates outperform general variance reduction methods.
Deep Reinforcement Learning (DRL) algorithms have been successfully applied to a range of challenging control tasks. However, these methods typically suffer from three core difficulties: temporal credit assignment with sparse rewards, lack of effective exploration, and brittle convergence properties that are extremely …
A new ES method improves reinforcement learning speed and accuracy.
problem Slow convergence and local maxima in reinforcement learning.
method Directional Gaussian Smoothing Evolution Strategy (DGS-ES)
result DGS-ES accelerates RL training with high accuracy and nonlocal search direction.
This work develops agents to learn generalizable policies for dynamic network environments.
problem Real-world network topologies change due to attackers, defenders, or system failures, leading to failures in adaptive ACD systems.
method Developing agents to learn generalizable policies across dynamic network environments.
result Agents can learn robust policies for dynamic network topologies and diverse attackers.
Study uses artificial counterfactuals to show lockdowns reduced US case and death counts.
problem Impact of lockdowns on US case and death counts during the pandemic.
method Artificial counterfactual approach comparing states with and without lockdowns.
result Average two times more cases would have occurred without lockdowns.
New algorithms solve robust MDPs efficiently, significantly faster than existing methods.
problem Computing robust MDP solutions with uncertainty in transition probabilities is computationally expensive.
method Partial policy iteration and fast robust Bellman operator computation methods.
result The proposed methods are many orders of magnitude faster than state-of-the-art approaches.
MetaGenRL learns a general objective function from diverse agents.
problem Generalizing to new environments in reinforcement learning.
method MetaGenRL distills experiences from many agents into a low-complexity neural objective function.
result MetaGenRL can generalize to new environments and outperforms human-engineered algorithms.
The success of popular algorithms for deep reinforcement learning, such as policy-gradients and Q-learning, relies heavily on the availability of an informative reward signal at each timestep of the sequential decision-making process. When rewards are only sparsely available during an episode, or a rewarding feedback i…
Quantum-enhanced metrology aims to estimate an unknown parameter such that the precision scales better than the shot-noise bound. Single-shot adaptive quantum-enhanced metrology (AQEM) is a promising approach that uses feedback to tweak the quantum process according to previous measurement outcomes. Techniques and form…
NGE uses neural graphs to efficiently design robots.
problem Designing robots is hard due to combinatorial search space and evaluation costs.
method Formulated as graph search, NGE uses neural networks for policy parameterization and graph mutation with uncertainty.
result NGE significantly outperforms previous methods, discovering kinematically preferred structures.
We explore the use of Evolution Strategies (ES), a class of black box optimization algorithms, as an alternative to popular MDP-based RL techniques such as Q-learning and Policy Gradients. Experiments on MuJoCo and Atari show that ES is a viable solution strategy that scales extremely well with the number of CPUs avail…
Study shows activist board representation improves Japanese companies' performance.
problem Lack of innovation and improvement in Japanese companies.
method Examined two Japanese companies with activist board representation, analyzing performance metrics.
result Companies with activist board representation experienced significant improvements in stock returns and operational metrics.
Study optimal retirement time and consumption with habitual persistence.
problem Understanding retirement consumption patterns with habitual persistence.
method Established concise habitual evolution, used martingale and duality methods.
result Optimal consumption declines sharply at retirement but excess consumption increases.
We investigate the use of attentional neural network layers in order to learn a `behavior characterization' which can be used to drive novelty search and curiosity-based policies. The space is structured towards answering a particular distribution of questions, which are used in a supervised way to train the attentiona…
The paper examines how slightly biasing towards under-represented groups in sequential selection processes can lead to long-term fairness.
problem Designing fair sequential decision-making processes for long-term social fairness.
method Proposes Multi-agent Fair-Greedy policy to balance score maximization and fairness.
result Proves convergence to long-term fairness target set by agents when score distributions are identical.
Model predicts time evolution of supply chain networks under varying costs.
problem Regulating downstream relationships for sustainable SMEs.
method Time varying SCN model based on Lagrangian mechanics, incorporating EDES cost kernels.
result Model predicts bankruptcy and break-even states under different cost scenarios.
Robust policies improve ICU transfer outcomes by predicting patient deterioration.
problem Higher mortality rates for unplanned ICU transfers.
method Markov Decision Process model to predict patient severity and optimize transfer policies.
result Robust policies are more aggressive in transferring patients than nominal policies, improving overall patient care.
Our knowledge about the evolution of guarantee network in downturn period is limited due to the lack of comprehensive data of the whole credit system. Here we analyze the dynamic Chinese guarantee network constructed from a comprehensive bank loan dataset that accounts for nearly 80% total loans in China, during 01/200…
CPS methods improve sample efficiency in robotics.
problem Improving sample efficiency in reinforcement learning for robotics.
method Empirical evaluation of C-CMA-ES with active covariance matrix adaptation and comparison-based surrogate model.
result Improvements in sample efficiency with C-CMA-ES extensions.
Study best arm identification in restless Markov multi-armed bandits with state-dependent transitions.
problem Identify the best arm in a multi-armed bandit with time-varying states.
method Propose a sequential policy to select arms without knowing their exact TPMs.
result Upper and lower bounds on expected time to find the best arm match in a special case.
Novel RL-based NPG improves multi-objective NAS efficiency and performance.
problem Discovering optimal neural architectures with multiple conflicting objectives.
method Non-stationary policy gradient with adaptive reward functions and shared model.
result Framework efficiently approximates full Pareto front and achieves superior performance.
This paper presents a general theory that aims at explaining timescales observed empirically in technology transitions and predicting those of future transitions. This framework is used further to derive a theory for exploring the dynamics that underlie the complex phenomenon of irreversible and path dependent price or…
Prototype controls stochastic drug resistance in cells.
problem Emergence of drug-resistant cells from random mutations.
method Deep reinforcement learning for adaptive drug dosing.
result 100% success rate in suppressing cell proliferation.
Study uses exchangeable GPs for staggered-adoption policy evaluation in panel data.
problem Evaluating the impact of staggered treatments in panel data settings.
method Exchangeable multi-task Gaussian processes (GPs) with flexible kernels.
result Flexible tool for policy evaluation in panel data settings.
Optimizes consumption under regime-switching economic states with risk-sensitive preferences.
problem Optimizing consumption in an economy with uncertain states and random shocks.
method Risk-sensitive optimization of consumption-utility with a Markov chain model of economic states and i.i.d. random shocks.
result Existence of unique optimal policy and value function in stationary policies.
Modeling option market making with hedging-induced price impact.
problem Tackles the challenge of market making in options markets with price impact.
method Models option order flow using Cox processes and studies the dynamics of inventory and price under hedging-induced impact.
result Establishes the well-posedness of the mixed control problem involving quoting and hedging.
Robot learns agile leg movements from simulations to real robots.
problem Training legged robots for dynamic maneuvers is challenging and expensive.
method Simulation-based reinforcement learning for real legged systems.
result ANYmal robot can follow high-level commands, run faster, and recover from falls.
Study minimax rates for online learning with time-varying dynamics.
problem Online learning with time-varying state and cost dynamics.
method Non-constructive upper and lower bounds, complexity and stability terms.
result Characterization of minimax rates and necessary conditions for learnability.
Analyzes COVID-19 data to predict mortality, forecast spread, and optimize resource allocation.
problem Challenges in patient triage, treatment, and care management during the pandemic.
method Integrated four-step approach combining descriptive, predictive, and prescriptive analytics.
result Optimized resource allocation and informed policy decisions.
Unified coverage analysis for linear off-policy evaluation in reinforcement learning.
problem Lack of a unified understanding of coverage parameters in linear off-policy evaluation.
method Developed a novel finite-sample analysis for LSTDQ algorithm, introducing feature-dynamics coverage.
result Unified understanding of coverage parameters in linear off-policy evaluation.
We have modeled the employment/population ratio in the largest developed countries. Our results show that the evolution of the employment rate since 1970 can be predicted with a high accuracy by a linear dependence on the logarithm of real GDP per capita. All empirical relationships estimated in this study need a struc…