NeuPL learns diverse policies in strategy games efficiently.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Framework improves policy generalizability under biased training data.
Estimates personalized policies robust to shifts in target populations.
Combines trial and observational data to improve policy evaluation.
GNMC reduces XCSF population size while preserving function approximation and policy accuracy.
Policy mirror ascent achieves Nash equilibrium in mean field games without a population generative model.
Investigates optimal pension policies in PAYG systems with forward utility and ageing population.
We develop asymptotically optimal policies for the multi armed bandit (MAB), problem, under a cost constraint. This model is applicable in situations where each sample (or activation) from a population (bandit) incurs a known bandit dependent cost. Successive samples from each population are iid random variables with u…
Unified q-learning for mean-field jump-diffusion models with unobservable population distribution.
Policy learning can be used to extract individualized treatment regimes from observational data in healthcare, civics, e-commerce, and beyond. One big hurdle to policy learning is a commonplace lack of overlap in the data for different actions, which can lead to unwieldy policy evaluation and poorly performing learned …
Continuous control tasks in reinforcement learning are important because they provide an important framework for learning in high-dimensional state spaces with deceptive rewards, where the agent can easily become trapped into suboptimal solutions. One way to avoid local optima is to use a population of agents to ensure…
In this paper, a new population-guided parallel learning scheme is proposed to enhance the performance of off-policy reinforcement learning (RL). In the proposed scheme, multiple identical learners with their own value-functions and policies share a common experience replay buffer, and search a good policy in collabora…
One shirt size cannot fit everybody, while we cannot make a unique shirt that fits perfectly for everyone because of resource limitation. This analogy is true for the policy making. Policy makers cannot establish a single policy to solve all problems for all regions because each region has its own unique issue. In the …
Optimizes COVID-19 testing policy using a Multi-Armed Bandit approach.
A key challenge in leveraging data augmentation for neural network training is choosing an effective augmentation policy from a large search space of candidate operations. Properly chosen augmentation policies can lead to significant generalization improvements; however, state-of-the-art approaches such as AutoAugment …
The paper improves QD policy ensembles using distribution ratio estimators.
Robust RL improves controller robustness to dynamics variations using adversarial populations.
New approach to off-policy evaluation connects causal graph to policy effects.
A new method simulates large, diverse populations of learning agents evolving in games.
Estimates Mozambique's population using remote sensing and microcensus data.
This paper combines and develops the models in Lastrapes (2002) and Mankiw & Weil (1989), which enables us to analyze the effects of interest rate and population growth shocks on housing price in one integrated framework. Based on this model, we carry out policy simulations to examine whether the housing (stock or flow…
We consider the \mnk{classical} problem of a controller activating (or sampling) sequentially from a finite number of populations, specified by unknown distributions. Over some time horizon, at each time , the controller wishes to select a population to sample, with the goal of sampling fro…
Algorithm optimizes lockdown policies balancing health and economy.
New framework for online control in evolving populations.
Framework generates precise synthetic populations for scalable modeling.
Consider the problem of a controller sampling sequentially from a finite number of populations, specified by random variables , and ; where denotes the outcome from population the time it is sampled. It is assumed that for each fixed , $\{…
Study bridges welfare maximization and CATE estimation in policy learning.
A government has to finance a risk for its population. It shares the charges among the population with a fixed scale based on economic criteria. Various organisms have to collect and to redistribute fairly the subsidies. Under these conditions, when the size of the organisms is varied, the distribution's laws of the cr…
Study models risks for low-carbon economy in Balkan countries, focusing on shadow economy and populism.
Model shows significant income inequality emerges from equal opportunities in a simple economy.
Social Security and other public policies can be viewed as a series of cash in and outflows that depend on parameters such as the age distribution of the population and the retirement age. Given forecasts of these parameters, policies can be designed to be financially stable, i.e., to terminate with a zero balance. If …
Bayesian networks learn sub-population differences from data.
PyCFRL helps ensure fair reinforcement learning policies from offline data.
The emergence of complex life on Earth is often attributed to the arms race that ensued from a huge number of organisms all competing for finite resources. We present an artificial intelligence research environment, inspired by the human game genre of MMORPGs (Massively Multiplayer Online Role-Playing Games, a.k.a. MMO…
LHIEM model predicts health, income, and employment over years.
Deep Reinforcement Learning (DRL) algorithms have been successfully applied to a range of challenging control tasks. However, these methods typically suffer from three core difficulties: temporal credit assignment with sparse rewards, lack of effective exploration, and brittle convergence properties that are extremely …
EPC curriculum improves MARL performance as agent population grows.
Consider the problem of finding a population or a probability distribution amongst many with the largest mean when these means are unknown but population samples can be simulated or otherwise generated. Typically, by selecting largest sample mean population, it can be shown that false selection probability decays at an…
We consider the problem of sequential sampling from a finite number of independent statistical populations to maximize the expected infinite horizon average outcome per period, under a constraint that the expected average sampling cost does not exceed an upper bound. The outcome distributions are not known. We construc…
Study uses machine learning to estimate effective policies in settings with hidden individual actions.
Consider the problem of sampling sequentially from a finite number of populations, specified by random variables , and ; where denotes the outcome from population the time it is sampled. It is assumed that for each fixed , $\{ X^i_k \}_{k …
Census data provide detailed information about population characteristics at a coarse resolution. Nevertheless, fine-grained, high-resolution mappings of population counts are increasingly needed to characterize population dynamics and to assess the consequences of climate shocks, natural disasters, investments in infr…
Study finds strict collection policies improve portfolio quality of microfinance banks.
Efficiently audits model fairness with continuous monitoring and flexible data collection.
Recent studies on fairness in automated decision making systems have both investigated the potential future impact of these decisions on the population at large, and emphasized that imposing ''typical'' fairness constraints such as demographic parity or equality of opportunity does not guarantee a benefit to disadvanta…
Lapse-supported life insurance exacerbates adverse selection risks.
Recent years have witnessed a tremendous improvement of deep reinforcement learning. However, a challenging problem is that an agent may suffer from inefficient exploration, particularly for on-policy methods. Previous exploration methods either rely on complex structure to estimate the novelty of states, or incur sens…
ABPS improves RL training efficiency by sharing policies and evolving hyper-params.