New algorithm bounds regret in mediator feedback bandit problems.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New complexity measure helps in agnostic reinforcement learning with or without access to MDP dynamics.
Study optimal treatment assignment policies under strategic agent responses.
We consider the generic approach of using an experience memory to help exploration by adapting a restart distribution. That is, given the capacity to reset the state with those corresponding to the agent's past observations, we help exploration by promoting faster state-space coverage via restarting the agent from a mo…
New insights into experience replay in RL algorithms.
Proposes PIC and POIC for measuring task difficulty in RL.
Study online learning with delays and capacity constraints, achieving optimal regret bounds.
Double descent found in DRL, improving generalization with model capacity.
Graph neural networks optimize radio resource management policies for wireless networks.
Deep imagination optimizes decision-making in large trees with limited resources.
Optimizes COVID-19 testing policy using a Multi-Armed Bandit approach.
Paper tackles inventory management with deep learning, improving performance and adherence to constraints.
Feed in tariff (FiT) is one of the most efficient ways that many governments throughout the world use to stimulate investment in renewable energies (REs) technology. For governments, financial management of the policy is very challenging as that it needs a considerable amount of budget to support RE producers during th…
New complete panel dataset for LMICs helps analyze innovation and development.
We propose a new method to study the internal memory used by reinforcement learning policies. We estimate the amount of relevant past information by estimating mutual information between behavior histories and the current action of an agent. We perform this estimation in the passive setting, that is, we do not interven…
We consider the dynamic assortment optimization problem under the multinomial logit model (MNL) with unknown utility parameters. The main question investigated in this paper is model mis-specification under the -contamination model, which is a fundamental model in robust statistics and machine learning. In…
Estimating individual treatment effects from data of randomized experiments is a critical task in causal inference. The Stable Unit Treatment Value Assumption (SUTVA) is usually made in causal inference. However, interference can introduce bias when the assigned treatment on one unit affects the potential outcomes of t…
Survival analysis models predict economic convergence across Americas.
Entropy regularization is an important idea in reinforcement learning, with great success in recent algorithms like Soft Q Network (SQN) and Soft Actor-Critic (SAC1). In this work, we extend this idea into the on-policy realm. We propose the soft policy gradient theorem (SPGT) for on-policy maximum entropy reinforcemen…
FACMAC combines deep policy gradients with factored critic for multi-agent reinforcement learning.
Successful implementation of California's Renewable Portfolio Standard (RPS) mandating 33 percent renewable energy generation by 2020 requires inclusion of a robust strategy to mitigate increased risk of energy deficits (blackouts) due to short time-scale (sub 1 hour) intermittencies in renewable energy sources. Of the…
Develops neural network framework for risk-reward optimization problems.
Deep Sets improve reinforcement learning agent's object-centered navigation and generalization.
Solves POMDPs with recurrent neural networks and natural policy gradient.
Policy distillation in deep reinforcement learning provides an effective way to transfer control policies from a larger network to a smaller untrained network without a significant degradation in performance. However, policy distillation is underexplored in deep reinforcement learning, and existing approaches are compu…
Study sharp decay of capacity for subharmonic functions on compact Hermitian manifolds.
Sharp upper bounds derived for capacities in hyperbolic and Euclidean spaces.
A new framework uses deep reinforcement learning to improve aircraft separation in busy airspace.
We introduce a methodology for efficiently computing a lower bound to empowerment, allowing it to be used as an unsupervised cost function for policy learning in real-time control. Empowerment, being the channel capacity between actions and states, maximises the influence of an agent on its near future. It has been sho…
In this paper, we investigate the common scenario where every candidate item for recommendation is characterized by a maximum capacity, i.e., number of seats in a Point-of-Interest (POI) or size of an item's inventory. Despite the prevalence of the task of recommending items under capacity constraints in a variety of s…
Study semicontinuity of capacity in non-smooth spaces using intrinsic flat convergence.
Reinforcement learning has attracted great attention recently, especially policy gradient algorithms, which have been demonstrated on challenging decision making and control tasks. In this paper, we propose an active multi-step TD algorithm with adaptive stepsizes to learn actor and critic. Specifically, our model cons…
A framework for reinforcement learning tackles CVRP with competitive results.
The paper develops methods to estimate optimal treatment sequences under policy constraints.
Overparameterized models generalize well in offline contextual bandits, but policy-based algorithms struggle.
The electric capacity of a conductor in the 3-dimensional Euclidean space is defined as a ratio of a given positive charge on the conductor to the value of potential on the surface. This definition of the capacity is independent of the given charge. The capacity of a set as a mathematical notion was defined firs…
RichID learns optimal control policies from nonlinear observations.
Agents trained with deep reinforcement learning algorithms are capable of performing highly complex tasks including locomotion in continuous environments. We investigate transferring the learning acquired in one task to a set of previously unseen tasks. Generalization and overfitting in deep reinforcement learning are …
This note develops certain sharp inequalities relating the fractional Sobolev capacity of a set to its standard volume and fractional perimeter.
New algorithm optimizes reward while ensuring safety in complex decision-making problems.
An online learning framework optimizes pricing and capacity in service systems.
Study capacity constraints in continual learning with a simple model.
The real estate is a pillar industry of China's national economy. Due to changes in policy and market conditions, the real estate companies are facing greater pressures to survive in a competitive environment. They must improve their financial competitiveness. Based on the conceptual framework of financial competitiven…
Training an agent to solve control tasks directly from high-dimensional images with model-free reinforcement learning (RL) has proven difficult. A promising approach is to learn a latent representation together with the control policy. However, fitting a high-capacity encoder using a scarce reward signal is sample inef…
Enhances large language models' reasoning through simpler off-policy reinforcement learning.
We consider assortment optimization over a continuous spectrum of products represented by the unit interval, where the seller's problem consists of determining the optimal subset of products to offer to potential customers. To describe the relation between assortment and customer choice, we propose a probabilistic choi…
Here, the concept of electric capacity on Finsler spaces is introduced and the fundamental conformal invariant property is proved, i.e. the capacity of a compact set on a connected non-compact Finsler manifold is conformal invariant. This work enables mathematicians and theoretical physicists to become more familiar wi…
We introduce the concept of hereditarily non uniformly perfect sets, compact sets for which no compact subset is uniformly perfect, and compare them with the following: Hausdorff dimension zero sets, logarithmic capacity zero sets, Lebesgue 2-dimensional measure zero sets, and porous sets. In particular, we give an exa…