Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

1.9%3.9%5.8%7.8% · Jun 201919922001200920182026
48 results for anti-recidivism policies

New method uses ML to estimate causal effects from past data.

problem Estimating causal effects from retrospective data with ambiguity and bias.
method Machine learning ensemble targeting a specific causal effect (RIE).
result Validates policy options for reducing recidivism in Colombia.

The paper shows how to improve policies on- and off-policy using bounds.

problem Improving reinforcement learning policies on- and off-policy.
method Lower bounding the performance difference of two policies to ensure monotonic improvement from mixture samples.
result An optimization procedure that applies the proposed bound can be seen as an off-policy natural policy gradient method.

Paper tackles efficient evaluation of natural stochastic policies in offline RL.

problem Efficiency issues in evaluating natural stochastic policies due to unknown evaluation policy.
method Derive efficiency bounds for tilting and modified treatment policies, propose nonparametric estimators.
result Proposed estimators attain efficiency bounds under lax conditions and enjoy partial double robustness.

Study shows accurate OPE depends on calibrated behaviour policy models.

problem Estimating a behaviour policy for OPE when true policy is unknown.
method Empirical studies comparing parametric vs non-parametric models.
result Simple non-parametric k-nearest neighbors model produces better calibrated behaviour policy estimates.

Stabilizes policy optimization with off-policy data using divergence augmentation.

problem Premature convergence and instability in policy optimization with off-policy data.
method Incorporates Bregman divergence between behavior and current policies to ensure safe policy updates.
result Empirically shows better performance in data-scarce scenarios compared to other algorithms.

New method estimates state-action stationary distribution for better off-policy policy evaluation.

problem Accurately estimating state-action stationary distribution for off-policy policy evaluation.
method Estimated Mixture Policy (EMP) for state and state-action stationary distribution corrections.
result Empirical validation shows improved accuracy over state-of-the-art methods.

New methods estimate policy value and gradients for deterministic policies from off-policy data.

problem Estimating policy value and gradients for deterministic policies from off-policy data.
method Proposed new doubly robust estimators based on kernelization approaches.
result Demonstrated a rate independent of horizon length for policy value and gradient estimation.

Paper explores using action-value gradients for policy improvement in off-policy actor-critic methods.

problem Improving policies using action-value gradients in off-policy stochastic actor-critic methods.
method Discusses and analyzes the use of action-value gradients for policy improvement, and proposes an incremental approach.
result Demonstrates the feasibility and incremental approach for following the policy gradient.

DSPI connects natural policy gradient to policy iteration, proving global convergence.

problem Optimizing policies in reinforcement learning.
method DSPI framework, combining smoothed policy iteration and natural policy gradient.
result DSPI achieves geometric convergence and optimal complexity for policy optimization.

New method improves off-policy policy evaluation by balancing policy values.

problem Accurate off-policy policy evaluation in reinforcement learning.
method Developed an MDP model with balanced representation to estimate both individual and average policy values.
result Substantially lower Mean Squared Error (MSE) in various benchmarks and a real-world HIV treatment simulation.

Normalizing flows policy improves trust region policy optimization.

problem Improving exploration and avoiding local optima in policy optimization.
method Constructing trust region with KL divergence constraints and using normalizing flows policy.
result Normalizing flows policy significantly improves policy optimization, especially on high-dimensional tasks.

POTEC tackles off-policy learning in large action spaces, improving effectiveness.

problem Existing OPL methods fail in large discrete action spaces due to bias or variance issues.
method Two-stage algorithm: cluster selection via policy-based approach, action selection via regression-based approach.
result POTEC provides substantial improvements in off-policy learning effectiveness, especially in large and structured action spaces.

A new theorem and algorithm solve off-policy policy gradient problems.

problem Solving the theoretical gap in off-policy policy gradient methods.
method Introduced an off-policy policy gradient theorem using emphatic weightings and developed the ACE algorithm.
result Demonstrated ACE finds the optimal solution in off-policy learning, unlike previous methods.

Protects proprietary policies from imitation learning by training adversarial policy ensembles.

problem Protecting policies from external observers cloning them.
method Introduces a reinforcement learning framework that trains an ensemble of near-optimal policies, making demonstrations useless for external observers.
result Demonstrates the existence of 'non-clonable' ensembles and provides a solution to the optimization problem.

PS framework selects best policy from library for CSO problems.

problem Policy selection in CSO with heterogeneous performance across covariate space.
method PS framework constructs library of candidate policies and learns a meta-policy to select the best one.
result PS consistently outperforms best single policy in heterogeneous CSO problems.

Entropy regularization improves policy optimization in reinforcement learning.

problem Improving policy optimization in reinforcement learning.
method Entropy regularization is introduced to soften the greedy policy towards a more diverse softmax policy, leading to a continuously parameterized algorithm that interpolates between policy gradient and Q-learning.
result An intermediate algorithm can improve performance in reinforcement learning.

PBVFs generalize across policies using learned value functions.

problem RL algorithms forget information about old policies when updating value functions to track the learned policy.
method Introduce Parameter-Based Value Functions (PBVFs) that include policy parameters in their inputs, enabling them to generalize across different policies.
result PBVFs enable zero-shot learning of new policies that outperform any policy seen during training.

Policy gradient aims to maximize expected return using gradient ascent.

problem Finding a policy that maximizes expected return in a given class of policies.
method Gradient ascent applied to a differentiable model of the policy, estimating the gradient of expected return.
result Policy gradient methods require on-policy data for gradient estimation, limiting sample efficiency.

Study optimizes portfolio allocation policies using off-policy data and constraints.

problem Optimizing portfolio allocation policies under constraints using off-policy data.
method Solves a minimax objective with off-policy estimators and online learning to control constraint violations.
result Constructs near-optimal allocation policies for various regimes of operation and constraints.

Paper introduces a new policy optimization method using importance sampling.

problem Stable and low variance policy learning with small policy updates.
method Derives an alternative objective using importance sampling and introduces an approximation to balance bias and variance.
result The new algorithm improves on-policy policy optimization on continuous control benchmarks.

This paper introduces a new method to evaluate multiple policies simultaneously.

problem Estimating the value of many policies for a single set of states.
method Developed a scalable, differentiable fingerprinting mechanism to represent complex policies.
result The method can produce policies that outperform those that generated the training data, in zero-shot manner.

New algorithm improves reinforcement learning policies without degrading performance.

problem Policy updates may degrade performance in reinforcement learning with general function approximators.
method Derives a new policy improvement bound with an average divergence instead of sup norm, leading to Easy Monotonic Policy Iteration.
result Generates sequences of policies with guaranteed non-decreasing returns.

The paper interprets policy-gradient algorithms using continuation theory.

problem Optimizing nonconvex functions in reinforcement learning.
method Formulates policy optimization as optimization by continuation, interprets policy-gradient algorithms as implicitly optimizing deterministic policies.
result Exploration in policy-gradient algorithms is seen as computing a continuation of the return of the policy.

This work uses MMD to find diverse policies in reinforcement learning.

problem Finding multiple near-optimal policies for a task.
method Formalizes policy difference as trajectory distribution discrepancy, uses MMD for optimization.
result Derives gradient-based optimization for diverse policy identification.

New method reduces state distribution mismatch in off-policy RL.

problem State distribution mismatch in off-policy RL algorithms.
method Develops a novel constrained off-policy gradient objective to minimize state distribution shift.
result Minimizing state distribution shift improves performance in off-policy RL algorithms.

Paper develops a new multi-agent reinforcement learning algorithm.

problem Improving policies in a network of communicating agents.
method Develops a multi-agent off-policy actor-critic algorithm using emphatic temporal difference learning.
result Proves convergence of the algorithm under linear function approximation.

This work analyzes the gap between off-policy and on-policy policy gradient methods and provides conditions to reduce this gap.

problem The gap between off-policy and on-policy policy gradient methods and conditions to reduce it.
method Theoretical analysis and empirical evidence of conditions to reduce the on-off gap.
result Conditions to reduce the on-off gap between off-policy and on-policy policy gradient methods.

We make policy optimization algorithms batch size-invariant by decoupling proximal and behavior policies.

problem Some policy optimization algorithms do not have batch size-invariance, leading to inefficiencies.
method We decouple the proximal policy from the behavior policy to achieve batch size-invariance.
result Our approach makes policy optimization algorithms more efficient and allows them to use stale data more effectively.