BRPO optimizes batch RL policies to better exploit state-action differences.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New method for distributional off-policy evaluation using Bellman residual minimization.
This paper aims at theoretically and empirically comparing two standard optimization criteria for Reinforcement Learning: i) maximization of the mean value and ii) minimization of the Bellman residual. For that purpose, we place ourselves in the framework of policy search algorithms, that are usually designed to maximi…
KBB algorithm reduces sample complexity for policy evaluation in general state spaces.
RFC enhances humanoid control to imitate complex human motions.
Proposes log density gradient to improve reinforcement learning sample complexity.
New estimator improves off-policy evaluation for large action spaces.
Recent work in fairness in machine learning has proposed adjusting for fairness by equalizing accuracy metrics across groups and has also studied how datasets affected by historical prejudices may lead to unfair decision policies. We connect these lines of work and study the residual unfairness that arises when a fairn…
New method uses observational data to improve trial design efficiency.
CEFOL uses deep learning for dynamic programming with recursive utility.
Minimal assumptions analysis of Q-learning with time-varying policies.
New method improves deep policy gradient algorithms by learning relative state values.
New method for evaluating and learning in complex decision-making scenarios.
A method to reduce bias in model-based policy evaluation by shifting operators.
Temporal difference learning and Residual Gradient methods are the most widely used temporal difference based learning algorithms; however, it has been shown that none of their objective functions is optimal w.r.t approximating the true value function . Two novel algorithms are proposed to approximate the true value…
Develop a decision-calibrated conformal framework for pacing decisions in streaming advertising.
VA-OPE improves OPE by incorporating variance information, achieving tighter error bounds.
Experience reuse is key to sample-efficient reinforcement learning. One of the critical issues is how the experience is represented and stored. Previously, the experience can be stored in the forms of features, individual models, and the average model, each lying at a different granularity. However, new tasks may requi…
DAGR improves navigation by refining goal representations conditioned on the current state.
DOLCE improves off-policy evaluation and learning by decomposing effects.
We present a new approach to the problems of evaluating and learning personalized decision policies from observational data of past contexts, decisions, and outcomes. Only the outcome of the enacted decision is available and the historical policy is unknown. These problems arise in personalized medicine using electroni…
Transfer reinforcement learning (RL) aims at improving the learning efficiency of an agent by exploiting knowledge from other source agents trained on relevant tasks. However, it remains challenging to transfer knowledge between different environmental dynamics without having access to the source environments. In this …
In this paper, two Q-learning (QL) methods are proposed and their convergence theories are established for addressing the model-free optimal control problem of general nonlinear continuous-time systems. By introducing the Q-function for continuous-time systems, policy iteration based QL (PIQL) and value iteration based…
The control design problem is considered for nonlinear systems with unknown internal system model. It is known that the nonlinear control problem can be transformed into solving the so-called Hamilton-Jacobi-Isaacs (HJI) equation, which is a nonlinear partial differential equation that is genera…
A new method for MARL with partial observations reduces communication overhead.
Develops statistical framework for resolving reward function ambiguity in inverse reinforcement learning.
We seek to learn an effective policy for a Markov Decision Process (MDP) with continuous states via Q-Learning. Given a set of basis functions over state action pairs we search for a corresponding set of linear weights that minimizes the mean Bellman residual. Our algorithm uses a Kalman filter model to estimate those …
This paper improves Thompson Sampling for complex decision-making problems.
This paper addresses the model-free nonlinear optimal problem with generalized cost functional, and a data-based reinforcement learning technique is developed. It is known that the nonlinear optimal control problem relies on the solution of the Hamilton-Jacobi-Bellman (HJB) equation, which is a nonlinear partial differ…
Abstract: Non-residually finite hyperbolic groups imply non-residually finite rigid hyperbolic groups.
Residual finiteness is known to be an important property of groups appearing in combinatorial group theory and low dimensional topology. In a recent work [2] residual finiteness of quandles was introduced, and it was proved that free quandles and knot quandles are residually finite. In this paper, we extend these resul…
In this note, residual finiteness of quandles is defined and investigated. It is proved that free quandles and knot quandles of tame knots are residually finite and Hopfian. Residual finiteness of quandles arising from residually finite groups (conjugation, core and Alexander quandles) is established. Further, residual…
Every non-trivial knot group is fully residually perfect.
Residual flows are shown to approximate MMD well.
Let be a prime. In this paper, we classify the geometric 3-manifolds whose fundamental groups are virtually residually . Let be a virtually fibered 3-manifold. It is well-known that is residually solvable and even residually finite solvable. We prove that is always virtually residually …
Researchers identify critical protein residues using advanced graph theory.
Defines Wodzicki residue using groupoids and fibered distributions.
New method for personalized pricing using invalid instrumental variables.
In this work we prove a Baum-Bott type residue theorem for flags of holomorphic foliations. We prove some relations between the residues of the flag and the residues of their correspondent foliations. We define the Nash residue for flags and we give a partial answer to the Baum-Bott type rationality conjecture in this …
We revisit residual algorithms in both model-free and model-based reinforcement learning settings. We propose the bidirectional target network technique to stabilize residual algorithms, yielding a residual version of DDPG that significantly outperforms vanilla DDPG in the DeepMind Control Suite benchmark. Moreover, we…
Wide residual networks generalize well with uniform convergence to RNTK as width increases.
We have modeled the employment/population ratio in the largest developed countries. Our results show that the evolution of the employment rate since 1970 can be predicted with a high accuracy by a linear dependence on the logarithm of real GDP per capita. All empirical relationships estimated in this study need a struc…
The paper studies residues of manifolds and their applications in geometry.
Given a prime , a group is called residually if the intersection of its -power index normal subgroups is trivial. A group is called virtually residually if it has a finite index subgroup which is residually . It is well-known that finitely generated linear groups over fields of characteristic zero are …
We show that Out(G) is residually finite if G is a one-ended group that is hyperbolic relative to virtually polycyclic subgroups. More generally, if G is one-ended and hyperbolic relative to proper residually finite subgroups, the group of outer automorphisms preserving the peripheral structure is residually finite. We…
Simplifies residual flows to make flow-based modeling more practical.
Study on endomorphism and automorphism groups of specific quandles.
Okun's law for the biggest developed countries is re-estimated using the most recent data on real GDP per capita and the rate of unemployment. Our results show that the change in unemployment rate can be predicted with a high accuracy. The link needs the introduction of a structural break which might be caused by the c…