Unified Bayesian framework for efficient off-policy evaluation and learning in large action spaces.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Simplifies large action space bandits by selecting representative actions.
New RL algorithm tackles complex discrete action spaces.
This is the second of two papers in which we prove that a cell model of the moduli space of curves with marked points and tangent vectors at the marked points acts on the Hochschild co--chains of a Frobenius algebra. We also prove that a there is dg--PROP action of a version of Sullivan Chord diagrams which acts on the…
Models analyze strategic risk-taking in continuous action games.
Proposes meTS for efficient exploration in correlated bandits.
Adaptive correlated MC improves sequence generation stability.
Diffusion models mimic human actions in sequential tasks.
Many decision-making problems naturally exhibit pronounced structures inherited from the characteristics of the underlying environment. In a Markov decision process model, for example, two distinct states can have inherently related semantics or encode resembling physical state configurations. This often implies locall…
A new algorithm for bandits with hierarchical rewards.
V-learning tackles multiagent reinforcement learning by reducing sample complexity.
Paper develops efficient algorithms for learning rationalizable equilibria in multiplayer games.
A new method for disentangling action sequences improves model stability.
Generalizes Hodge correlators using quantum master equation concepts.
A new reinforcement learning method reduces action complexity for robust control.
Researchers develop geodesics for a new metric on correlation matrices.
In this paper, we present an approach for identification of actions within depth action videos. First, we process the video to get motion history images (MHIs) and static history images (SHIs) corresponding to an action video based on the use of 3D Motion Trail Model (3DMTM). We then characterize the action video by ex…
We develop a normative framework for hierarchical model-based policy optimization based on applying second-order methods in the space of all possible state-action paths. The resulting natural path gradient performs policy updates in a manner which is sensitive to the long-range correlational structure of the induced st…
Study how untrained policies explore in RL environments.
In this paper we extend our correlation functions to the open/closed case. This gives rise to actions of an open/closed version of the Sullivan PROP as well as an action of the relevant moduli space. There are several unexpected structures and conditions that arise in this extension which are forced upon us by consider…
Advanced driver assistance systems (ADAS) can be significantly improved with effective driver action prediction (DAP). Predicting driver actions early and accurately can help mitigate the effects of potentially unsafe driving behaviors and avoid possible accidents. In this paper, we formulate driver action prediction a…
Reinforcement learning (RL) agents performing complex tasks must be able to remember observations and actions across sizable time intervals. This is especially true during the initial learning stages, when exploratory behaviour can increase the delay between specific actions and their effects. Many new or popular appro…
Method preserves correlations in synthetic data.
New setting combines state evolution and corrupted context for better decision-making.
Gaussian copulas are widely used in the industry to correlate two random variables when there is no prior knowledge about the co-dependence between them. The perturbed Gaussian copula approach allows introducing the skew information of both random variables into the co-dependence structure. The analytical expression of…
Improved CEM for fast real-time planning in high-dimensional control tasks.
The simplest field theory description of the multivariate statistics of forward rate variations over time and maturities, involves a quadratic action containing a gradient squared rigidity term. However, this choice leads to a spurious kink (infinite curvature) of the normalized correlation function for coinciding matu…
Recurrent networks learn beliefs from history in partially observable environments.
A variety of machine learning models have been proposed to assess the performance of players in professional sports. However, they have only a limited ability to model how player performance depends on the game context. This paper proposes a new approach to capturing game context: we apply Deep Reinforcement Learning (…
New metrics defined for full-rank correlation matrices, ensuring unique operations.
A statistical generalization is made of microeconomics in the spirit of going from classical to statistical mechanics. The price and quantity of every commodity1 traded in the market, at each instant of time, is considered to be an independent random variable: all prices and quantities are considered to be stochastic p…
New algorithm tackles nonstationary linear bandits with latent dynamics.
This paper improves sample efficiency for learning equilibria in multi-player games.
This is the first of two papers in which we prove that a cell model of the moduli space of curves with marked points and tangent vectors at the marked points acts on the Hochschild co--chains of a Frobenius algebra. We also prove that a there is dg--PROP action of a version of Sullivan Chord diagrams which acts on the …
The paper explains credit decisions using Shapley decomposition for adverse actions.
We consider the problem of learning to play a repeated multi-agent game with an unknown reward function. Single player online learning algorithms attain strong regret bounds when provided with full information feedback, which unfortunately is unavailable in many real-world scenarios. Bandit feedback alone, i.e., observ…
We consider the action on moduli spaces of quadratic differentials. If is an -invariant probability measure, crucial information about the associated representation on (and in particular, fine asymptotics for decay of correlations of the diagonal action, the Teichmüller flow) is encoded …
New algorithms reduce regret in online learning with imperfect hints.
Study optimal offline RL with uncertainty sets and distribution shifts.
We present a scheme for online, unsupervised state discovery and detection from streaming, multi-featured, asynchronous data in high-frequency financial markets. Online feature correlations are computed using an unbiased, lossless Fourier estimator. A high-speed maximum likelihood clustering algorithm is then used to f…
Algorithm optimizes bandit decisions with changing action sets using Gaussian processes.
We give geometric explanations and proofs of various mirror symmetry conjectures for -invariant Calabi-Yau manifolds when instanton corrections are absent. This uses fiberwise Fourier transformation together with base Legendre transformation. We discuss mirror transformations of (i) moduli spaces of complex stru…
Policy gradient methods have demonstrated success in reinforcement learning tasks that have high-dimensional continuous state and action spaces. However, policy gradient methods are also notoriously sample inefficient. This can be attributed, at least in part, to the high variance in estimating the gradient of the task…
New algorithm for partially observable contexts in finance.
StakeBench evaluates language understanding by linking comments to market commitments, improving model alignment with real-world outcomes.
We tackle the Multi-task Batch Reinforcement Learning problem. Given multiple datasets collected from different tasks, we train a multi-task policy to perform well in unseen tasks sampled from the same distribution. The task identities of the unseen tasks are not provided. To perform well, the policy must infer the tas…
The existence of stationary Markov perfect equilibria in stochastic games is shown under a general condition called "(decomposable) coarser transition kernels". This result covers various earlier existence results on correlated equilibria, noisy stochastic games, stochastic games with finite actions and state-independe…
Quantum field theory connects deep neural networks to criticality.