Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

59118177236 · Jun 202019922001200920172026
48 results for Correlated Actions

Unified Bayesian framework for efficient off-policy evaluation and learning in large action spaces.

problem Efficient off-policy evaluation and learning in systems with correlated actions.
method Unified Bayesian framework with structured priors and sDM approach.
result sDM leverages action correlations without compromising computational efficiency.

Simplifies large action space bandits by selecting representative actions.

problem Efficiently managing large action spaces with correlated outcomes.
method Random sampling and solving of bandit instances to identify representative actions.
result The algorithm selects a smaller set of representative actions that perform nearly as well as the full action space.

Many decision-making problems naturally exhibit pronounced structures inherited from the characteristics of the underlying environment. In a Markov decision process model, for example, two distinct states can have inherently related semantics or encode resembling physical state configurations. This often implies locall…

2019-09-11abs ↗pdf ↗

V-learning tackles multiagent reinforcement learning by reducing sample complexity.

problem Curse of multiagents in multiagent reinforcement learning.
method V-learning is a fully decentralized algorithm that learns Nash, correlated, and coarse correlated equilibria.
result V-learning achieves sample complexity that scales with the maximum number of actions per agent, not the joint action space.

Paper develops efficient algorithms for learning rationalizable equilibria in multiplayer games.

problem Learning rationalizable behavior in multiplayer games under bandit feedback.
method New algorithms for finding rationalizable Coarse Correlated Equilibria and Correlated Equilibria with polynomial sample complexity.
result Achieved polynomial sample complexity for learning rationalizable equilibria, improving over existing exponential complexity.

A new method for disentangling action sequences improves model stability.

problem Challenges in unsupervised disentanglement learning due to incomplete theories and abstract notions.
method Introducing disentangling action sequences and a novel fractional variational autoencoder (FVAE) framework.
result FVAE improves the stability of disentanglement for action sequences.

Generalizes Hodge correlators using quantum master equation concepts.

problem Developing a mathematical framework for non-acyclic Chern-Simons theory.
method Introduces a DG Lie algebra of uni-trivalent graphs with loops satisfying a Maurer-Cartan equation.
result Arithmetic analogue of effective action and quantum master equation.

A new reinforcement learning method reduces action complexity for robust control.

problem Deep reinforcement learning's susceptibility to spurious correlations.
method Minimizing trajectory entropy to encourage simple, predictable actions.
result Trajectory Entropy Reinforcement Learning achieves superior performance and robustness.

Researchers develop geodesics for a new metric on correlation matrices.

problem Lack of intrinsic tools for statistical analyses of correlation matrices.
method Developed geodesics for the quotient-affine metric on full-rank correlation matrices.
result Provided fundamental Riemannian operations for the quotient-affine metric.

Study how untrained policies explore in RL environments.

problem Challenges in reinforcement learning, especially sparse or adversarial reward structures.
method Theoretical and empirical analysis of untrained deep neural policies in a toy model.
result Untrained policies generate correlated actions and non-trivial state-visitation distributions.

New setting combines state evolution and corrupted context for better decision-making.

problem Decision-making in a changing state with unreliable context.
method Proposes a new algorithm using a referee to dynamically combine contextual bandit and multi-armed bandit policies.
result Improved empirical performance compared to existing algorithms.

Gaussian copulas are widely used in the industry to correlate two random variables when there is no prior knowledge about the co-dependence between them. The perturbed Gaussian copula approach allows introducing the skew information of both random variables into the co-dependence structure. The analytical expression of…

2010-02-27abs ↗pdf ↗

The simplest field theory description of the multivariate statistics of forward rate variations over time and maturities, involves a quadratic action containing a gradient squared rigidity term. However, this choice leads to a spurious kink (infinite curvature) of the normalized correlation function for coinciding matu…

2004-03-29abs ↗pdf ↗

Recurrent networks learn beliefs from history in partially observable environments.

problem Learning optimal policies in partially observable environments.
method Trained recurrent neural networks to approximate value functions, measuring mutual information between hidden states and beliefs.
result Recurrent networks' hidden states correlate with beliefs of relevant state variables, improving expected return.

New metrics defined for full-rank correlation matrices, ensuring unique operations.

problem No suitable problem statement as the abstract does not describe a problem to be solved.
method New Riemannian metrics defined on full-rank correlation matrices, providing unique operations.
result Unique Riemannian logarithm and Fréchet mean defined for full-rank correlation matrices.

A statistical generalization is made of microeconomics in the spirit of going from classical to statistical mechanics. The price and quantity of every commodity1 traded in the market, at each instant of time, is considered to be an independent random variable: all prices and quantities are considered to be stochastic p…

2012-11-30abs ↗pdf ↗

This paper improves sample efficiency for learning equilibria in multi-player games.

problem Sample-efficient learning of equilibria in games with many players.
method Designs algorithms for learning CCE and CE with polynomial sample complexity in the number of players.
result First to show polynomial sample complexity for learning CCE and CE in multi-player games.

The paper explains credit decisions using Shapley decomposition for adverse actions.

problem Identifying predictors responsible for adverse credit decisions.
method Develops a simple and intuitive approach based on Shapley decomposition for models with low-order interactions.
result Shows the approach generalizes to Shapley decomposition and Baseline Shapley.

We consider the problem of learning to play a repeated multi-agent game with an unknown reward function. Single player online learning algorithms attain strong regret bounds when provided with full information feedback, which unfortunately is unavailable in many real-world scenarios. Bandit feedback alone, i.e., observ…

2019-09-18abs ↗pdf ↗

We consider the SL(2,R)SL(2,R) action on moduli spaces of quadratic differentials. If μμ is an SL(2,R)SL(2,R)-invariant probability measure, crucial information about the associated representation on L2(μ)L^2(μ) (and in particular, fine asymptotics for decay of correlations of the diagonal action, the Teichmüller flow) is encoded …

2010-11-24abs ↗pdf ↗

New algorithms reduce regret in online learning with imperfect hints.

problem Designing algorithms to minimize regret in online learning with imperfect hints.
method Developed algorithms that are resilient to bad hints and interpolate between correlated and no-hints cases.
result Achieved nearly matching lower bounds for online learning with imperfect directional hints.

Algorithm optimizes bandit decisions with changing action sets using Gaussian processes.

problem Optimizing decisions in a bandit problem with time-varying action sets.
method Proposes an algorithm called O'CLOK-UCB using Gaussian processes to handle changing action sets and contexts.
result Achieves regret bound of ildeO(λ(K)KTγKT(tTXt)) ilde{O}(\sqrt{λ^*(K)KTγ_{KT}(\cup_{t\leq T}\mathcal{X}_t)} ) with high probability.

We give geometric explanations and proofs of various mirror symmetry conjectures for TnT^{n}-invariant Calabi-Yau manifolds when instanton corrections are absent. This uses fiberwise Fourier transformation together with base Legendre transformation. We discuss mirror transformations of (i) moduli spaces of complex stru…

2000-09-27abs ↗pdf ↗

StakeBench evaluates language understanding by linking comments to market commitments, improving model alignment with real-world outcomes.

problem Existing financial NLP benchmarks measure perceived language rather than market commitments.
method StakeBench uses observable market behavior to supervise models, testing their ability to detect commitments, identify sides, and project odds.
result Models partially recover position-side signals but struggle with later tasks, highlighting structural failures.

We tackle the Multi-task Batch Reinforcement Learning problem. Given multiple datasets collected from different tasks, we train a multi-task policy to perform well in unseen tasks sampled from the same distribution. The task identities of the unseen tasks are not provided. To perform well, the policy must infer the tas…

2019-09-25abs ↗pdf ↗

The existence of stationary Markov perfect equilibria in stochastic games is shown under a general condition called "(decomposable) coarser transition kernels". This result covers various earlier existence results on correlated equilibria, noisy stochastic games, stochastic games with finite actions and state-independe…

2013-11-07abs ↗pdf ↗

Quantum field theory connects deep neural networks to criticality.

problem Understanding the criticality and training dynamics of deep neural networks.
method Constructing quantum field theory for deep neural networks, computing corrections to correlation functions.
result Found precise analogy with O(N)O(N) vector model, providing corrections to correlation length.