Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

82165247329 · Jun 202019922001200920172026
48 results for action probabilities

We solve offline learning problems without action probabilities using imitation and policy improvement.

problem Offline learning in automated decision systems with missing action probabilities.
method We use policy improvement and imitation regularization to estimate action policies.
result Our method improves upon existing approaches by reducing uncertainty and improving accuracy.

The paper introduces metrics to rank potential outcomes for better decision-making.

problem Optimal action selection in uncertain situations using causal reasoning.
method Introducing two new metrics: probabilities of potential outcome ranking (PoR) and probability of achieving the best potential outcome (PoB). Establishing identification theorems and deriving bounds for these metrics, and presenting estimation methods.
result The estimators' finite-sample properties and their application to a real-world dataset are demonstrated.

Study positive entropy actions by higher-rank lattices, proving rigidity and conjugacy results.

problem Positive entropy actions by higher-rank lattices in Lie groups.
method Analysis of sub-actions, fiber entropy upper semicontinuity, and conjugacy arguments.
result Actions by higher-rank lattices in SL(n,R)\mathrm{SL}(n,\mathbb{R}) are conjugate to affine actions on (infra-)tori.

Safe actions learned in finite trials, without infinite exploration.

problem Learning safe actions in unknown environments efficiently.
method Defining a handicap metric and using sequential probability ratio test for discarding unsafe actions.
result Achieves constant handicap, discarding unsafe machines with probability one in finite rounds.

Develops a machine learning framework for computing most probable paths in stochastic systems.

problem Computing the most probable paths in stochastic dynamical systems.
method Reformulates the boundary value problem of Hamiltonian systems and uses a neural network to solve the Euler-Lagrange equation for the Onsager-Machlup action functional.
result Demonstrates the efficacy and accuracy of the machine learning approach in computing most probable paths for stochastic systems with various types of noise.

Deep Reinforcement Learning improves with Weighted Q-Learning to reduce bias and uncertainty.

problem Overestimation and high variance in Q-Learning cause learning algorithms to diverge in complex environments.
method Deep Weighted Q-Learning (Deep WQL) uses Dropout and Monte Carlo sampling to approximate WQL's weights and reduce bias.
result Deep WQL reduces bias and improves performance on benchmarks compared to existing methods.

In this note we prove the a pointwise ergodic theorem for functions taking values in a separable complete CAT(0)-space, analogous to Lindenstrauss' pointwise ergodic theorem for real-valued integrable functions on a probability space subject to a probability-preserving action of an amenable l.c.s.c. group, where in the…

2009-05-05abs ↗pdf ↗

Paper studies zero-sum games with noisy observations and identifies equilibrium conditions.

problem Zero-sum games with noisy observations of the leader's actions.
method Analyzes the equilibrium of games with noisy action observability, identifies necessary conditions for uniqueness, and investigates the cardinality of best responses.
result The noisy observations significantly impact the cardinality of the follower's set of best responses, and under certain conditions, this set becomes a singleton almost surely.

The study proves curvature bounds for quotient spaces of isometric actions.

problem Proving curvature bounds for quotient spaces of isometric actions.
method Disintegrate absolutely continuous measures and define a functional to prove curvature bounds.
result Necessary and sufficient conditions for Ricci curvature to be bounded below.

This paper explores the possibility that asset prices, especially those traded in large volume on public exchanges, might comply with specific physical laws of motion and probability. The paper first examines the basic dynamics of asset price displacement and finds one can model this dynamic as a harmonic oscillator at…

2017-05-28abs ↗pdf ↗

Let M1M_1 and M2M_2 be two nn-dimensional smooth manifolds with boundary. Suppose we glue M1M_1 and M2M_2 along some boundary components (which are, therefore, diffeomorphic). Call the result N.N. If we have a group GG acting continuously on M1,M_1, and also acting continuously on M2,M_2, such that the actions are comp…

2012-10-08abs ↗pdf ↗

LITE efficiently estimates Gaussian PoM with linear time and memory complexity.

problem Estimating the probability of maximality (PoM) of Gaussian vectors efficiently.
method LITE: entropy-regularized UCB approach for almost-linear time and memory complexity.
result Achieves state-of-the-art accuracy with significantly faster performance than existing methods.

A 3-manifold is Haken if it contains a topologically essential surface. The Virtual Haken Conjecture posits that every irreducible 3-manifold with infinite fundamental group has a finite cover which is Haken. In this paper, we study random 3-manifolds and their finite covers in an attempt to shed light on this difficul…

2005-02-27abs ↗pdf ↗

The study explores geodesics and KL-divergence on Hölder equilibrium probabilities.

problem Finding the probability that minimizes KL-divergence from a fixed probability in a convex set of probabilities.
method Analyzes geodesics paths on the manifold of Hölder equilibrium probabilities and uses KL-divergence as a metric.
result Explicit equations for the solution of the minimization problem are derived.

We propose an analytically tractable variation of the minority game in which rational agents use probabilistic strategies. In our model, NN agents choose between two alternatives repeatedly, and those who are in the minority get a pay-off 1, others zero. The agents optimize the expectation value of their discounted fu…

2012-12-29abs ↗pdf ↗

This work formalizes robustness criteria for reinforcement learning actions and improves performance in perturbed environments.

problem Improving reinforcement learning policies to perform well in uncertain or adversarial action scenarios.
method Formalized two robustness criteria for reinforcement learning actions, considering adversarial actions and action perturbations. Developed algorithms for tabular and deep reinforcement learning settings.
result Action-robust reinforcement learning policies improve performance in perturbed environments and are a form of implicit regularization.

Given a Kähler manifold (Z,J,ω)(Z,J,ω) and a compact real submanifold MZM\subset Z, we study the properties of the gradient map associated with the action of a noncompact real reductive Lie group G{\rm G} on the space of probability measures on M.M. In particular, we prove convexity results for such map when G{\rm G} is A…

2017-01-17abs ↗pdf ↗

The isotropy action on certain symmetric spaces is shown to be equivariantly formal.

problem Understanding the equivariant formality of isotropy actions on symmetric spaces.
method Developed a new approach to prove equivariant formality for (Z2Z2)(\mathbb{Z}_2\oplus \mathbb{Z}_2)-symmetric spaces.
result Symmetric spaces with (Z2Z2)(\mathbb{Z}_2\oplus \mathbb{Z}_2)-symmetry are equivariantly formal and formal in the Sullivan sense.

Develops a dynamic mean field theory for reinforcement learning.

problem Finite state and action Bayesian reinforcement learning in large state spaces.
method Analogies with statistical physics, interpreting probabilities as couplings and values as spins, solving mean field equations.
result State-action values are statistically independent in the asymptotic state space limit, with exact or approximate equations for computation.

We study the sparse entropy-regularized reinforcement learning (ERL) problem in which the entropy term is a special form of the Tsallis entropy. The optimal policy of this formulation is sparse, i.e.,~at each state, it has non-zero probability for only a small number of actions. This addresses the main drawback of the …

2018-02-10abs ↗pdf ↗

Let S be a non-exceptional oriented surface of finite type. We discuss the action of subgroups of the mapping class group of S on the CAT(0)-boundary of the completion of Teichmueller space with respect to the Weil-Petersson metric. We show that the set of invariant Borel probability measures for the Weil-Petersson flo…

2009-01-27abs ↗pdf ↗

Study on information cascade fragility under mismatched revealing probabilities.

problem Analyzing the fragility of information cascades in decision-making processes with imperfect knowledge of revealing probabilities.
method Examined sequential decision-making models with players having private information and imitating previous decisions. Studied the effect of a mismatch between players' beliefs and actual revealing probabilities.
result Derived closed-form expressions for optimal learning rates and identified phase transitions in the behavior of asymptotic learning rates.

A new framework for offline RL improves policy flexibility and regularity.

problem Lack of environmental interactions in offline RL leads to poor policy performance.
method Proposes a behavior-regularized implicit policy framework with modified policy-matching methods.
result The framework improves policy effectiveness and robustness beyond static datasets.

A new method estimates multi-dimensional value distributions using Hilbert space embeddings.

problem Estimating value distributions in complex, multi-dimensional reinforcement learning settings.
method Hilbert space mappings and kernel mean embeddings to estimate the kernel mean embedding of multi-dimensional value distributions.
result Uniform convergence guarantees and robust off-policy evaluation demonstrated in simulations.

This paper uses probability tensors for efficient path planning in complex scenarios.

problem Efficient path planning in complex environments with obstacles and multiple goals.
method Probability tensors are used to model agent motion and decision-making, incorporating past and future information.
result The model finds solutions in complex scenarios, demonstrating realistic emergent behaviors.

We introduce a new class of reinforcement learning methods referred to as {\em episodic multi-armed bandits} (eMAB). In eMAB the learner proceeds in {\em episodes}, each composed of several {\em steps}, in which it chooses an action and observes a feedback signal. Moreover, in each step, it can take a special action, c…

2015-08-04abs ↗pdf ↗

Central limit theorem for Green metrics on hyperbolic groups.

problem Proving a central limit theorem for Green metrics on hyperbolic groups.
method Proving a central limit theorem for Green metrics on hyperbolic groups using probability measures and ordering elements.
result Proved a central limit theorem for Green metrics on hyperbolic groups.

Efficient algorithms identify true hypothesis from many options with minimal actions.

problem Identifying true hypothesis from a large set of options with minimal actions.
method Greedy approximation algorithms for active sequential hypothesis testing.
result First approximation guarantees for ASHT, independent of the number of hypotheses.

Uplift modeling is an area of machine learning which aims at predicting the causal effect of some action on a given individual. The action may be a medical procedure, marketing campaign, or any other circumstance controlled by the experimenter. Building an uplift model requires two training sets: the treatment group, w…

2018-07-20abs ↗pdf ↗

A new method combines online and offline learning to tackle contextual bandits with missing action support.

problem Learning optimal policies with logged data when the logging policy has deficient support.
method Hybrid approach using online exploration to exploit supported actions and offline learning to avoid unnecessary explorations.
result Determines an optimal policy with theoretical guarantees using minimal online explorations.

New complexity analysis for estimating normalizing constants in high dimensions.

problem Estimating the normalizing constant of unnormalized probability densities in high dimensions.
method Analyze and derive the oracle complexity of annealed importance sampling.
result Oracle complexity of $\widetilde{O}\left(\frac{dβ^2{\mathcal{A}}^2}{\varepsilon^4} ight)$ for estimating ZZ within ε\varepsilon relative error.

TRAiL is a linear bandit algorithm that ensures optimal regret and guarantees inference quality.

problem Optimal regret and inference quality in linear bandits with convex action sets.
method TRAiL estimates the parameter through regularized least squares and perturbs the action set along the tangent plane.
result TRAiL achieves an Ω(T)Ω(\sqrt{T}) upper bound on cumulative regret with high probability.

Develops CLTs for Markov chain transition probabilities and policies.

problem Estimating transition probabilities and policies in controlled Markov chains.
method Non-parametric estimator for transition matrices; CLTs for value, Q-, and advantage functions; goodness-of-fit tests.
result Asymptotic normality of estimators under specific logging policies.

DG improves policy gradients by weighting actions with a sigmoid of advantage and surprisal.

problem Pathologies in standard policy gradients, leading to poor updates and over-allocation of gradient budget.
method Introduces Delightful Policy Gradient (DG) that gates each term with a sigmoid of advantage and surprisal.
result DG provably improves directional accuracy in a single context and shifts the expected gradient closer to the oracle across multiple contexts.

RegFlow models future states with flexible probability distributions.

problem Predicting future states under complex, non-deterministic scenarios.
method Hypernetwork architecture and continuous normalizing flow model.
result RegFlow achieves state-of-the-art results on benchmark datasets.

Safe imitation learning with a safety layer for flexible training.

problem Flexible yet safe imitation learning for complex tasks.
method Theory and modular method with a safety layer for continuous policy, adversarial training, and worst-case safety guarantees.
result Robustness advantage of safety layer during training compared to test time.

Let (M,ω)(M,ω) be a Kähler manifold and let KK be a compact group that acts on MM in a Hamiltonian fashion. We study the action of KCK^\mathbb{C} on probability measures on MM. First of all we identify an abstract setting for the momentum mapping and give numerical criteria for stability, semi-stability and polystabili…

2015-12-13abs ↗pdf ↗

Study geometric and representation theory of statistical transformation models.

problem Understand relationships between induced structures and actions on measure spaces.
method Investigate geometric properties and symplectic actions on induced structures.
result Show equivariance of action and relationships between tangent bundles and projectivizations.