Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

4590135180 · Jun 202019922001200920182026
48 results for action identification

Study lenient regret and good-action identification in Gaussian process bandits.

problem Optimizing function values above a certain threshold in Gaussian process bandits.
method Study lenient regret notions and introduce algorithms for finding good actions.
result Upper and lower bounds on lenient regret for GP-UCB and elimination algorithms.

Paper improves action selection for accurate parameter estimation in linear bandits.

problem Best action identification in stochastic linear bandits with fixed confidence constraints.
method Designs a sequential adaptive policy to estimate underlying parameter efficiently.
result The designed policy achieves the same estimation error scaling as a lower bound.

Study robust best-arm identification in linear bandits with lower bounds and algorithms.

problem Identify a near-optimal robust arm in linear bandits with adversarial actions.
method Propose instance-dependent lower bounds and both static and adaptive bandit algorithms.
result Sample complexity matches the lower bound and algorithms effectively identify robust arms.

Simplifies causal effect identification by pruning unnecessary variables.

problem Complex expressions from causal models can lead to unnecessary variables and computational burden.
method Graphical criteria and improved identifiability algorithm to detect and prune redundant variables.
result Improved computational efficiency and reduced bias in estimating causal effects.

Paper tackles RL with continuous actions and unmeasured confounders.

problem Offline policy learning with continuous actions and unmeasured confounders.
method Developed a novel identification result and a minimax estimator for nonparametric policy value estimation.
result Introduced a policy-gradient-based algorithm to identify the optimal policy.

The Farrell-Jones Conjecture holds for groups acting acylindrically on trees.

problem Verifying the Farrell-Jones Conjecture for groups acting on trees.
method Analyzing acylindrical actions on simplicial trees and using the Farrell-Jones Conjecture.
result The Farrell-Jones Conjecture holds for groups acting acylindrically on trees.

The action of a Lie pseudogroup GG on a smooth manifold MM induces a prolonged pseudogroup action on the jet spaces JnJ^n of submanifolds of MM. We prove in this paper that both the local and global freeness of the action of GG on JnJ^n persist under prolongation in the jet order nn. Our results underlie the const…

2009-12-22abs ↗pdf ↗

The paper introduces metrics to rank potential outcomes for better decision-making.

problem Optimal action selection in uncertain situations using causal reasoning.
method Introducing two new metrics: probabilities of potential outcome ranking (PoR) and probability of achieving the best potential outcome (PoB). Establishing identification theorems and deriving bounds for these metrics, and presenting estimation methods.
result The estimators' finite-sample properties and their application to a real-world dataset are demonstrated.

Study best arm identification in restless bandits with unknown TPMs.

problem Identify the best arm with fixed confidence in restless bandits with unknown TPMs.
method Proposed a policy for best arm identification and proved its expected stopping time matches the lower bound.
result The state-action visitation proportions match the optimal proportions under any asymptotically optimal policy.

A new algorithm for identifying the best arm in linear feedback with safety constraints.

problem Identifying the best arm in linear feedback with safety constraints.
method A gap-based algorithm that ensures safety while minimizing sample complexity.
result The algorithm achieves meaningful sample complexity while ensuring safety.

The paper provides a non-asymptotic error bound for linear system identification under nonlinear policies.

problem System identification for linear systems with nonlinear and/or time-varying policies under i.i.d. random excitation noises.
method Least square estimation with non-asymptotic error bound for bounded state and action trajectories.
result The error bound is consistent with linear policies and generalizes existing guarantees.

Paper uses sparse learning to estimate quasi-potential and drift components in stochastic systems.

problem Estimating quasi-potential and drift components in stochastic systems.
method Sparse identification of non-linear dynamics (SINDy) combined with action minimization methods.
result Evaluation of quasi-potential landscape from a single trajectory.

Paper presents a method for recognizing human actions using GLAC features from motion and static images.

problem Action recognition in 3D depth videos.
method 3D Motion Trail Model (3DMTM) for MHIs and SHIs, GLAC features extraction, l2-regularized Collaborative Representation Classifier (l2-CRC) for classification.
result The method outperforms other approaches in recognizing human actions.

Algorithm identifies best arm in combinatorial bandits with semi-bandit feedback.

problem Identifying the best arm in combinatorial bandits with semi-bandit feedback.
method Interpreted as a sequential zero-sum game, developed a CombGame meta-algorithm with finite time guarantees.
result First computationally efficient algorithm that is asymptotically optimal and has competitive empirical performance.

New RL method handles hidden actions in offline learning.

problem Learning from unseen actions in real-world RL datasets.
method LURE (Learning from the Unseen: Robust Estimator) method using next-state variable as proxy.
result Valid statistical inference and improved RL conclusions with hidden actions.

Bootstrap policies improve regret in continuous state-action reinforcement learning.

problem Improving regret in reinforcement learning for continuous state and action spaces.
method Bootstrap-based policies for stochastic linear systems with quadratic cost functions.
result Bootstrap policies achieve a square root scaling of regret with respect to time.

New RLHF algorithm identifies optimal policies from human feedback without explicit reward inference.

problem Training large language models with human feedback without reward inference.
method Model-free RLHF algorithm BSAD\mathsf{BSAD} that identifies optimal policies directly from human preference.
result Provable, instance-dependent sample complexity ildeO(cMSA3H3Mlog1δ) ilde{\mathcal{O}}(c_{\mathcal{M}}SA^3H^3M\log\frac{1}δ).

The paper introduces pseudo-quotients for algebraic actions and applies them to character varieties.

problem Characterizing algebraic actions and their quotients.
method Introducing pseudo-quotients as a weak version of quotients for algebraic actions, focusing on purely topological properties.
result Pseudo-quotients are unique up to virtual class in characteristic zero and can be used to compute character varieties.

Improved sample and time complexity for identifying mixtures of product distributions.

problem Identifying a mixture of kk product distributions from statistics.
method Combining robust tensor decomposition and Hadamard extensions to bound the condition number of key matrices.
result Achieved sample complexity and run-time complexity of (1/ζ)O(k)(1/ζ)^{O(k)} for n2k1n \geq 2k-1.

Predictive State Representations (PSRs) are an expressive class of models for controlled stochastic processes. PSRs represent state as a set of predictions of future observable events. Because PSRs are defined entirely in terms of observable data, statistically consistent estimates of PSR parameters can be learned effi…

2013-09-26abs ↗pdf ↗

The paper proves a pseudo-Kähler structure on a torus's projective space.

problem Existence of a pseudo-Kähler structure on a torus's projective space.
method Proved the existence of a pseudo-Kähler structure using complex, symplectic, and Riemannian compatibility.
result Existence of a moment map for the SL(2, R) action over the deformation space.

The Knowledge Gradient policy is improved for MABs by avoiding dominated actions.

problem Weaknesses in KG policy for MABs, including taking dominated actions.
method Proposed variants of KG that avoid taking dominated actions, including an index heuristic.
result New policies perform well over a range of MABs, including those for which index policies are not optimal.

A new pricing controller handles resource constraints to infer target prices effectively.

problem Resource constraints prevent fixed-price inference, leading to support exclusion.
method Formalizes support-exclusion failure, designs a target-aware controller, and uses a realized information clock.
result The controller can certify feasible target bands and log continuous local densities, leading to polynomial rates of inference.

Renormalized circle diffeos with breaks converge to Moebius maps with a symplectic structure.

problem Analyzing the renormalization of circle diffeomorphisms with breaks.
method Proving convergence to invariant piecewise Moebius maps and identifying the renormalization operator with a sub-action of the mapping class group.
result Renormalization identifies with a symplectic form preserved by the mapping class group.

Study identifies change points in piecewise constant reward functions with fixed exploration budget.

problem Locating abrupt changes in piecewise constant reward functions under bandit feedback.
method Fixed exploration budget, piecewise constant bandit problem, lower bounds, near optimal algorithms.
result Established lower bounds and near matching upper bounds for both small and large budgets.

Paper explores using EEG for better speaker identification, even in noisy environments.

problem Speaker identification performance degrades in background noise.
method Uses EEG signals to enhance speaker identification systems, comparing with acoustic features.
result Speaker identification system using only EEG features outperforms one using only acoustic features in high background noise.

IDS integrates physics engines into deep learning for efficient, interpretable system identification.

problem Lack of generalization and interpretability in learning-based models of physical systems.
method Interactive Differentiable Simulation (IDS) that allows efficient, accurate inference of physical properties.
result Automatic task-based robot design and parameter estimation for nonlinear dynamical systems.

New method identifies critical states to improve RL agent explainability and speed.

problem Challenges in RL agent explainability and action selection timing.
method Identify critical states based on action-based variance in Q-function, prioritize exploitation on these states.
result Identified critical states accelerate RL in grid worlds and deep RL tasks.

StakeBench evaluates language understanding by linking comments to market commitments, improving model alignment with real-world outcomes.

problem Existing financial NLP benchmarks measure perceived language rather than market commitments.
method StakeBench uses observable market behavior to supervise models, testing their ability to detect commitments, identify sides, and project odds.
result Models partially recover position-side signals but struggle with later tasks, highlighting structural failures.

CoCAI uses copulas for accurate multivariate time-series forecasting and anomaly detection.

problem Accurate multivariate time-series forecasting and robust anomaly detection.
method Copula-based conformal prediction for multivariate time-series analysis.
result CoCAI provides statistically valid predictive regions and robust anomaly scores.