Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

12.5%25.0%37.5%50.0% · Sep 199319922001200920182026
48 results for Action Elimination

We study an original problem of pure exploration in a strategic bandit model motivated by Monte Carlo Tree Search. It consists in identifying the best action in a game, when the player may sample random outcomes of sequentially chosen pairs of actions. We propose two strategies for the fixed-confidence setting: Maximin…

2016-02-15abs ↗pdf ↗

Paper tackles online learning in large MDPs with low Bellman rank using AVE algorithm.

problem Online learning of MDPs with large state spaces.
method Develops AVE algorithm inspired by OLIVE, using contextual bandit problems and elimination steps.
result Achieves n\sqrt{n}-regret for learning optimal value function in MDPs with function approximation and low Bellman rank.

We initiate the study of multi-stage episodic reinforcement learning under adversarial corruptions in both the rewards and the transition probabilities of the underlying system extending recent results for the special case of stochastic bandits. We provide a framework which modifies the aggressive exploration enjoyed b…

2019-11-20abs ↗pdf ↗

New algorithm reduces regret in stochastic linear bandits with heteroscedastic noise.

problem Optimizing performance in stochastic linear bandits with varying noise levels.
method Variance-adaptive algorithm VAEE with active exploration strategy.
result Achieves simple regret with a nearly harmonic-mean dependent rate.

Partial connections are (singular) differential systems generalizing classical connections on principal bundles, yielding analogous decompositions for manifolds with nonfree group actions. Connection forms are interpreted as maps determining projections of the tangent bundle onto the partial connection; this approach e…

2003-09-14abs ↗pdf ↗

The paper tackles fair sequential decision making with biased linear bandit feedback.

problem Fair sequential decision making with biased linear bandit feedback.
method Phased elimination algorithm to correct unfair evaluations, establishing upper bounds on regret.
result The worst-case regret is smaller than O(κ1/3log(T)1/3T2/3)\mathcal{O}(κ_*^{1/3}\log(T)^{1/3}T^{2/3}).

Study lenient regret and good-action identification in Gaussian process bandits.

problem Optimizing function values above a certain threshold in Gaussian process bandits.
method Study lenient regret notions and introduce algorithms for finding good actions.
result Upper and lower bounds on lenient regret for GP-UCB and elimination algorithms.

A federated learning algorithm tackles unknown contexts in multi-arm bandits.

problem Learning optimal actions in federated multi-arm bandits with unobserved contexts.
method Elimination-based algorithm for linearly parametrized reward functions.
result Proved regret bound for linearly parametrized reward functions.

Neural eliminators reduce unreliable classification by eliminating improbable classes.

problem Unreliable classification due to noise, insufficient data, overlapping distributions, and unclear class definitions.
method Construct eliminators using classifiers with modified error functions, assigning cases to multiple classes instead of one.
result Elimination of improbable classes improves classification accuracy in real-life medical applications.

We establish a new connection between value and policy based reinforcement learning (RL) based on a relationship between softmax temporal value consistency and policy optimality under entropy regularization. Specifically, we show that softmax consistent action values correspond to optimal entropy regularized policy pro…

2017-02-28abs ↗pdf ↗

Improved elimination strategies for adaptive bandit identification reduce sample complexity and computational burden.

problem Inefficient elimination strategies in bandit identification.
method Adaptive elimination methods that update sampling rules frequently and reduce problem size.
result Adaptive elimination methods achieve better sample complexity and computational efficiency.

New algorithm eliminates arms to minimize regret in complex bandit problems.

problem Minimizing regret in combinatorial bandit problems with explicit exploration.
method Introduces a novel arm elimination scheme that partitions arms into three categories and incorporates explicit exploration.
result Achieves near-optimal regret in combinatorial multi-armed and linear contextual bandit problems.

A new ML method predicts long-time-step molecular dynamics, preserving symplectic and time-reversible properties.

problem Limited computational efficiency in long-time-step molecular dynamics simulations.
method Learning data-driven structure-preserving maps to generate long time-step classical dynamics.
result The method eliminates artifacts like lack of energy conservation and loss of equipartition.

A new method for RL with continuous actions improves stability and scalability.

problem Stability and scalability issues in existing RL methods.
method Soft policy gradient with entropy regularization, combined with double sampling for soft Bellman equation.
result Outperforms off-policy prior methods in continuous action RL tasks.

CoverNet predicts urban driving trajectories using diverse sets of possible actions.

problem Multimodal probabilistic trajectory prediction for urban driving.
method Frame trajectory prediction as classification over a diverse set of trajectories; dynamically generate sets based on current state.
result Outperforms state-of-the-art methods on real-world self-driving datasets.

We introduce Neural Choice by Elimination, a new framework that integrates deep neural networks into probabilistic sequential choice models for learning to rank. Given a set of items to chose from, the elimination strategy starts with the whole item set and iteratively eliminates the least worthy item in the remaining …

2016-02-17abs ↗pdf ↗

New BE dimension measure reveals rich RL problems with sample-efficient algorithms.

problem Finding sample-efficient algorithms for complex RL problems.
method Introducing Bellman Eluder (BE) dimension and designing GOLF and OLIVE algorithms.
result GOLF and OLIVE algorithms learn near-optimal policies for low BE dimension problems with polynomial samples.

Approximate models help RL by reducing policy search space.

problem How much does an approximate model help in learning near-optimal policies in RL?
method Study sample complexity in RL with an approximate model, providing an algorithm and a lower bound.
result An approximate model can reduce sample complexity by eliminating sub-optimal actions.

We develop a model to study the role of rationality in economics and biology. The model's agents differ continuously in their ability to make rational choices. The agents' objective is to ensure their individual survival over time or, equivalently, to maximize profits. In equilibrium, however, rational agents who maxim…

2015-07-14abs ↗pdf ↗

We simplify Khovanov homology for torus braids using Gaussian elimination.

problem Computing Khovanov homology for torus braids is complex and computationally intensive.
method Applying Gaussian elimination to reduce the number of generators in the Khovanov chain complex.
result We provide a bound on the number of generators in the whittled complex at fixed homological degree.

AMBER method selects features efficiently using autoencoders and model-based elimination.

problem Efficiently selecting relevant features for classification.
method Greedy backward elimination using a ranker model and autoencoders.
result AMBER outperforms other feature selection methods in classification accuracy.

Logic approach finds real singularities in differential equations.

problem Finding geometric singularities of implicit ODEs over the reals.
method Vessiot theory, parametric Gaussian elimination, heuristic simplification, real quantifier elimination.
result Effective computation of geometric singularities using logic methods.

We develop an approach for feature elimination in statistical learning with kernel machines, based on recursive elimination of features.We present theoretical properties of this method and show that it is uniformly consistent in finding the correct feature space under certain generalized assumptions.We present four cas…

2013-04-18abs ↗pdf ↗

Method models other agents' behaviors without requiring direct observation.

problem Understanding and interacting effectively with other agents in reinforcement learning.
method Extracts representations from local observations of the controlled agent using encoder-decoder architectures.
result The method achieves higher returns than baseline methods in multi-agent environments.

AlphaForgeBench evaluates LLMs as quantitative researchers, not trading agents, to address instability in financial decision-making.

problem Behavioral instability of LLMs in sequential decision-making under financial uncertainty.
method Proposes AlphaForgeBench, a framework that requires LLMs to generate executable alpha factors and compose factor-based trading strategies.
result Eliminates execution-induced instability and provides a rigorous benchmark for evaluating financial reasoning.

A new algorithm improves recommendation systems by considering repeated exposure to actions.

problem Improving recommendation systems by accounting for human memory decay.
method Introducing Weighted Tallying Bandits (WTB) and studying them under Repeated Exposure Optimality (REO).
result A simple modification of the successive elimination algorithm achieves nearly optimal complete policy regret.

Let D be an irreducible lattice in a connected, semisimple Lie group G with finite center. Assume that the real rank of G is at least two, that G/D is not compact, and that G has more than one noncompact simple factor. We show that D has no orientation-preserving actions on the real line. (In algebraic terms, this mean…

2006-04-28abs ↗pdf ↗

Study reduces complexity and uncertainty in human atrial cell models.

problem Uncertainty in parameter estimates from gating kinetics models.
method Approximate Bayesian computation to re-calibrate models, investigate two approaches: more complete datasets and less complex formulations.
result Less complex model with fewer parameters gives better fit and lower uncertainty.

This paper proposes using variational autoencoders to model opponents in multi-agent systems.

problem Understanding and interacting with opponents in multi-agent systems.
method Variational autoencoders for opponent modeling, with a modification to use local information.
result Our opponent modeling methods achieve equal or greater episodic returns.

PoWER-BERT speeds up BERT inference by eliminating redundant word-vectors.

problem Improving BERT inference speed without sacrificing accuracy.
method Eliminating redundant word-vectors using a self-attention-based significance measure and learning the number of vectors to eliminate.
result Up to 4.5x reduction in inference time with <1% loss in accuracy on GLUE benchmark.

GPE algorithm optimizes nonparametric contextual bandits with efficient regret bounds.

problem Optimizing nonparametric contextual bandits with efficient regret bounds.
method Inspired by Policy Elimination, GPE uses oracle-efficient techniques for nonparametric classes with infinite VC-dimension.
result GPE is regret-optimal for policy classes with integrable entropy, and for larger entropy, it provides an ε\varepsilon-greedy algorithm with matching regret bounds.

New RL algorithm reduces policy switching cost to loglog(T) with similar regret.

problem Low policy switching cost in real-life RL applications.
method Stage-wise exploration and adaptive policy elimination.
result Regret of O(HSAloglogT)O(HSA \log\log T) with O(HSAloglogT)O(HSA \log\log T) switching cost.