Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

265379105 · Jun 202019922001200920182026
48 results for pseudo rewards

PHE adds pseudo-rewards to history to minimize regret in stochastic bandits.

problem Minimizing cumulative regret in stochastic multi-armed bandits.
method PHE algorithm that adds O(t)O(t) i.i.d. pseudo-rewards to history and pulls the best arm based on the perturbed history.
result Near-optimal regret bounds derived for PHE.

A new bandit algorithm Giro learns to explore by sampling its history with pseudo rewards.

problem Learning to explore in multi-armed bandits with limited feedback.
method Giro uses a non-parametric bootstrap of its history with pseudo rewards to guide exploration.
result Giro achieves a regret bound of O(KΔ1logn)O(K Δ^{-1} \log n), demonstrating effective exploration.

We consider an agent's uncertainty about its environment and the problem of generalizing this uncertainty across observations. Specifically, we focus on the problem of exploration in non-tabular reinforcement learning. Drawing inspiration from the intrinsic motivation literature, we use density models to measure uncert…

2016-06-06abs ↗pdf ↗

A fair policy for hiring candidates from different groups is proposed in a linear contextual bandit problem.

problem Selecting candidates from different sensitive groups in a fair manner.
method A greedy policy that constructs a ridge regression estimate and computes relative rank using empirical cumulative distribution function.
result The greedy policy achieves fair pseudo-regret of order dT\sqrt{dT} after TT rounds, satisfying demographic parity.

Paper shows RL's policy gradient methods can be viewed as supervised learning problems.

problem Connecting RL and SL approaches.
method Proves PGM as supervised learning problem, replaces true labels with discounted rewards.
result PGM can be seen as supervised learning, emphasizing similarities.

New algorithm for RL with horizon-free reward-free exploration for linear MDPs.

problem Reward-free reinforcement learning with long planning horizons.
method Uncertainty-weighted value-targeted regression with exploration-driven pseudo-reward and moment estimator.
result Horizon-free sample complexity of O(d2ε2)O(d^2\varepsilon^{-2}) for finding an ε\varepsilon-optimal policy.

Study shows fast rates for inverse reinforcement learning with linear rewards.

problem Entropy-regularized min-max inverse reinforcement learning in finite-horizon MDPs.
method Structural and statistical analysis of Min-Max-IRL with pseudo-self-concordance.
result Both trajectory-level KL divergence and parameter error decay at O(n1)\mathcal{O}(n^{-1}).

Successor Options discovers reusable skills using landmark states.

problem Discovering reusable skills in reinforcement learning.
method Leverages Successor Representations to build a state space model and learns intra-option policies using a novel pseudo-reward.
result Demonstrates the approach's efficacy on grid-worlds and high-dimensional robotic control environments.

We study, to the best of our knowledge, the first Bayesian algorithm for unimodal Multi-Armed Bandit (MAB) problems with graph structure. In this setting, each arm corresponds to a node of a graph and each edge provides a relationship, unknown to the learner, between two nodes in terms of expected reward. Furthermore, …

2016-11-17abs ↗pdf ↗

Improves neural networks' ability to learn new tasks without forgetting earlier ones.

problem Preventing catastrophic forgetting in neural networks.
method Introducing a second discriminator in the GAN to generate important features for task retention.
result Significant reduction in catastrophic forgetting compared to standard methods.

A novel bandit problem with delayed arms, showing optimal strategies and lower bounds.

problem Optimizing reward in a stochastic multi-armed bandit setting with delayed arms.
method Mapping to PINWHEEL scheduling problem, simple greedy algorithm, UCB based algorithm, lower bounds.
result Simple greedy algorithm is asymptotically (11/e)(1-1/e) optimal and UCB based algorithm has clogT+o(logT)c \log T + o(\log T) cumulative regret.

A method for continual learning using world models in reinforcement learning.

problem Catastrophic forgetting in lifelong learning with neural networks.
method Interleaving internally generated episodes of past experiences (pseudo-rehearsal) with external environment's observations.
result Consistent reduction in temporal prediction loss compared to non-interleaved learning.

New algorithm robust to probabilistic unbounded adversarial attacks in bandit problems.

problem Powerful adversaries that can catastrophically perturb the revealed reward in bandit problems.
method Proposes med-E-UCB and med-εε-greedy algorithms based on sample median for robustness.
result Achieves O(logT)\mathcal{O}(\log T) pseudo-regret under arbitrary and unbounded reward perturbation.

SUPE combines unlabeled data with RL to efficiently explore tasks.

problem Efficient exploration in reinforcement learning with sparse rewards.
method Extract low-level skills using VAE, pseudo-label unlabeled data, and use as off-policy data for online RL.
result SUPE outperforms prior methods across 42 long-horizon tasks.

Paper proposes a deep RL method for hedging variable annuities, outperforming misspecified models.

problem Model miscalibration in variable annuity contracts with GMMB and GMDB riders.
method Two-phase deep reinforcement learning approach: training phase in a controlled environment, online learning phase in real market.
result Trained reinforcement learning agent hedges equally well as correct Delta in training phase and outperforms misspecified Deltas.

In this work, we study the pseudo-Riemannian submanifolds of a pseudo-sphere with 1-type pseudo-spherical Gauss map. First, we classify the Lorentzian surfaces in a 4-dimensional pseudo-sphere Ss4(1)\mathbb{S}^4_s(1) with index s, s=1,2s=1, 2, and having harmonic pseudo-spherical Gauss map. Then we give a characterization the…

2015-10-28abs ↗pdf ↗

Extends tangle theory to include undetermined crossings in periodic structures.

problem Classical tangle theory's limitations in handling undetermined crossings.
method Introduces pseudo DP tangles, defined as liftings of pseudo motifs in the thickened torus, and analyzes them through diagrammatic methods.
result Defines equivalence for pseudo DP tangles and proves an analogue of Reidemeister theorem.

Paper proposes DG-ETC for online submodular maximization with stochastic bandit feedback.

problem Online unconstrained submodular maximization with stochastic bandit feedback.
method Double-Greedy - Explore-then-Commit (DG-ETC) approach.
result DG-ETC achieves logarithmic regret O(dlog(dT))O(d\log(dT)) for 1/21/2-approximate pseudo-regret.

The paper studies prolongations of Lie algebras associated with pseudo HH-type Lie algebras.

problem Investigating prolongations of Lie algebras associated with pseudo HH-type Lie algebras.
method Analyzing prolongations of associated fundamental graded Lie algebra and associated conformal pseudo-subriemannian fundamental graded Lie algebra.
result The prolongation of the associated conformal pseudo-subriemannian fundamental graded Lie algebra coincides with that of the associated fundamental graded Lie algebra under certain conditions.

In this paper, we derived biharmonic equations for pseudo-Riemannian submanifolds of pseudo-Riemannian manifolds which includes the biharmonic equations for submanifolds of Riemannian manifolds as a special case. As applications, we proved that a pseudo-umbilical biharmonic pseudo-Riemannian submanifold of a pseudo-Rie…

2015-12-08abs ↗pdf ↗

The paper studies pseudo links in genus g handlebodies, generalizing knot theory.

problem Modeling DNA knots with missing crossing information.
method Introducing pseudo links as mixed pseudo links in S^3, generalizing Kauffman bracket polynomial and Alexander theorem.
result The theory of pseudo links is closely related to singular links and can be applied to study singular links in genus g handlebodies.

Paper constructs a HOMFLYPT-type invariant for pseudo links.

problem Inability to construct polynomial invariants for pseudo links using Hecke algebra techniques.
method Using a resolution homomorphism and pseudo Hecke algebra of type \(A\), the paper constructs a HOMFLYPT-type invariant for oriented pseudo links.
result The constructed invariant satisfies a natural pseudo skein relation and admits a state-sum formulation.

The paper explores knotoids, pseudo knotoids, braidoids, and pseudo braidoids on the torus.

problem The study of knotoids, pseudo knotoids, braidoids, and pseudo braidoids on the torus.
method Introducing new knotoid and braidoid concepts, isotopy theorems, state sum formulas, and Alexander and Markov theorems.
result Formulation and proof of Alexander and Markov theorems for mixed knotoids and mixed pseudo knotoids.

The paper characterizes metallic pseudo-Riemannian manifolds using conjugate connections and tensor structures.

problem Characterizing metallic pseudo-Riemannian manifolds.
method Using conjugate connections and tensor structures, the paper derives new characterizations and conditions for these manifolds.
result A necessary and sufficient condition for a non-integrable metallic pseudo-Riemannian manifold to be a quasi metallic pseudo-Riemannian manifold is derived.

The paper proves Liouville theorems for holomorphic maps on pseudo-Hermitian manifolds.

problem Proving Liouville theorems for holomorphic maps on pseudo-Hermitian manifolds.
method Analyzing maps between different classes of pseudo-Hermitian manifolds, using curvature assumptions.
result Holomorphic maps are constant under certain curvature conditions.

Reward hacking exploits misspecified rewards, affecting agent capabilities and true performance.

problem Reward hacking in RL models exploiting reward misspecifications.
method Constructed four RL environments with misspecified rewards; analyzed agent capabilities and behavior.
result More capable agents exploit reward misspecifications, achieving higher proxy reward but lower true reward.

The paper aims to initiate a systematic study of conformal mappings between Finsler spacetimes and, more generally, between pseudo-Finsler spaces. This is done by extending several results in pseudo-Riemannian geometry which are necessary for field-theoretical applications and by proposing a technique which reduces a s…

2017-06-06abs ↗pdf ↗

Study on pseudo-Riemannian algebraic Ricci solitons in 4D Lie groups.

problem Investigating conditions for pseudo-Riemannian algebraic Ricci solitons on 4D Lie algebras.
method Analyzing the algebraic Ricci soliton equation for each 4D Lie algebra.
result Complete description of pseudo-Riemannian algebraic Ricci solitons in dimension four.

Paper introduces PRMs to learn non-Markovian stochastic rewards for reinforcement learning.

problem Lack of structured representation for non-Markovian stochastic rewards in reinforcement learning.
method Introduces probabilistic reward machines (PRMs) and presents an algorithm to learn them from decision processes.
result Algorithm proves correct and convergent for learning PRMs from decision processes.