Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

21416282 · Jun 202019922001200920172026
48 results for rotting rewards

Study tackles infinitely many-armed bandits with rotting rewards, achieving tight regret bounds.

problem Infinitely many-armed bandits with rotting rewards.
method Adaptive sliding window UCB algorithm for slow and abrupt rotting scenarios.
result Achieves tight regret bounds for both slow and abrupt rotting scenarios.

In stochastic multi-armed bandits, the reward distribution of each arm is assumed to be stationary. This assumption is often violated in practice (e.g., in recommendation systems), where the reward of an arm may change whenever is selected, i.e., rested bandit setting. In this paper, we consider the non-parametric rott…

2018-11-27abs ↗pdf ↗

The Multi-Armed Bandits (MAB) framework highlights the tension between acquiring new knowledge (Exploration) and leveraging available knowledge (Exploitation). In the classical MAB problem, a decision maker must choose an arm at each time step, upon which she receives a reward. The decision maker's objective is to maxi…

2017-02-23abs ↗pdf ↗

Graph-Triggered Bandits unify rested and restless bandits with graph-defined arm interactions.

problem Modeling sequential decision-making problems with evolving arm rewards.
method Graph-Triggered Bandits (GTBs) framework that generalizes rested and restless bandits using a graph.
result Rested and restless bandits are special cases of GTBs for suitable graphs.

ROTS improves sentence similarity by incorporating structural information.

problem Measuring sentence similarity with theoretical insights and structural awareness.
method Recursive Optimal Transport (ROT) framework to incorporate structural information.
result ROTS outperforms weakly supervised approaches in sentence similarity tasks.

We investigate Legendrian graphs in (R3,ξstd)(\R^3, ξ_{std}). We extend the classical invariants, Thurston-Bennequin number and rotation number to Legendrian graphs. We prove that a graph can be Legendrian realized with all its cycles Legendrian unknots with tb=1tb=-1 and rot=0rot=0 if and only if it does not contain K4K_4 as a mi…

2011-08-10abs ↗pdf ↗

Enhanced rotation prediction improves SSL models by capturing both shape and texture information.

problem Rotation prediction misses texture information, limiting model performance.
method Introduces image enhanced rotation prediction (IE-Rot) that combines rotation and image enhancement tasks.
result IE-Rot models outperform Rotation on various benchmarks.

Researchers predict butt rot volume using harvester data and remote sensing.

problem Predicting butt rot volume in Norway spruce stands for optimal forest management.
method Used random forest models with harvester information, remote sensing, and environmental data.
result Remotely sensed predictor variables were more important than environmental variables.

Global existence of Willmore flow with boundary via Li-Yau inequality.

problem Global existence of Willmore flow with boundary conditions.
method Extending Li-Yau inequality to surfaces with boundary and using geometric measure theory.
result Global existence of Willmore flow with Dirichlet boundary data below a specific energy threshold.

We prove two results on the classification of trivial Legendrian embeddings g:G(S3,ξstd)g: G \rightarrow (S^3,ξ_{std}) of planar graphs. First, the oriented Legendrian ribbon RgR_g and rotation invariant rotg\text{rot}_g are a complete set of invariants. Second, if GG is 3-connected or contains K4K_4 as a minor, then the unique t…

2016-04-04abs ↗pdf ↗

In this paper, as the second in our series of papers on differential geometry of microlinear Frolicher spaces, we study differenital forms. The principal result is that the exterior differentiation is uniquely determined geometrically, just as grad (ient), div (ergence) and rot (ation) are uniquely determined geometric…

2010-03-23abs ↗pdf ↗

This paper presents a unified framework for smooth convex regularization of discrete optimal transport problems. In this context, the regularized optimal transport turns out to be equivalent to a matrix nearness problem with respect to Bregman divergences. Our framework thus naturally generalizes a previously proposed …

2016-10-20abs ↗pdf ↗

Let MnM_n be the topological moduli space of all parallel n-cables of long framed oriented knots in 3-space. We construct in a combinatorial way for each natural number n>1n>1 a 1-cocycle RnR_n which represents a non trivial class in H1(Mn;Z[x1,x2,...,x11,x21,...])H^1(M_n; \mathbb{Z} [x_1,x_2,...,x_1^{-1},x_2^{-1},...]), where the number of variabl…

2017-09-28abs ↗pdf ↗

Let (Mn,g,f)(M^n,g,\nabla f), n3n\geq 3, be an expanding gradient Ricci soliton with nonnegative sectional curvature whose asymptotic cone is isometric to C(Sn1(c))C(\mathbb{S}^{n-1}(c)) where Sn1(c)\mathbb{S}^{n-1}(c) is the standard (n1)(n-1)-sphere of curvature 1/c21/c^2, with c(0,1)c\in(0,1). We prove that if the convergence to the asympto…

2013-03-14abs ↗pdf ↗

The following three geometrical structures on a manifold are studied in detail: (1) Leibnizian: a non-vanishing 1-form ΩΩ plus a Riemannian metric $\h$ on its annhilator vector bundle. In particular, the possible dimensions of the automorphism group of a Leibnizian G-structure are characterized. (2) Galilean: Leibnizi…

2002-11-08abs ↗pdf ↗

Let ΩΩ be a smooth compact oriented 3-dimensional Riemannian manifold with boundary. A quaternion field is a pair q={α,u}q=\{α,u\} of a function αα and a vector field uu on ΩΩ. A field qq is {\it harmonic} if α,uα, u are continuous in ΩΩ and α=rotu,divu=0\nablaα={\rm rot\,}u,\,{\rm div\,}u=0 holds into ΩΩ. The space ${\mathscr Q…

2019-01-26abs ↗pdf ↗

We present numerical visualizations of Ricci Flow of surfaces and 3-dimensional manifolds of revolution. Ricci_rot is an educational tool which visualizes surfaces of revolution moving under Ricci flow. That these surfaces tend to remain embedded in R3 is what makes direct visualization possible. The numerical lessons …

2004-06-09abs ↗pdf ↗

Reward hacking exploits misspecified rewards, affecting agent capabilities and true performance.

problem Reward hacking in RL models exploiting reward misspecifications.
method Constructed four RL environments with misspecified rewards; analyzed agent capabilities and behavior.
result More capable agents exploit reward misspecifications, achieving higher proxy reward but lower true reward.

Paper introduces PRMs to learn non-Markovian stochastic rewards for reinforcement learning.

problem Lack of structured representation for non-Markovian stochastic rewards in reinforcement learning.
method Introduces probabilistic reward machines (PRMs) and presents an algorithm to learn them from decision processes.
result Algorithm proves correct and convergent for learning PRMs from decision processes.

Self-supervised reward prediction improves RL in sparse reward settings.

problem Data efficiency and sparse reward signals in reinforcement learning.
method Learning a state representation for reward prediction and using it to shape rewards.
result Self-supervised reward prediction enhances RL algorithms in single-goal environments.

The study categorizes reward errors in reinforcement learning, finding some can be beneficial.

problem Training language models with imperfect proxy rewards.
method Theoretical analysis of policy gradient optimization and categorization of reward errors.
result Reward errors can be benign or even beneficial, preventing policy from stalling.

Reward collapse occurs when ranking-based reward models yield uniform rewards for different prompts.

problem Reward collapse in aligning large language models with human preferences.
method Introduced a prompt-aware optimization scheme to derive closed-form expressions for reward distributions.
result Our prompt-aware utility functions significantly alleviate reward collapse during training.

Proposes a method to boost deep reinforcement learning with sparse rewards.

problem Challenges in learning complex behaviors with long horizons and sparse rewards.
method Predictive coding for reward shaping.
result Achieves better learning by providing reward signals that understand environment dynamics and emphasize useful features.

Action guidance helps agents learn true objectives in games with sparse rewards.

problem Training agents in games with sparse rewards requires significant exploration.
method Action guidance, a novel technique that combines exploration with reward shaping.
result Action guidance enables agents to optimize true objectives efficiently.

New RL method uses distance between states instead of rewards for sparse reward environments.

problem Sparse rewards or non-reward environments in reinforcement learning.
method Uses goal-distance gradient and bridge point planning for policy improvement.
result Significantly better performance on sparse reward and local optimal problems in complex environments.

Paper proposes RRD to learn proxy rewards for sparse delayed rewards in episodic reinforcement learning.

problem Learning from sparse and delayed rewards in reinforcement learning.
method Randomized Return Decomposition (RRD) algorithm to redistribute rewards.
result Substantial improvement over baseline algorithms in experiments.

Learning reward functions from data is a promising path towards achieving scalable Reinforcement Learning (RL) for robotics. However, a major challenge in training agents from learned reward models is that the agent can learn to exploit errors in the reward model to achieve high reward behaviors that do not correspond …

2019-11-01abs ↗pdf ↗

Enhances reward specification in RL with a novel language-based approach.

problem Reward specification in RL can lead to unintended, potentially harmful behaviours.
method Developed a novel class of language-based Reward Machines using RML's built-in memory.
result Can specify non-regular, non-Markovian reward functions for complex tasks.

Reward tweaking optimizes behavior for long-term goals by adjusting the reward function.

problem Optimizing behavior for long-term goals in reinforcement learning with unstable long planning horizons.
method Reward tweaking learns a surrogate reward function that induces optimal behavior for the original task.
result Reward tweaking guides agents towards better long-term returns while planning for short horizons.

We propose a generic, Bayesian, information geometric approach to the exploration--exploitation trade-off in multi-armed bandit problems. Our approach, BelMan, uniformly supports pure exploration, exploration--exploitation, and two-phase bandit problems. The knowledge on bandit arms and their reward distributions is su…

2018-05-04abs ↗pdf ↗

Extends reinforcement learning alignment to scalar rewards, improving math reasoning.

problem Designing reinforcement learning algorithms for general LLM alignment.
method Introduces f-GRPO and f-HAL, estimating f-divergences between reward-aligned and unaligned distributions.
result Improves math reasoning RLVR tasks and mitigates reward hacking.

This work characterizes reward function partial identifiability and its impact on policy optimization.

problem Reward function partial identifiability in complex tasks.
method Formal characterisation of partial identifiability using various reward learning data sources.
result Unified framework for comparing data sources and downstream tasks by their invariances.