Study tackles infinitely many-armed bandits with rotting rewards, achieving tight regret bounds.
problem Infinitely many-armed bandits with rotting rewards.
method Adaptive sliding window UCB algorithm for slow and abrupt rotting scenarios.
result Achieves tight regret bounds for both slow and abrupt rotting scenarios.
New algorithm reduces regret in infinitely many-armed bandits with decreasing rewards.
problem Infinitely many-armed bandits with rotting rewards.
method UCB index and adaptive threshold for unknown rotting rate, UCB index alone for known rotting rate.
result Matching upper bounds on regret achieved for different scenarios.
New algorithm handles both decaying and non-decaying bandit problems.
problem Decaying rewards in bandit problems.
method RAW-UCB algorithm for both rested and restless rotting bandits.
result Achieves near-optimal regret in both rested and restless rotting bandits.
In stochastic multi-armed bandits, the reward distribution of each arm is assumed to be stationary. This assumption is often violated in practice (e.g., in recommendation systems), where the reward of an arm may change whenever is selected, i.e., rested bandit setting. In this paper, we consider the non-parametric rott…
The Multi-Armed Bandits (MAB) framework highlights the tension between acquiring new knowledge (Exploration) and leveraging available knowledge (Exploitation). In the classical MAB problem, a decision maker must choose an arm at each time step, upon which she receives a reward. The decision maker's objective is to maxi…
New method tracks shifts in infinite-armed bandits without prior knowledge.
problem Tracking shifts in non-stationary infinite-armed bandits.
method Blackbox conversion of finite-armed MAB to infinite-armed non-stationary, randomized elimination.
result First parameter-free optimal regret bounds for all reservoir regularity regimes.
Graph-Triggered Bandits unify rested and restless bandits with graph-defined arm interactions.
problem Modeling sequential decision-making problems with evolving arm rewards.
method Graph-Triggered Bandits (GTBs) framework that generalizes rested and restless bandits using a graph.
result Rested and restless bandits are special cases of GTBs for suitable graphs.
TURB-Rot provides a large database of turbulent rotating flow snapshots for research.
problem Lack of large-scale, high-resolution datasets for turbulent rotating flows.
method Direct Numerical Simulations of Navier-Stokes equations with rotation.
result Provides a diverse set of 300K complex images and fields for testing.
ROTS improves sentence similarity by incorporating structural information.
problem Measuring sentence similarity with theoretical insights and structural awareness.
method Recursive Optimal Transport (ROT) framework to incorporate structural information.
result ROTS outperforms weakly supervised approaches in sentence similarity tasks.
Thomas Rot discusses a topology theorem at a math conference.
problem Topology theorem presentation at a math conference.
method No specific method mentioned; presentation of a theorem.
result No specific key result mentioned; presentation of a theorem.
We investigate Legendrian graphs in (R3,ξstd). We extend the classical invariants, Thurston-Bennequin number and rotation number to Legendrian graphs. We prove that a graph can be Legendrian realized with all its cycles Legendrian unknots with tb=−1 and rot=0 if and only if it does not contain K4 as a mi…
New R-equivalence classes found for torus knot diagrams.
problem Classifying colorings of torus knots.
method Introducing R-equivalence relation on quandle colorings.
result Determined R-equivalence classes for RotE2-colorings of torus knots. Enhanced rotation prediction improves SSL models by capturing both shape and texture information.
problem Rotation prediction misses texture information, limiting model performance.
method Introduces image enhanced rotation prediction (IE-Rot) that combines rotation and image enhancement tasks.
result IE-Rot models outperform Rotation on various benchmarks.
Develops a novel global pooling framework using optimal transport.
problem Sub-optimal performance in global pooling operations.
method Regularized Optimal Transport (ROT) for generalized global pooling.
result ROTP layers can improve performance in various machine learning scenarios.
Researchers predict butt rot volume using harvester data and remote sensing.
problem Predicting butt rot volume in Norway spruce stands for optimal forest management.
method Used random forest models with harvester information, remote sensing, and environmental data.
result Remotely sensed predictor variables were more important than environmental variables.
Global existence of Willmore flow with boundary via Li-Yau inequality.
problem Global existence of Willmore flow with boundary conditions.
method Extending Li-Yau inequality to surfaces with boundary and using geometric measure theory.
result Global existence of Willmore flow with Dirichlet boundary data below a specific energy threshold.
We prove two results on the classification of trivial Legendrian embeddings g:G→(S3,ξstd) of planar graphs. First, the oriented Legendrian ribbon Rg and rotation invariant rotg are a complete set of invariants. Second, if G is 3-connected or contains K4 as a minor, then the unique t…
In this paper, as the second in our series of papers on differential geometry of microlinear Frolicher spaces, we study differenital forms. The principal result is that the exterior differentiation is uniquely determined geometrically, just as grad (ient), div (ergence) and rot (ation) are uniquely determined geometric…
This paper presents a unified framework for smooth convex regularization of discrete optimal transport problems. In this context, the regularized optimal transport turns out to be equivalent to a matrix nearness problem with respect to Bregman divergences. Our framework thus naturally generalizes a previously proposed …
Let Mn be the topological moduli space of all parallel n-cables of long framed oriented knots in 3-space. We construct in a combinatorial way for each natural number n>1 a 1-cocycle Rn which represents a non trivial class in H1(Mn;Z[x1,x2,...,x1−1,x2−1,...]), where the number of variabl…
We advocate an optimization-centric view on and introduce a novel generalization of Bayesian inference. Our inspiration is the representation of Bayes' rule as infinite-dimensional optimization problem (Csiszar, 1975; Donsker and Varadhan; 1975, Zellner; 1988). First, we use it to prove an optimality result of standard…
Let (Mn,g,∇f), n≥3, be an expanding gradient Ricci soliton with nonnegative sectional curvature whose asymptotic cone is isometric to C(Sn−1(c)) where Sn−1(c) is the standard (n−1)-sphere of curvature 1/c2, with c∈(0,1). We prove that if the convergence to the asympto…
New examples of non-simple knots in Lens spaces show rich botany.
problem Understanding the variety of Legendrian and transversal knots in Lens spaces.
method Presented new families of non-simple knots in tight Lens spaces.
result More non-isotopic Legendrians topologically isotopic to the n-twist knot in Lens spaces than in S3. New Legendrian bounds for non-fibered knots in 3-manifolds.
problem Understanding Legendrian representatives of non-fibered knots.
method Analyzing Thurston-Bennequin bounds and contact invariants.
result Non-fibered knots have Legendrian representatives with tb=0. The paper proves n-transitivity for equivariant diffeomorphisms of manifolds.
problem Proving n-transitivity for equivariant diffeomorphisms. method Analyzing the group of equivariant diffeomorphisms on proper smooth G-manifolds. result The group of equivariant diffeomorphisms acts n-transitively on M. The following three geometrical structures on a manifold are studied in detail: (1) Leibnizian: a non-vanishing 1-form Ω plus a Riemannian metric $\h$ on its annhilator vector bundle. In particular, the possible dimensions of the automorphism group of a Leibnizian G-structure are characterized. (2) Galilean: Leibnizi…
Let Ω be a smooth compact oriented 3-dimensional Riemannian manifold with boundary. A quaternion field is a pair q={α,u} of a function α and a vector field u on Ω. A field q is {\it harmonic} if α,u are continuous in Ω and ∇α=rotu,divu=0 holds into Ω. The space ${\mathscr Q…
We classify Legendrian unknots in overtwisted contact structures on S3. In particular, we show that up to contact isotopy for every pair (n,±(n−1)) with n>0 there are exactly two oriented non-loose Legendrian unknots in S3 with Thurston-Bennequin invariant n and rotation number ±(n−1). (Only one overt…
We propose a novel visual context-aware filter generation module which incorporates contextual information present in images into Convolutional Neural Networks (CNNs). In contrast to traditional CNNs, we do not employ the same set of learned convolution filters for all input image instances. Our proposed input-conditio…
We present numerical visualizations of Ricci Flow of surfaces and 3-dimensional manifolds of revolution. Ricci_rot is an educational tool which visualizes surfaces of revolution moving under Ricci flow. That these surfaces tend to remain embedded in R3 is what makes direct visualization possible. The numerical lessons …
Reward hacking exploits misspecified rewards, affecting agent capabilities and true performance.
problem Reward hacking in RL models exploiting reward misspecifications.
method Constructed four RL environments with misspecified rewards; analyzed agent capabilities and behavior.
result More capable agents exploit reward misspecifications, achieving higher proxy reward but lower true reward.
Paper introduces PRMs to learn non-Markovian stochastic rewards for reinforcement learning.
problem Lack of structured representation for non-Markovian stochastic rewards in reinforcement learning.
method Introduces probabilistic reward machines (PRMs) and presents an algorithm to learn them from decision processes.
result Algorithm proves correct and convergent for learning PRMs from decision processes.
This work analyzes the value of future reward information in RL.
problem Analyzing the impact of knowing future rewards in reinforcement learning.
method Competitive analysis and worst-case reward distribution.
result Exact ratios between standard RL agents and those with future-reward lookahead.
Self-supervised reward prediction improves RL in sparse reward settings.
problem Data efficiency and sparse reward signals in reinforcement learning.
method Learning a state representation for reward prediction and using it to shape rewards.
result Self-supervised reward prediction enhances RL algorithms in single-goal environments.
The study categorizes reward errors in reinforcement learning, finding some can be beneficial.
problem Training language models with imperfect proxy rewards.
method Theoretical analysis of policy gradient optimization and categorization of reward errors.
result Reward errors can be benign or even beneficial, preventing policy from stalling.
Reward collapse occurs when ranking-based reward models yield uniform rewards for different prompts.
problem Reward collapse in aligning large language models with human preferences.
method Introduced a prompt-aware optimization scheme to derive closed-form expressions for reward distributions.
result Our prompt-aware utility functions significantly alleviate reward collapse during training.
Reward models need more than just accuracy for effective RLHF.
problem The effectiveness of reward models in RLHF is not fully understood.
method An optimization perspective to evaluate reward models.
result Reward models with low reward variance can lead to a flat optimization landscape, hindering performance.
Proposes a method to boost deep reinforcement learning with sparse rewards.
problem Challenges in learning complex behaviors with long horizons and sparse rewards.
method Predictive coding for reward shaping.
result Achieves better learning by providing reward signals that understand environment dynamics and emphasize useful features.
Action guidance helps agents learn true objectives in games with sparse rewards.
problem Training agents in games with sparse rewards requires significant exploration.
method Action guidance, a novel technique that combines exploration with reward shaping.
result Action guidance enables agents to optimize true objectives efficiently.
New RL method uses distance between states instead of rewards for sparse reward environments.
problem Sparse rewards or non-reward environments in reinforcement learning.
method Uses goal-distance gradient and bridge point planning for policy improvement.
result Significantly better performance on sparse reward and local optimal problems in complex environments.
Paper proposes RRD to learn proxy rewards for sparse delayed rewards in episodic reinforcement learning.
problem Learning from sparse and delayed rewards in reinforcement learning.
method Randomized Return Decomposition (RRD) algorithm to redistribute rewards.
result Substantial improvement over baseline algorithms in experiments.
Learning reward functions from data is a promising path towards achieving scalable Reinforcement Learning (RL) for robotics. However, a major challenge in training agents from learned reward models is that the agent can learn to exploit errors in the reward model to achieve high reward behaviors that do not correspond …
Enhances reward specification in RL with a novel language-based approach.
problem Reward specification in RL can lead to unintended, potentially harmful behaviours.
method Developed a novel class of language-based Reward Machines using RML's built-in memory.
result Can specify non-regular, non-Markovian reward functions for complex tasks.
Reward tweaking optimizes behavior for long-term goals by adjusting the reward function.
problem Optimizing behavior for long-term goals in reinforcement learning with unstable long planning horizons.
method Reward tweaking learns a surrogate reward function that induces optimal behavior for the original task.
result Reward tweaking guides agents towards better long-term returns while planning for short horizons.
We propose a generic, Bayesian, information geometric approach to the exploration--exploitation trade-off in multi-armed bandit problems. Our approach, BelMan, uniformly supports pure exploration, exploration--exploitation, and two-phase bandit problems. The knowledge on bandit arms and their reward distributions is su…
Many reinforcement-learning researchers treat the reward function as a part of the environment, meaning that the agent can only know the reward of a state if it encounters that state in a trial run. However, we argue that this is an unnecessary limitation and instead, the reward function should be provided to the learn…
Extends reinforcement learning alignment to scalar rewards, improving math reasoning.
problem Designing reinforcement learning algorithms for general LLM alignment.
method Introduces f-GRPO and f-HAL, estimating f-divergences between reward-aligned and unaligned distributions.
result Improves math reasoning RLVR tasks and mitigates reward hacking.
This work characterizes reward function partial identifiability and its impact on policy optimization.
problem Reward function partial identifiability in complex tasks.
method Formal characterisation of partial identifiability using various reward learning data sources.
result Unified framework for comparing data sources and downstream tasks by their invariances.