Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,982 papers · 148 categories

Trend · papers per month

1.3%2.6%3.9%5.2% · Feb 202519922001200920172026
48 results for intermediate rewards

Natural language instructions improve reinforcement learning efficiency.

problem Designing effective reward functions for reinforcement learning is difficult and time-consuming.
method Proposes LanguagE-Action Reward Network (LEARN) to map natural language instructions to intermediate rewards.
result Language-based rewards lead to successful task completion 60% more often than without language.

A new recommendation system model tackles unreliable user behavior.

problem Creating effective recommendation systems in the presence of unreliable user behavior.
method A novel modification of Multi-Armed Bandits with an unreliable intermediate.
result Proved fundamental theorems and developed an Explore-Commit algorithm close to optimal performance.

MFMs enable efficient reward alignment for generative models.

problem Computational bottleneck in controlling generative models.
method Meta Flow Maps (MFMs) extend consistency models and flow maps to stochastic regime for efficient value function estimation.
result MFMs enable inference-time steering and unbiased, off-policy fine-tuning to general rewards efficiently.

A robot learns to classify images with limited perception using a layered reinforcement learning approach.

problem Image classification for robots with partial perception.
method Three-layer architecture using deep reinforcement learning, including meta-layer, action-layer, and classification-layer.
result The method achieves high accuracy on the MNIST dataset and provides explainability of the agent's decision-making process.

Develops a new model for RLHF accounting for partially observed states and intermediate feedback.

problem Lack of models for partially observed states and intermediate feedback in RLHF.
method PORRL model with cardinal and dueling feedback methods.
result Demonstrates improved learning and alignment with new model-based and model-free methods.

This work shows how approximate reward models can significantly improve inference-time scaling.

problem Improving the efficiency of inference for large language models.
method Identifying the Bellman error of approximate reward models and using Sequential Monte Carlo (SMC) for inference.
result Approximate reward models can reduce computational complexity from exponential to polynomial in TT.

Improved text generation using transferable rewards from related tasks.

problem Non-differentiable task-specific scores limit the use of policy gradient methods in text generation.
method Transferable Reward Learner that uses model-based rewards for sentence-level and phrase-level similarity.
result Improved performance on semantic evaluation measures in image captioning tasks.

Mathematical Reinforcement Learning faces a 'Two-Hump' problem due to sparse rewards and a scarcity of intermediate 'hard-but-solvable' instances.

problem Mathematical search problems in Reinforcement Learning
method Novel data generation techniques and algorithmic enhancements
result Substantial performance improvements over previous baselines

Paper develops methods to optimize policies directly from human feedback without reward inference.

problem Challenges in RLHF, including reward model overfitting and distribution shift.
method Develops two algorithms for RLHF without reward inference, using zeroth-order gradient approximators.
result Establishes polynomial convergence rates and outperforms existing methods in numerical experiments.

We introduce a simple extension of the minority game in which the market rewards contrarian (resp. trend-following) strategies when it is far from (resp. close to) efficiency. The model displays a smooth crossover from a regime where contrarians dominate to one where trend-followers dominate. In the intermediate phase,…

2004-03-26abs ↗pdf ↗

New algorithms ensure policies perform at least as good as a baseline in reinforcement learning.

problem Learning policies that are guaranteed to perform at least as well as a baseline in reinforcement learning.
method Introduce conservative exploration for average reward and finite horizon problems, presenting two optimistic algorithms.
result Guaranteed performance of policies at least as good as a baseline, without hindering learning ability.

Training-free method improves large language model sequence quality via reward-guided sampling.

problem Optimizing large language model sequence quality over token likelihood.
method Reward-augmented target distribution combined with Sequential Monte Carlo sampling.
result Significant gains in sequence generation and mathematical reasoning tasks.

We consider the classical problem of sequential resource allocation where a decision maker must repeatedly divide a budget between several resources, each with diminishing returns. This can be recast as a specific stochastic optimization problem where the objective is to maximize the cumulative reward, or equivalently …

2019-02-12abs ↗pdf ↗

Quantum model outperforms classical in training but underperforms in real-world metrics.

problem Mismatch between proxy reward signals and true investment objectives in financial domains.
method Hybrid quantum-classical reinforcement learning framework with automated feature engineering.
result Quantum models achieve higher training rewards but underperform in real-world metrics.

New RLHF algorithm identifies optimal policies from human feedback without explicit reward inference.

problem Training large language models with human feedback without reward inference.
method Model-free RLHF algorithm BSAD\mathsf{BSAD} that identifies optimal policies directly from human preference.
result Provable, instance-dependent sample complexity ildeO(cMSA3H3Mlog1δ) ilde{\mathcal{O}}(c_{\mathcal{M}}SA^3H^3M\log\frac{1}δ).

This paper improves sample efficiency for off-policy evaluation with preference data.

problem Improving sample efficiency for off-policy evaluation with preference data.
method Using a deep neural network to learn the value function and leveraging manifold structure.
result Established a provably efficient guarantee for off-policy evaluation with RLHF.

New method optimizes diffusion models without fine-tuning, integrating soft value functions.

problem Optimizing natural design spaces of images, molecules, DNA, RNA, and protein sequences.
method Iterative sampling method integrating soft value functions into diffusion model inference.
result Directly utilizes non-differentiable features/reward feedback, applies to discrete diffusion models.

Tutorial on optimizing diffusion model samples for specific metrics.

problem Optimizing diffusion model samples for specific downstream metrics.
method Review and exploration of inference-time guidance and alignment methods.
result Unified perspective on inference-time algorithms and novel methods.

Paper introduces Latent-CLIP for efficient text-image comparison in latent space.

problem Efficiently compare text and images in latent space without costly decoding.
method Trains CLIP model in latent space, uses Latent-CLIP rewards for noise optimization, and guides generation away from harmful content.
result Latent-CLIP matches CLIP performance on text-image classification and harmful content detection.

Transformers learn sparse Boolean functions through RL and SFT, revealing distinct learning behaviors.

problem Learning sparse Boolean functions with Transformers.
method Reinforcement Learning (RL) with process rewards and Supervised Fine-Tuning (SFT).
result RL learns the whole CoT chain simultaneously, while SFT learns step by step.

This paper generalizes the envelope of mid-lines to intermediate lines for a plane curve.

problem Understanding the envelope of intermediate lines for a plane curve.
method Using singularity theory techniques to analyze the local behavior of the envelope of intermediate lines.
result The envelope of intermediate lines (EILEIL) is formed by three disconnected sets: AEIL, the curve itself, and IPTL.

Entropy regularization improves policy optimization in reinforcement learning.

problem Improving policy optimization in reinforcement learning.
method Entropy regularization is introduced to soften the greedy policy towards a more diverse softmax policy, leading to a continuously parameterized algorithm that interpolates between policy gradient and Q-learning.
result An intermediate algorithm can improve performance in reinforcement learning.

The study connects manifold topology to metrics with positive intermediate curvature.

problem Understanding the relationship between manifold topology and metrics with positive intermediate curvature.
method Formulated a conjecture and proved it for specific dimensions and conditions.
result Closed, aspherical 6-manifolds cannot admit metrics with positive 4-intermediate curvature.

A new algorithm learns optimal personalized treatment plans online with low regret.

problem Learning optimal dynamic treatment regimes in an online setting.
method Developed a novel algorithm balancing exploration and exploitation for rate-optimal regret.
result Guaranteed rate-optimal regret for linear transition and reward models.

New rigidity results for manifolds with maximal symmetry rank and positive intermediate Ricci curvature.

problem Understanding the structure of manifolds with maximal symmetry rank and positive intermediate Ricci curvature.
method Recovering stronger topological rigidity results using higher intermediate Ricci curvatures and nontrivial fundamental groups.
result Stronger topological rigidity results for manifolds with maximal symmetry rank and positive intermediate Ricci curvature.

This work proposes a RL approach to learn versatile robotic manipulation tasks.

problem Challenging manipulation tasks in robotics and vision.
method Reinforcement learning (RL) to combine primitive skills, no intermediate rewards, few demonstrations, and efficient skill learning.
result Versatile robotic manipulation in challenging settings with temporary occlusions and dynamic scene changes.

Study on spaces of metrics with intermediate curvature bounds.

problem Understanding spaces of metrics with lower bounds on intermediate curvatures.
method Analyzing spaces of Riemannian metrics with specific curvature bounds on high-dimensional Spin-manifolds.
result Spaces of metrics with positive p-curvature and k-positive Ricci curvature have non-trivial homotopy groups.

Extends Perelman's theorem to positive intermediate curvature conditions.

problem Positive intermediate curvature conditions and their implications.
method Generalization of Perelman's gluing theorem to positive intermediate curvature conditions.
result Observer moduli space can have non-trivial higher homotopy groups.

The study proves that certain manifolds with boundary cannot have metrics with positive intermediate curvatures.

problem Proving the nonexistence of metrics with positive intermediate curvatures on manifolds with boundary.
method Curvature obstruction theorems for manifolds with boundary.
result Topologically nontrivial compact manifolds with boundary cannot have metrics of positive mm-intermediate curvature if the boundary is mm-convex.

Study finds metrics with positive intermediate Ricci curvature on specific low-dimensional manifolds.

problem Existence of invariant metrics with positive intermediate Ricci curvature on low-dimensional cohomogeneity one manifolds.
method Construction of invariant metrics with positive intermediate Ricci curvature on specific manifolds.
result Invariant metrics with positive 4th-intermediate Ricci curvature exist but not for 3rd-intermediate Ricci curvature on certain manifolds.

Sharp dimension constraints for positive intermediate curvature metrics are established.

problem Proving sharp dimension constraints for metrics with positive intermediate curvature.
method Constructing counterexamples and extending rigidity results.
result Sharp dimension constraints for positive intermediate curvature metrics are established.

This paper introduces a novel approach to measuring privacy risks in deep computer vision models based on intermediate outputs.

problem The exposure of intermediate results in hidden layers of deep computer vision models poses significant privacy concerns.
method The approach leverages Degrees of Freedom (DoF) to evaluate the amount of information retained in each layer and combines this with the rank of the Jacobian matrix to assess sensitivity to input variations.
result The proposed framework provides deeper insights into privacy risks associated with intermediate representations without requiring adversarial attack simulations.

New method evaluates AI stock prediction systems based on decision-making processes.

problem Lack of evaluation for AI systems' decision-making processes.
method Scores intermediate decision process using large language models and closed-loop reinforcement learning feedback.
result Composite behavioral score correlates with Sharpe ratio and reduces prediction error.

The paper proves manifold splitting theorems with nonnegative intermediate curvature.

problem Proving rigidity results for manifolds with nonnegative intermediate curvatures.
method New recursion theorem for spectral intermediate curvatures and cylindrical splitting theorems.
result Smooth metrics with uniformly positive intermediate curvature constructed.

A framework to explain decoder-only sequence classification models using intermediate predictions.

problem Explaining predictions of decoder-only sequence classification models.
method Progressive Inference framework with Single Pass-Progressive Inference and Multi Pass-Progressive Inference methods.
result Significantly better attributions compared to prior work on text classification tasks.

Polyhedral semantics for intermediate logics; Nerve Criterion ensures completeness.

problem Characterize polyhedrally-complete intermediate logics.
method Developed Nerve Criterion to characterize polyhedrally-complete logics combinatorially.
result Nerve Criterion provides a necessary and sufficient condition for polyhedrally-completeness.

We use a local argument to prove if an rr-dimensional torus acts isometrically and effectively on a connected nn-dimensional manifold which has positive kthk^\mathrm{th}-intermediate Ricci curvature at some point, then rn+k2r \leq \lfloor \frac{n+k}{2} \rfloor. This symmetry rank bound generalizes those established by Gr…

2019-01-15abs ↗pdf ↗