Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

23466992 · May 202619922001200920172026
48 results for genuine rewards

New metrics improve scRNA-seq perturbation modeling by reducing mode collapse.

problem Outperformed by simple mean prediction in scRNA-seq perturbation modeling.
method Introduce DEG-aware metrics (WMSE, Rw2(Δ)R^{2}_{w}(Δ)) and negative/positive baselines.
result WMSE loss function reduces mode collapse and improves model performance.

The paper examines how updates to probabilistic models influence behavior based on evidence.

problem Understanding how updates to probabilistic models influence behavior based on evidence.
method Study of KL-regularized soft updates as Bayesian posterior updates within a single probabilistic model.
result Posterior updates determine relative incentives but not absolute rewards, which are ambiguous up to context-specific baselines.

We show that if a closed atoroidal 3-manifold M contains a genuine lamination, then it is group negatively curved in the sense of Gromov. Specifically, we exploit the structure of the non-product complementary regions of the genuine lamination and then apply the first author's Ubiquity Theorem to show that M satisfies …

1998-05-11abs ↗pdf ↗

We extend to the conformal realm the concept of genuine deformations of submanifolds, introduced by Dajczer and the first author for the isometric case. Analogously to that case, we call a conformal deformation of a submanifold MnM^n genuine if no open subset of MnM^n can be included as a submanifold of a higher dimens…

2008-06-03abs ↗pdf ↗

New method decomposes Markov chain rewards into persistent and transient components.

problem Ambiguity in classical evaluation methods for Markov chains with reducible and periodic states.
method Minimal exact quotient by the real peripheral invariant subspace, decomposing rewards into persistent and transient components.
result Exact comparison with classical methods shows that the new decomposition reallocates the same information, making persistent modes explicit.

Signed compression progress on a sealed audit is goodhart-resistant.

problem Intrinsic motivation for agents to improve their world models by compressing experience.
method Rewarding agents for the signed decrease of a fixed sealed-audit loss.
result Cumulative reward telescopes exactly to endpoint audit improvement, preventing infinite reward push while true audit performance stagnates.

We classify hypersurfaces of rank two of Euclidean space Rn+1\R^{n+1} that admit genuine isometric deformations in Rn+2\R^{n+2}. That an isometric immersion f^ ⁣:MnRn+2\hat f\colon\,M^n\to\R^{n+2} is a genuine isometric deformation of a hypersurface f ⁣:MnRn+1f\colon\, M^n\to\R^{n+1} means that f^\hat f is nowhere a composition $\hat f=\ha…

2010-10-14abs ↗pdf ↗

We extend the concept of genuine rigidity of submanifolds by allowing mild singularities, mainly to obtain new global rigidity results and unify the known ones. As one of the consequences, we simultaneously extend and unify Sacksteder and Dajczer-Gromoll theorems by showing that any compact nn-dimensional submanifold …

2018-03-16abs ↗pdf ↗

In this paper we classify Euclidean hypersurfaces f ⁣:MnRn+1f\colon M^n \rightarrow \mathbb{R}^{n+1} with a principal curvature of multiplicity n2n-2 that admit a genuine conformal deformation f~ ⁣:MnRn+2\tilde{f}\colon M^n \rightarrow \mathbb{R}^{n+2}. That f~ ⁣:MnRn+2\tilde{f}\colon M^n \rightarrow \mathbb{R}^{n+2} is a genuine conformal defo…

2018-05-17abs ↗pdf ↗

A basic question in submanifold theory is whether a given isometric immersion f ⁣:MnRn+pf\colon M^n\to\R^{n+p} of a Riemannian manifold of dimension n3n\geq 3 into Euclidean space with low codimension pp admits, locally or globally, a genuine infinitesimal bending. That is, if there exists a genuine smooth variation of ff by…

2019-04-23abs ↗pdf ↗

Develops a method for solving optimal stopping problems with multiple exercise rights.

problem Optimal stopping with multiple exercise rights under model uncertainty.
method Pathwise duality approach based on robust martingale dual representation.
result Establishes upper and lower bounds that converge to the true solution.

We construct a pair of transverse genuine laminations on an atoroidal 3-manifold admitting transversely orientable uniform 1-cochain. The laminations are induced by the uniform 1-cochain and they are indeed the "straightening" of the coarse laminations defined in [Ca], by using minimal surface techniques. Moreover, whe…

2003-04-07abs ↗pdf ↗

In the Minority, Majority and Dollar Games (MG, MAJG, $G), synthetic agents compete for rewards, at each time-step acting in accord with the previously best-performing of their limited sets of strategies. Different components and/or aspects of real-world financial markets are modelled by these games. In the MG, agents …

2008-02-28abs ↗pdf ↗

As advances in signature recognition have reached a new plateau of performance at around 2% error rate, it is interesting to investigate alternative approaches. The approach detailed in this paper looks at using Variational Auto-Encoders (VAEs) to learn a latent space representation of genuine signatures. This is then …

2019-04-18abs ↗pdf ↗

We examine the difference between several notions of curvature homogeneity and show that the notions introduced by Kowalski and Vanžurová are genuine generalizations of the ordinary notion of kk-curvature homogeneity. The homothety group plays an essential role in the analysis.

2013-09-20abs ↗pdf ↗

We study left-invariant symmetric Killing 2-tensors on 2-step nilpotent Lie groups endowed with a left-invariant Riemannian metric, and construct genuine examples, which are not linear combinations of parallel tensors and symmetric products of Killing vector fields.

2018-11-22abs ↗pdf ↗

We show that among the Euclidean submanifolds with codimension two the ones of rank two that are parabolic but nonruled are isometrically rigid. This generalizes the result in [10] that these submanifolds are genuinely rigid. In addition, we give a parametric classifications of all parabolic submanifolds.

2009-03-31abs ↗pdf ↗

BCPO optimizes offline RL policies by converting uncertainty into conservative bounds.

problem Offline RL's fragility under distribution shifts and model errors.
method Bayesian approach with credible lower bounds and KL regularization.
result BCPO yields an uncertainty-calibrated policy that avoids exploiting model errors.

AI needs causal inference to avoid being just a correlation machine.

problem AI's inability to distinguish correlation from causation.
method Develops a unified framework connecting various causal statistical estimators and proves a Statistical Necessity Theorem for causal generalization.
result AI systems without causal grounding are brittle and biased, highlighting the need for causal statistics.

Study on fake stationary Volterra Heston model for non-stationary processes.

problem Non-stationary nature of true Volterra equations.
method Weak notion of stationarity (fake stationary regime) for inhomogeneous affine Stochastic Volterra equations.
result Existence of limiting distributions in the long run, which may depend on initial state.

Paper extends SI method for detecting CPs in complex systems' frequency domain.

problem Identifying change points in complex systems' frequency domain.
method Extends SI framework to frequency domain using DFT properties and develops valid p-values.
result Reliable detection of genuine CPs with strong statistical guarantees.

Study categorizes time series anomaly detection metrics based on evaluation challenges.

problem Challenges in evaluating time series anomaly detection due to diverse application objectives and metric assumptions.
method Problem-oriented framework categorizing metrics into six dimensions based on evaluation challenges.
result Quantifies each metric's discriminative ability and reveals limitations of widely used metrics.

Using ideas from an article of P. Bieliavsky, M. Rooman and Ph. Spindel on BTZ black holes, I construct a family of interesting examples of quasi-Poisson actions as defined by A. Alekseev and Y. Kosmann-Schwarzbach. As an application, I obtain a genuine Poisson structure on SL(2,R)SL(2,R) which induces a Poisson structure o…

2004-09-29abs ↗pdf ↗

Reward hacking exploits misspecified rewards, affecting agent capabilities and true performance.

problem Reward hacking in RL models exploiting reward misspecifications.
method Constructed four RL environments with misspecified rewards; analyzed agent capabilities and behavior.
result More capable agents exploit reward misspecifications, achieving higher proxy reward but lower true reward.

Paper introduces PRMs to learn non-Markovian stochastic rewards for reinforcement learning.

problem Lack of structured representation for non-Markovian stochastic rewards in reinforcement learning.
method Introduces probabilistic reward machines (PRMs) and presents an algorithm to learn them from decision processes.
result Algorithm proves correct and convergent for learning PRMs from decision processes.

In this paper we prove that, in the category of chain complexes, partial algebras can be functorially replaced by quasi-isomorphic algebras. In particular, partial algebras contain all of the important homological and homotopical information that genuine algebras do. Applying this result to McClure's partial algebra in…

2004-10-18abs ↗pdf ↗

Self-supervised reward prediction improves RL in sparse reward settings.

problem Data efficiency and sparse reward signals in reinforcement learning.
method Learning a state representation for reward prediction and using it to shape rewards.
result Self-supervised reward prediction enhances RL algorithms in single-goal environments.

The study categorizes reward errors in reinforcement learning, finding some can be beneficial.

problem Training language models with imperfect proxy rewards.
method Theoretical analysis of policy gradient optimization and categorization of reward errors.
result Reward errors can be benign or even beneficial, preventing policy from stalling.

Reward collapse occurs when ranking-based reward models yield uniform rewards for different prompts.

problem Reward collapse in aligning large language models with human preferences.
method Introduced a prompt-aware optimization scheme to derive closed-form expressions for reward distributions.
result Our prompt-aware utility functions significantly alleviate reward collapse during training.