Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

176351527702 · Jun 202019922001200920172026
48 results for neutral reward function

AlphaZeroBeta uses deep reinforcement learning for market-neutral portfolios, outperforming traditional methods.

problem Traditional portfolio management methods often fail during market regime shifts or when assumptions break down.
method Combines a composite reward function and CNN-GRU policy trained end-to-end via Recurrent PPO.
result Achieves higher Sharpe ratios than baselines while maintaining near-zero benchmark correlations.

Online learning has traditionally focused on the expected rewards. In this paper, a risk-averse online learning problem under the performance measure of the mean-variance of the rewards is studied. Both the bandit and full information settings are considered. The performance of several existing policies is analyzed, an…

2018-07-24abs ↗pdf ↗

Study optimal stopping for variable annuity contracts with discontinuous rewards.

problem Optimal timing to surrender a variable annuity contract with guaranteed minimum benefit.
method Analytical study of an optimal stopping problem with a discontinuous reward function, considering general fee and surrender charge functions.
result Characterization of the surrender region and its interrelation with fee and surrender charge functions.

The paper has 2 main goals: 1. We propose a variant of the CAPM based on coherent risk. 2. In addition to the real-world measure and the risk-neutral measure, we propose the third one: the extreme measure. The introduction of this measure provides a powerful tool for investigating the relation between the first two mea…

2006-05-02abs ↗pdf ↗

The paper addresses human-like decision-making in multi-agent systems using bounded risk-sensitive Markov Games.

problem Modeling human-like decision-making in multi-agent systems with risk-seeking and loss-aversion behaviors.
method Forward policy design and inverse reward learning with iterative reasoning and cumulative prospect theory.
result The proposed algorithms demonstrate both risk-averse and risk-seeking behaviors in multi-agent systems.

Unified framework for risk-aware policy learning in contextual bandits.

problem Optimizing decision rules in high-stakes domains with adverse outcomes.
method Distributional framework for Lipschitz-continuous risk functionals, with novel empirical concentration inequalities.
result Data-dependent suboptimality bounds with an ildeO(1/n) ilde{\mathcal{O}}(1/\sqrt{n}) rate, matching risk-neutral offline policy optimization.

Neural nets optimize dynamic hedging strategies with transaction costs.

problem Optimal hedging strategy in presence of transaction costs and discrete time.
method Convolutional neural network trained to infer optimal hedging frequencies.
result Dynamic multiscale hedging strategy reduces risk and maximizes profit.

Proposes a method to construct risk-neutral marginals from arbitrage-free option prices.

problem Lack of risk-neutral marginals that are free of arbitrage and easy to use.
method Explicit construction of risk-neutral marginals from discrete arbitrage-free option prices.
result Explicit construction guarantees risk-neutral marginals free of butterfly and calendar arbitrage.

Efficient RL in partially observable risk-sensitive environments with hindsight observations.

problem Risk-sensitive reinforcement learning in partially observable environments.
method Integrates hindsight observations into POMDP framework, develops novel RL algorithm.
result Achieves polynomial regret with provable efficiency, outperforming existing methods.

The paper studies nilpotent structures in oriented neutral vector bundles and neutral hyperKähler structures.

problem Nilpotent structures in oriented neutral vector bundles and their relation to neutral hyperKähler structures.
method Defined HH-nilpotent structures for Lie subgroups of SO(2n,2n)SO(2n, 2n) related to neutral hyperKähler structures.
result Existence of complex and paracomplex structures forming neutral hyperKähler structures if and only if there exists an HH-nilpotent structure.

The purpose of this article is to review some recent results on the geometry of neutral signature metrics in dimension four and their twistor spaces. The following topics are considered: Neutral Kähler and hyperkähler surfaces, Walker metrics, Neutral anti-self-dual 4-manifolds and projective structures, Twistor spaces…

2008-04-14abs ↗pdf ↗

Word embedding models have become a fundamental component in a wide range of Natural Language Processing (NLP) applications. However, embeddings trained on human-generated corpora have been demonstrated to inherit strong gender stereotypes that reflect social constructs. To address this concern, in this paper, we propo…

2018-08-29abs ↗pdf ↗

This work characterizes reward function partial identifiability and its impact on policy optimization.

problem Reward function partial identifiability in complex tasks.
method Formal characterisation of partial identifiability using various reward learning data sources.
result Unified framework for comparing data sources and downstream tasks by their invariances.

This paper considers aspects of 4-manifold topology from the point of view of the null cone of a neutral metric, a point of view we call neutral causal topology. In particular, we construct and investigate neutral 4-manifolds with null boundaries that arise from canonical 3- and 4-dimensional settings. A null hypersurf…

2016-05-31abs ↗pdf ↗

The paper shows how to calculate risk-neutral default probabilities from bid and ask CDS quotes.

problem Calculating risk-neutral default probabilities from market quotes.
method Using conic finance framework and Poisson process to formulate and solve the calibration problem.
result A unique solution for risk-neutral default probabilities and implied liquidity.

We present a novel method for learning a set of disentangled reward functions that sum to the original environment reward and are constrained to be independently obtainable. We define independent obtainability in terms of value functions with respect to obtaining one learned reward while pursuing another learned reward…

2019-01-24abs ↗pdf ↗

Enhances reward specification in RL with a novel language-based approach.

problem Reward specification in RL can lead to unintended, potentially harmful behaviours.
method Developed a novel class of language-based Reward Machines using RML's built-in memory.
result Can specify non-regular, non-Markovian reward functions for complex tasks.

The paper proposes a new method to estimate interest rates consistently under both risk-neutral and real-world measures.

problem Consistent estimation of interest rates under both risk-neutral and real-world measures.
method Proposes a framework using progressive and square-integrable functions to specify the change of measure, and introduces two time-dependent candidates: step and linear functions.
result The proposed methods produce more stable and realistic long-term interest rate forecasts compared to using a constant function.

New algorithm for reward-free RL with linear function approximation, reducing sample complexity.

problem Efficiently learning optimal policies without prior reward information in complex environments.
method Developed an algorithm for reward-free RL in linear Markov decision processes, proving sample complexity bounds.
result Polynomial sample complexity in feature dimension and planning horizon, independent of states and actions.

The study finds that specific distributions can be used for risk-neutral valuation in Heston's SV model.

problem Valuation of European options under Heston's stochastic volatility model.
method Analyzing scale-parameter distributions and proving their equivalence to Heston's solution.
result Any RND with mean as the forward spot price that satisfies Heston's option valuation solution must be a member of a scale-family of distributions.

New lower bounds for combinatorial multi-armed bandits for general reward functions.

problem Maximizing reward in sequential decisions with sets of arms.
method Proved tight regret lower bounds for all smooth reward functions under mild assumptions.
result Lower bounds are tight up to log-factors for monotone reward functions.

Designers of AI agents often iterate on the reward function in a trial-and-error process until they get the desired behavior, but this only guarantees good behavior in the training environment. We propose structuring this process as a series of queries asking the user to compare between different reward functions. Thus…

2018-09-09abs ↗pdf ↗

The aim of this paper is to give examples of compact neutral 4-manifolds (M,g)(M,g) whose Ricci tensor ρρ satisfies the relation Xρ(X,X)=13Xτg(X,X)\nabla_Xρ(X,X) =\frac13Xτg(X,X). We present also a family of new Einstein bi-Hermitian neutral metrics on ruled surfaces of genus g>1g>1.

2008-01-14abs ↗pdf ↗

PQR estimates reward functions from actions and states without assuming state-only rewards.

problem Estimating reward functions from actions and states without state-only assumptions.
method Deep learning approach that sequentially estimates policy, Q-function, and reward.
result PQR uniquely recovers true reward with known transitions and bounds error with unknown transitions.

In many sequential decision making tasks, it is challenging to design reward functions that help an RL agent efficiently learn behavior that is considered good by the agent designer. A number of different formulations of the reward-design problem, or close variants thereof, have been proposed in the literature. In this…

2018-04-17abs ↗pdf ↗

One obstacle to applying reinforcement learning algorithms to real-world problems is the lack of suitable reward functions. Designing such reward functions is difficult in part because the user only has an implicit understanding of the task objective. This gives rise to the agent alignment problem: how do we create age…

2018-11-19abs ↗pdf ↗

Paper studies pricing and hedging of nonreplicable insurance contracts using benchmark-neutral approach.

problem Pricing and hedging of long-term insurance contracts like variable annuities.
method Benchmark-neutral pricing framework using stock growth optimal portfolio as numéraire.
result Prices can be significantly lower than risk-neutral ones, offering attractive long-term risk-management.