The paper addresses human-like decision-making in multi-agent systems using bounded risk-sensitive Markov Games.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The theory of convex risk functions has now been well established as the basis for identifying the families of risk functions that should be used in risk averse optimization problems. Despite its theoretical appeal, the implementation of a convex risk function remains difficult, as there is little guidance regarding ho…
Develops a new method to compute risk-sharing allocations using Laplace transforms.
Paper uses inverse optimization to measure risk preference from investment portfolios.
Active learning from demonstration allows a robot to query a human for specific types of input to achieve efficient learning. Existing work has explored a variety of active query strategies; however, to our knowledge, none of these strategies directly minimize the performance risk of the policy the robot is learning. U…
Study on estimating invertible functions with minimax analysis.
Researchers calculated EVaR for various distributions using Lambert function.
One typical assumption in inverse reinforcement learning (IRL) is that human experts act to optimize the expected utility of a stochastic cost with a fixed distribution. This assumption deviates from actual human behaviors under ambiguity. Risk-sensitive inverse reinforcement learning (RS-IRL) bridges such gap by assum…
We address the problem of inverse reinforcement learning in Markov decision processes where the agent is risk-sensitive. In particular, we model risk-sensitivity in a reinforcement learning framework by making use of models of human decision-making having their origins in behavioral psychology, behavioral economics, an…
Reduces risk of model inversion by reducing sensitive feature influence.
A new model forecasts Value-at-Risk using NIG distribution and dynamic scores.
Bayesian Robust Optimization for Imitation Learning (BROIL) balances risk and reward.
This paper tackles regularization parameter learning in inverse problems using data-driven bilevel optimization.
In this paper, we obtain analytical expression for the distribution of the occupation time in the red (below level ) up to an (independent) exponential horizon for spectrally negative Lévy risk processes and refracted spectrally negative Lévy risk processes. This result improves the existing literature in which only…
New methods improve portfolio risk minimization by estimating covariance matrix more accurately.
We model human decision-making behaviors in a risk-taking task using inverse reinforcement learning (IRL) for the purposes of understanding real human decision making under risk. To the best of our knowledge, this is the first work applying IRL to reveal the implicit reward function in human risk-taking decision making…
Framework uses IRL and RL to elicit and optimize risk preferences robustly to noise.
The authors aim to develop numerical schemes of the two representative quadratic hedging strategies: locally risk minimizing and mean-variance hedging strategies, for models whose asset price process is given by the exponential of a normal inverse Gaussian process, using the results of Arai et al. \cite{AIS}, and Arai …
Optimal insurance minimizes ruin probability with non-decreasing functions.
Paper uses SGD for solving linear inverse problems, improving empirical performance.
We consider the issue of solution uniqueness for portfolio optimization problem and its inverse for asset returns with a finite number of possible scenarios. The risk is assessed by deviation measures introduced by [Rockafellar et al., Mathematical Programming, Ser. B, 108 (2006), pp. 515-540] instead of variance as in…
CREDO assesses decision optimality under uncertainty without assuming a model.
We construct a data-driven statistical indicator for quantifying the tail risk perceived by the EURGBP option market surrounding Brexit-related events. We show that under lognormal SABR dynamics this tail risk is closely related to the so-called martingale defect and provide a closed-form expression for this defect whi…
The paper models cryptocurrency price and volatility with jumps and fractional volatility.
The paper identifies and critiques problems with risk matrices using ordinal scales.
Learning near-optimal behaviour from an expert's demonstrations typically relies on the assumption that the learner knows the features that the true reward function depends on. In this paper, we study the problem of learning from demonstrations in the setting where this is not the case, i.e., where there is a mismatch …
The paper solves classical problems in option pricing.
Study recovers investor preferences from portfolio data using synthetic data and robust optimization.
This paper solves the inversion problem for jump processes using Markovian projections.
Study utility indifference pricing in a Bachelier model with small linear price impact.
Robo-advisors estimate clients' risk aversion using interactive questionnaires.
Study optimal liquidation under high risk aversion and small price impact.
Deep models memorize training data in geophysical inversion, leading to biased posterior distributions.
Develops new instance-optimality concepts in differential privacy.
We analyze a nonlinear equation proposed by F. Black (1968) for the optimal portfolio function in a log-normal model. We cast it in terms of the risk tolerance function and provide, for general utility functions, existence, uniqueness and regularity results, and we also examine various monotonicity, concavity/convexity…
In the field of reinforcement learning there has been recent progress towards safety and high-confidence bounds on policy performance. However, to our knowledge, no practical methods exist for determining high-confidence policy performance bounds in the inverse reinforcement learning setting---where the true reward fun…
Inverse classification, the process of making meaningful perturbations to a test point such that it is more likely to have a desired classification, has previously been addressed using data from a single static point in time. Such an approach yields inflated probability estimates, stemming from an implicitly made assum…
New framework calibrates decision robustness using inverse conformal risk control.
In this paper, we obtain a property of the expectation of the inverse of compound Wishart matrices which results from their orthogonal invariance. Using this property as well as results from random matrix theory (RMT), we derive the asymptotic effect of the noise induced by estimating the covariance matrix on computing…
We present a Bayesian view of counterfactual risk minimization (CRM) for offline learning from logged bandit feedback. Using PAC-Bayesian analysis, we derive a new generalization bound for the truncated inverse propensity score estimator. We apply the bound to a class of Bayesian policies, which motivates a novel, pote…
Inverse reinforcement learning has proved its ability to explain state-action trajectories of expert agents by recovering their underlying reward functions in increasingly challenging environments. Recent advances in adversarial learning have allowed extending inverse RL to applications with non-stationary environment …
This article examines arbitrage investment in a mispriced asset when the mispricing follows the Ornstein-Uhlenbeck process and a credit-constrained investor maximizes a generalization of the Kelly criterion. The optimal differentiable and threshold policies are derived. The optimal differentiable policy is linear with …
The standard probabilistic perspective on machine learning gives rise to empirical risk-minimization tasks that are frequently solved by stochastic gradient descent (SGD) and variants thereof. We present a formulation of these tasks as classical inverse or filtering problems and, furthermore, we propose an efficient, g…
Innovative extensions to option pricing models using asymmetric Brownian motion and random walk approaches.
The study finds that specific distributions can be used for risk-neutral valuation in Heston's SV model.
New algorithm corrects bias in LDP-released data for better analysis.
DO-IQS recovers optimal stopping region from expert trajectories, addressing specific challenges.
A non-parametric method for evaluation of the aggregate loss distribution (ALD) by combining and numerically inverting the empirical characteristic functions (CFs) is presented and illustrated. This approach to evaluate ALD is based on purely non-parametric considerations, i.e., based on the empirical CFs of frequency …