Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

20406080 · Jun 202019922001200920172026
48 results for Risk-Averse Bandits

Motivated by applications in clinical trials and finance, we study the problem of online convex optimization (with bandit feedback) where the decision maker is risk-averse. We provide two algorithms to solve this problem. The first one is a descent-type algorithm which is easy to implement. The second algorithm, which …

2018-10-01abs ↗pdf ↗

This paper unifies risk-averse Thompson sampling for continuous risk functionals.

problem Designing and analyzing risk-averse Thompson sampling algorithms for continuous risk functionals.
method Developed analytical toolkits to prove asymptotically optimal regret bounds for various risk measures.
result Proved asymptotic optimality of ρρ-MTS for Bernoulli distributions and a class of risk measures.

The paper tackles risk-averse multi-armed bandit with linear payoffs.

problem Risk-averse contextual multi-armed bandit problem with linear payoffs.
method Apply Thompson Sampling algorithm for disjoint model and provide comprehensive regret analysis.
result Proved an O((1+ρ+1ρ)dlnTlnKδdKT1+2εlnKδ1ε)O((1+ρ+\frac{1}ρ) d\ln T \ln \frac{K}δ\sqrt{d K T^{1+2ε} \ln \frac{K}δ \frac{1}ε}) regret bound for mean-variance criterion.

The paper examines how loss aversion impacts multi-armed bandit decisions over long periods.

problem The impact of loss aversion on multi-armed bandit decisions over long periods.
method A new central limit theorem for measures with history-dependent variances, derived under risk aversion in gains and risk loving in losses.
result Consequences of loss aversion for asymptotic properties are derived in analytical results.

Nonparametric Thompson Sampling achieves optimal regret for risk-averse bandits with sub-Gaussian rewards.

problem Optimizing risk-averse bandit problems with sub-Gaussian rewards.
method Anchor-free nonparametric Thompson Sampling algorithm ρextNPTSSGρ ext{-}NPTS_{\mathrm{SG}}.
result Achieves regret matching the instance-dependent lower bound to leading order in logn\log n.

In this paper, we study multi-armed bandit problems in explore-then-commit setting. In our proposed explore-then-commit setting, the goal is to identify the best arm after a pure experimentation (exploration) phase and exploit it once or for a given finite number of times. We identify that although the arm with the hig…

2019-04-30abs ↗pdf ↗

We consider the problem of minimizing the regret in stochastic multi-armed bandit, when the measure of goodness of an arm is not the mean return, but some general function of the mean and the variance.We characterize the conditions under which learning is possible and present examples for which no natural algorithm can…

2014-05-05abs ↗pdf ↗

New algorithms for best arm identification in bandits robust to misspecified parameters.

problem Inconsistent learning performance of traditional MAB algorithms when parameters are misspecified.
method Proposes two classes of asymptotically near-optimal algorithms for statistically robust MAB under fixed-budget pure exploration.
result Establishes fundamental performance limits and proposes algorithms that are asymptotically near-optimal.

A new method for risk-averse decision-making in Markov processes with improved regret bounds.

problem Risk-averse decision-making in Markov processes.
method Introduces mini-batch measures and multipattern risk-averse problems in a feature-based QQ-learning method.
result Proves a high-probability regret bound of O(H2NHK)\mathcal{O}\big(H^2 N^H \sqrt{ K}\big) for the QQ-learning method.

New framework shifts bandit algorithms from expected reward to preference metrics, optimizing mixtures of arms.

problem Traditional bandit algorithms focus on expected rewards, ignoring variability and risk.
method Introduces preference metrics (PMs) and designs algorithms to optimize mixtures of arms.
result Optimal policy selects mixtures of arms based on specific weights, not a single best arm.

We introduce the functional bandit problem, where the objective is to find an arm that optimises a known functional of the unknown arm-reward distributions. These problems arise in many settings such as maximum entropy methods in natural language processing, and risk-averse decision-making, but current best-arm identif…

2014-05-10abs ↗pdf ↗

Online learning has traditionally focused on the expected rewards. In this paper, a risk-averse online learning problem under the performance measure of the mean-variance of the rewards is studied. Both the bandit and full information settings are considered. The performance of several existing policies is analyzed, an…

2018-07-24abs ↗pdf ↗

Develops new methods for risk-aware decision-making in medical bandits.

problem Risk-averse decision-making in medical contexts with limited data.
method Safe, anytime-valid concentration bounds, risk-aware contextual bandits, nonparametric algorithms.
result Improved decision-making algorithms for postoperative patient follow-up.

We propose and analyze StoROO, an algorithm for risk optimization on stochastic black-box functions derived from StoOO. Motivated by risk-averse decision making fields like agriculture, medicine, biology or finance, we do not focus on the mean payoff but on generic functionals of the return distribution. We provide a g…

2019-04-17abs ↗pdf ↗

A new approach to hedging using contextual bandit models outperforms traditional methods.

problem Effective replication of financial contracts in incomplete markets with low transaction costs.
method Viewing hedging as a contextual kk-armed bandit problem, using reinforcement learning.
result The contextual bandit model provides more accurate and sample-efficient hedging than QQ-learning.

A new framework for risk-aware multi-armed bandits tackles volatile environments.

problem Volatility in healthcare and finance makes naive reward maximization unreliable.
method Risk-aware strategies with adaptive risk measures and change-point detection.
result Finite-time theoretical guarantees and asymptotic regret bound of order ildeO(KTT) ilde O(\sqrt{K_T T}).

Different models of capital exchange among economic agents have been proposed recently trying to explain the emergence of Pareto's wealth power law distribution. One important factor to be considered is the existence of risk aversion. In this paper we study a model where agents posses different levels of risk aversion,…

2003-11-06abs ↗pdf ↗

Robo-advisors estimate clients' risk aversion using interactive questionnaires.

problem Estimating risk aversion of non-expert clients using adaptive questionnaires.
method Model risk aversion with cost functions and spectral risk measures. Use inverse reinforcement learning to design questions maximizing distinguishing power.
result Designing questions by maximizing distinguishing power achieves satisfactory accuracy in learning risk aversion with fewer than 50 questions.

Risk aversion is a key element of utility maximizing hedge strategies; however, it has typically been assigned an arbitrary value in the literature. This paper instead applies a GARCH-in-Mean (GARCH-M) model to estimate a time-varying measure of risk aversion that is based on the observed risk preferences of energy hed…

2011-03-30abs ↗pdf ↗

Modeling informed trading with risk-averse market makers.

problem Understanding informed trading and its impact on market liquidity and risk premia.
method Connections between optimal transport theory and Kyle's model, including new characterizations of profits and duality.
result Liquidity is lower, assets exhibit short-term reversals, and risk premia depend on market maker inventories, which are mean reverting.

Study risk-averse insider's behavior in dynamic signal asset pricing.

problem Analyzing risk-averse insider's dynamic signal in asset pricing.
method Employing a weak conditioning methodology to construct a Schrödinger bridge, deriving necessary conditions for equilibrium.
result Derive explicit closed-form solutions for important cases.

The standard asset pricing models (the CCAPM and the Epstein-Zin non-expected utility model) counterintuitively predict that equilibrium asset prices can rise if the representative agent's risk aversion increases. If the income effect, which implies enhanced saving as a result of an increase in risk aversion, dominates…

2014-03-04abs ↗pdf ↗

Investigates how diversification preferences relate to risk attitudes.

problem Connecting diversification preferences to risk attitudes.
method Analyzes diversification preferences for various pairs of risks under different conditions.
result Diversification preferences for certain pairs of risks imply specific levels of risk aversion.

New insights into risk aversion for complex decision models.

problem Understanding risk aversion in non-monotone decision models.
method Characterization of probabilistic risk aversion for generalized rank-dependent functions.
result Probabilistic risk aversion is determined by the distortion function, which is convex or scaled quantile-spread mixtures.

Unified formula for optimal portfolio under piecewise hyperbolic risk aversion.

problem Optimizing portfolios with piecewise hyperbolic risk aversion utilities.
method Derive a unified closed-form formula for the optimal portfolio.
result Unified formula reflects risk aversion behaviors and risk-taking behaviors.

Investigates a Kyle model with imperfect information and risk aversion.

problem Tackles a Kyle model with imperfect information and risk-averse informed traders.
method Solves an optimal transport problem and a filtering problem under specific measures.
result Constructs an equilibrium for the Gaussian Kyle model with imperfect information and risk aversion.

New methods reduce bias in estimating optimality gaps for risk-averse stochastic programs.

problem Optimality gap estimation bias in risk-averse stochastic programs.
method Two independent samples, each estimating a different component of the optimality gap.
result Our method reduces bias in estimating optimality gaps for risk-averse problems.

Paper develops NPG for risk-averse RL with ECRMs, proving global convergence.

problem Ensuring reliable performance in stochastic RL problems with risk-averse policies.
method Developed natural policy gradient updates for ECRMs-based RL problems, proving global optimality and iteration complexity.
result Global convergence of risk-averse NPG algorithm with ECRMs.

New MFG model for MV portfolio management with peer-based risk aversion.

problem Time-inconsistent mean-variance portfolio management with peer-based risk aversion.
method Mean-field game, smooth regularization, fixed-point arguments, convergence analysis.
result Existence of mean-field equilibrium in time-inconsistent MFG.

Study optimal liquidation under high risk aversion and small price impact.

problem Optimal liquidation of options under high risk aversion and linear price impact.
method Analyzes Bachelier model with linear price impact, computes utility indifference prices, and finds asymptotically optimal portfolios.
result Establishes a scaling limit for vanishing price impact and computes corresponding utility indifference prices.

Is the elasticity of intertemporal substitution (EIS) more or less than one? This question can be answered by confronting theoretical results of asset pricing models with investor behaviour during episodes of stock market panic. If we consider these episodes as periods of high risk aversion, then lower asset prices are…

2015-05-27abs ↗pdf ↗

This study measures price risk aversion using indirect utility functions in a lab experiment.

problem Measuring risk aversion with uncertain prices in experimental economics.
method Using indirect utility functions and a multiple price list method in a lab experiment.
result Price risk aversion is statistically greater than payoff risk aversion.

New model considers wealth and time affecting risk aversion in portfolio selection.

problem Optimal investment strategy and consumption process depend on wealth and future income balance.
method Proposed a new mean-variance-utility framework with time and state-dependent risk aversion, solved using game theory.
result Equilibrium investment and consumption policies derived, aligning with investor behavior.

This paper solves optimal consumption-investment choices with wealth-driven risk aversion using neural networks.

problem Optimal consumption-investment choices under wealth-driven risk aversion.
method Neural network LSTM trained on jump-diffusion model data to optimize investment rate and consumption.
result Neural network approach shows promising results in solving the investment problem.

Spectral risk measures are attractive risk measures as they allow the user to obtain risk measures that reflect their risk-aversion functions. To date there has been very little guidance on the choice of risk-aversion functions underlying spectral risk measures. This paper addresses this issue by examining two popular …

2011-03-29abs ↗pdf ↗

The geometric Lévy model (GLM) is a natural generalisation of the geometric Brownian motion model (GBM) used in the derivation of the Black-Scholes formula. The theory of such models simplifies considerably if one takes a pricing kernel approach. In one dimension, once the underlying Lévy process has been specified, th…

2011-11-09abs ↗pdf ↗

In this paper the fractional trading ansatz of money management is reconsidered with special attention to chance and risk parts in the goal function of the related optimization problem. By changing the goal function with due regards to other risk measures like current drawdowns, the optimal fraction solutions reflect t…

2016-12-09abs ↗pdf ↗

The paper characterizes equilibrium strategies under random risk aversion, showing unique solutions based on risk aversion distribution.

problem Characterizing equilibrium strategies in a continuous-time portfolio selection problem under random risk aversion.
method Provided a complete characterization of all deterministic equilibrium strategies in closed form, analyzing the structure of the solution based on the distribution of random risk aversion.
result The equilibrium is unique (if exists) when the expectation of random risk aversion is finite, but infinite expectation leads to either infinitely many equilibria or a unique trivial one.