Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

285785113 · Jun 202019922001200920172026
48 results for multi-action scenarios

Paper tackles optimal policy learning with observational data in multi-action scenarios.

problem Optimal policy learning in multi-action settings with observational data.
method Review of estimation approaches, analysis of risk preference, discussion of potential failures.
result Average regret of a policy with multi-valued treatment is contingent on the decision-maker's attitude towards risk.

Neural Index Policy for multi-action bandits with heterogeneous budgets.

problem Real-world settings often involve multiple interventions with heterogeneous costs and constraints, breaking classical assumptions.
method Introduces a Neural Index Policy (NIP) that learns to assign budget-aware indices to arm-action pairs using a neural network and differentiable knapsack layer.
result Empirically achieves near-optimal performance while strictly enforcing heterogeneous budgets and scaling to hundreds of arms.

Lower bounds for PI on multi-action MDPs are established, showing complexity grows with action count.

problem Establishing the minimum number of iterations for PI to converge on MDPs with multiple actions.
method Developed lower bounds for a specific PI variant on multi-action MDPs, scaling with action count.
result A particular PI variant can take Ω(kn/2)Ω(k^{n/2}) iterations to terminate, scaling with action count.

In many settings, a decision-maker wishes to learn a rule, or policy, that maps from observable characteristics of an individual to an action. Examples include selecting offers, prices, advertisements, or emails to send to consumers, as well as the problem of determining which medication to prescribe to a patient. Whil…

2018-10-10abs ↗pdf ↗

Optimal policy for multi-armed multi-action bandits with unknown parameters.

problem Optimal sequential action selection for multi-armed multi-action bandits with unknown parameters.
method Occupancy-Measured-Reward Index Policy (OMRIP) and R(MA)^2B-UCB algorithm.
result Asymptotically optimal policy with sub-linear regret and low computational complexity.

Study proves duality in exotic option pricing under uncertain model and delayed information.

problem Pricing and hedging of multi-action exotic options under nondominated model uncertainty and delayed information.
method Reformulated superhedging problem as a European option problem, proving duality results.
result Superhedging price equals model-based price with future look-up power.

In this work we describe a novel deep reinforcement learning architecture that allows multiple actions to be selected at every time-step in an efficient manner. Multi-action policies allow complex behaviours to be learnt that would otherwise be hard to achieve when using single action selection techniques. We use both …

2018-03-14abs ↗pdf ↗

Policy learning can be used to extract individualized treatment regimes from observational data in healthcare, civics, e-commerce, and beyond. One big hurdle to policy learning is a commonplace lack of overlap in the data for different actions, which can lead to unwieldy policy evaluation and poorly performing learned …

2019-06-20abs ↗pdf ↗

ARL and Hawkes processes improve market-making strategies with variable volatility.

problem Enhancing market-making strategies to adapt to varying volatility levels and self-exciting behaviors.
method Integrates ARL, Hawkes processes, and variable volatility levels; shifts from Poisson to Hawkes process.
result 4-action MM trained in low-volatility environment adapts to high-volatility conditions, providing stable performance.

Paper derives policy rules from observational data for hepatitis C treatment.

problem Improving treatment guidelines for HIV/HCV co-infected patients.
method Weighted K-means algorithm for estimating CATEs, decision tree implementation.
result Identifies a subgroup with high spontaneous HCV clearance rate.

Research in deep learning for multi-speaker source separation has received a boost in the last years. However, most studies are restricted to mixtures of a specific number of speakers, called a specific scenario. While some works included experiments for different scenarios, research towards combining data of different…

2018-08-24abs ↗pdf ↗

Methodology measures financial impacts using existing credit loss infrastructure.

problem Measuring the impact of financial scenarios on expected credit losses.
method Captures scenario effects through changes in default probabilities; uses existing provisioning infrastructure.
result Methodology validated through standardized climate scenario exercise in Canada and Quebec.

Statistical depth metrics help identify risky power grid scenarios.

problem Identifying extreme scenarios for risk mitigation in power grid planning.
method Functional depth metrics for sub-selecting outlying scenarios.
result The proposed approach effectively identifies risky scenarios for operational risk mitigation.

A new method using energy distance for ensemble and scenario reduction.

problem Solving complex dynamic and stochastic programs, especially in energy systems.
method Proposes a new method based on energy distance for ensemble and scenario reduction.
result Reduced scenario sets exhibit better statistical properties for energy distance than Wasserstein distance.

Develops a method for reverse stress testing in multivariate scenarios.

problem Reconstructing a multivariate stress scenario from a single exogenous shock.
method Maximizing conditional density under three distributional assumptions.
result Simulated scenarios are economically coherent and reproduce risk-reward asymmetry.

Method generates plausible financial stress scenarios using large deviations.

problem Misleading risk management by overlooking or overemphasizing implausible scenarios.
method Exploits large-deviations principle to concentrate risk factors near most likely stress configurations.
result Can generate informative stress scenarios even with limited historical data.

Scenario discovery is the process of finding areas of interest, known as scenarios, in data spaces resulting from simulations. For instance, one might search for conditions, i.e., inputs of the simulation model, where the system is unstable. Subgroup discovery methods are commonly used for scenario discovery. They find…

2019-10-03abs ↗pdf ↗

Generates multimodal safety-critical scenarios for robustness evaluation of decision-making algorithms.

problem Lack of comprehensive evaluation of neural network robustness under real-world scenarios.
method Proposes a flow-based multimodal scenario generator using weighted likelihood maximization and gradient-based sampling.
result Demonstrates improved testing efficiency and multimodal modeling capability compared to traditional methods.

Adaptive framework generates challenging adversarial scenarios for autonomous vehicles.

problem Lack of efficient and adaptable evaluation methods for autonomous vehicles.
method Adaptive evaluation framework using ensemble models and nonparametric Bayesian clustering.
result Adversarial scenarios significantly degrade tested autonomous vehicles' performance.

Algorithm reduces historical expected shortfall computation by focusing on worst-case scenarios.

problem Computing the historical expected shortfall efficiently and accurately.
method Multi-step algorithm using Monte Carlo simulations to identify and reduce the number of worst-case scenarios.
result Non-asymptotic bounds for the L p-error of the expected shortfall estimator are derived.

Risk measures such as Expected Shortfall (ES) and Value-at-Risk (VaR) have been prominent in banking regulation and financial risk management. Motivated by practical considerations in the assessment and management of risks, including tractability, scenario relevance and robustness, we consider theoretical properties of…

2018-08-22abs ↗pdf ↗

Study optimal timing to divest from assets with uncertain future scenarios.

problem Optimal timing to divest from assets with uncertain future scenarios.
method Smooth model of decision making under ambiguity aversion, optimal stopping problem with learning.
result Proves a minimax result reducing the problem to standard optimal stopping problems with learning.

Researchers validate ML scenario generators by checking dependencies and detecting memorization effects.

problem Validation of machine learning-based scenario generators differs from classical methods due to data-driven dependencies.
method Two novel validation aspects: checking dependencies and detecting memorization effects. Novel memorization ratio introduced.
result Validation methods successfully detect dependencies and memorization effects in ML-based scenario generators.

Standard artificial neural networks suffer from the well-known issue of catastrophic forgetting, making continual or lifelong learning difficult for machine learning. In recent years, numerous methods have been proposed for continual learning, but due to differences in evaluation protocols it is difficult to directly c…

2019-04-15abs ↗pdf ↗

This paper presents a method for testing the decision making systems of autonomous vehicles. Our approach involves perturbing stochastic elements in the vehicle's environment until the vehicle is involved in a collision. Instead of applying direct Monte Carlo sampling to find collision scenarios, we formulate the probl…

2019-02-05abs ↗pdf ↗

The paper proposes an efficient nested simulation design using likelihood ratio method.

problem Designing nested simulations with fixed outer scenarios and minimizing simulation effort.
method Proposes a bi-level optimization problem to decide inner replications and pooling strategies.
result Optimized design achieves $\cO(Γ^{-1})$ mean squared error of estimators.

MARCD uses generative scenarios to improve portfolio decisions during regime shifts.

problem Improving portfolio decisions under regime shifts and drawdowns.
method MARCD employs a Gaussian HMM for regime inference, a diffusion generator for scenario production, and a CVaR allocator with tail-weighted and crisis-aware components.
result MARCD reduces maximum drawdowns by 34% compared to baseline methods over 2020-2025.

Detects causal scenarios with inequality constraints among classical correlations.

problem Classifying causal structures and identifying those with inequality constraints.
method Using d-separation, e-separation, incompatible supports, and HLP condition.
result Resolved all but three causal scenarios with up to 4 observed variables.

Generative Adversarial Network (GAN) simulates realistic multi-asset scenarios for tail risk.

problem Simulating realistic joint dynamics of multi-asset portfolios for tail risk estimation.
method Designing a GAN that preserves Value-at-Risk (VaR) and Expected Shortfall (ES) tail risk features.
result Correctly captures tail risk for a broad class of trading strategies and demonstrates strong generalization.

Consider an agent who enters a financial market on day t = 0 with an initial capital amount x. He invests this amount on stocks and the money market, and by day t = T, has generated a wealth W . He is given a convex class of probability measures (called scenarios) and a real-valued function (or floors) corresponding to…

2006-01-25abs ↗pdf ↗

Novel Bayesian meta-reinforcement learning framework improves traffic signal control robustness.

problem Lack of robustness and stability in adaptation for traffic signal control.
method Value-based Bayesian meta-reinforcement learning framework BM-DQN with fast-adaptation variation and DQN fast-update advantage.
result Framework adapts more quickly and robustly to new scenarios than previous methods.

We present a method to generate renewable scenarios using Bayesian probabilities by implementing the Bayesian generative adversarial network~(Bayesian GAN), which is a variant of generative adversarial networks based on two interconnected deep neural networks. By using a Bayesian formulation, generators can be construc…

2018-02-02abs ↗pdf ↗

There are a variety of Domain Adaptation (DA) scenarios subject to label sets and domain configurations, including closed-set and partial-set DA, as well as multi-source and multi-target DA. It is notable that existing DA methods are generally designed only for a specific scenario, and may underperform for scenarios th…

2019-12-08abs ↗pdf ↗

New method selects critical DER scenarios for distribution grid investment planning.

problem Determining critical DER adoption scenarios for risk assessment in distribution grids.
method Bayesian Optimization framework using Gaussian Process surrogates and Pareto-critical acquisition function.
result Statistical guarantee and significant speed-up over exhaustive search in selecting critical DER scenarios.

Paper introduces a new method for calibrating ESGs to both historical and forward-looking data.

problem Lack of a generally accepted methodology for calibrating ESGs to forward-looking information.
method Conditional Scenario Simulator framework for consistent calibration of economic and financial variables.
result Framework can embed various financial and macroeconomic models and demonstrate practical examples in frequentist and Bayesian settings.