A new algorithm avoids worst-case outcomes in risky contexts.
problem Risk-averse behavior in contextual bandits is challenging.
method Developed a first risk-averse contextual bandit algorithm with online regret guarantees.
result First algorithm with an online regret guarantee for risk-averse contextual bandits.
New algorithms minimize costs while adhering to risk constraints in adversarial contextual bandits.
problem Minimizing cumulative cost while satisfying long-term risk constraints in adversarial contextual bandits.
method Developed a meta algorithm using online mirror descent for the full information setting, extended to contextual bandit with risk constraints using expert advice.
result Achieved near-optimal regret in terms of minimizing total cost, with sublinear growth of cumulative risk constraint violation.
Unified framework for risk-aware policy learning in contextual bandits.
problem Optimizing decision rules in high-stakes domains with adverse outcomes.
method Distributional framework for Lipschitz-continuous risk functionals, with novel empirical concentration inequalities.
result Data-dependent suboptimality bounds with an i l d e O ( 1 / n ) ilde{\mathcal{O}}(1/\sqrt{n}) i l d e O ( 1/ n ) rate, matching risk-neutral offline policy optimization. Paper introduces risk assessment for contextual bandits without experiments.
problem Evaluate policies using logged data in context bandits.
method Lipschitz risk functionals and Off-Policy Risk Assessment (OPRA) framework.
result OPRA provides finite sample guarantees for various risk estimates.
The paper proposes a method to learn and leverage contextual preference distributions for better decision-making.
problem Heterogeneous and context-dependent human preferences in decision-making problems.
method A sequential learning-and-optimization pipeline using a bounded-variance score function gradient estimator to train a predictive model mapping contextual features to preference distributions.
result The approach reduces average post-decision surprise by up to 25 times compared to risk-averse baselines in a ridesharing environment.
The Allais and Ellsberg paradoxes show that the expected utility hypothesis and Savage's Sure-Thing Principle are violated in real life decisions. The popular explanation in terms of 'ambiguity aversion' is not completely accepted. On the other hand, we have recently introduced a notion of 'contextual risk' to mathemat…
Efficient algorithm for contextual bandits using unlabeled data.
problem Efficient algorithms for contextual bandits with i.i.d. covariates.
method BISTRO algorithm using unlabeled data and multiplicative approximation for ERM.
result BISTRO requires d calls to ERM oracle per round, improving computational efficiency.
The paper tackles risk-averse multi-armed bandit with linear payoffs.
problem Risk-averse contextual multi-armed bandit problem with linear payoffs.
method Apply Thompson Sampling algorithm for disjoint model and provide comprehensive regret analysis.
result Proved an O ( ( 1 + ρ + 1 ρ ) d ln T ln K δ d K T 1 + 2 ε ln K δ 1 ε ) O((1+ρ+\frac{1}ρ) d\ln T \ln \frac{K}δ\sqrt{d K T^{1+2ε} \ln \frac{K}δ \frac{1}ε}) O (( 1 + ρ + ρ 1 ) d ln T ln δ K d K T 1 + 2 ε ln δ K ε 1 ) regret bound for mean-variance criterion. Paper proposes a risk-aware decision-making framework for real-world sequential decisions.
problem Real-world sequential decision-making problems often have critical constraints that learning solutions often neglect.
method Actor multi-critic architecture with risk characterization.
result Our approach consistently satisfies system constraints with minimal performance toll.
New method for estimating and optimizing MDPs without stationarity.
problem Challenges in offline contextual MDP estimation without stationarity.
method Introduces a new adaptive estimation and cost optimization approach for contextual MDPs.
result First robust, theoretically backed method for offline contextual MDP estimation.
Risk-aware linear bandits optimize against adverse outcomes.
problem Optimizing decisions under risk in sequential decision-making problems.
method Proposed an optimistic UCB algorithm for contextual bandits with convex loss minimization.
result Regret guarantees similar to generalized linear bandits with convex optimization at each round.
A new algorithm for contextual combinatorial bandits reduces regret.
problem Balancing exploration and exploitation in sequential decisions with risk sensitivity.
method Contextual combinatorial conservative bandits framework, with algorithms and regret bounds proven.
result Proven regret bound of i l d e O ( d 2 + d T ) ilde O(d^2+d\sqrt{T}) i l d e O ( d 2 + d T ) for known conservative reward, and unknown reward case. Uncertainty in economics still poses some fundamental problems illustrated, e.g., by the Allais and Ellsberg paradoxes. To overcome these difficulties, economists have introduced an interesting distinction between 'risk' and 'ambiguity' depending on the existence of a (classical Kolmogorovian) probabilistic structure m…
This paper improves offline contextual bandits using distributional robustness.
problem Improving offline contextual bandits with robustness.
method Extends Distributionally Robust Optimization (DRO) for offline contextual bandits, introducing a convex reformulation of Counterfactual Risk Minimization.
result Automatic calibration of asymptotic confidence intervals for policy optimization.
Meta-learning reformulated as Bayesian risk minimization.
problem Learning models to quickly adapt to new tasks from small datasets.
method Formalized meta-learning as Bayesian risk minimization, using a probabilistic framework to compute predictive distributions from posterior distributions of latent variables conditioned on contextual datasets.
result A novel Gaussian approximation for the posterior distribution that converges to maximum likelihood estimates and outperforms Neural Process on benchmark datasets.
The study proposes a Bayesian model to avoid filter bubbles by recommending articles with high uncertainty.
problem Filter bubbles limit exposure to diverse viewpoints, harming long-term user experiences.
method A Bayesian model of uncertainty-aware scoring and ranking for news articles is proposed. The model uses a Beta-distributed random variable conditional on context features.
result The proposed estimator outperforms existing algorithms in identifying successful outliers, improving personalized targeting of exceptional articles.
New algorithms learn model complexity and stochasticity robustly in online prediction.
problem Learning model complexity and stochasticity in online prediction.
method Probabilistic structural risk minimization integrated into adaptive algorithms.
result Competitive regret bounds for model and stochasticity adaptivity.
Paper develops robust OPF method using contextual information.
problem Optimal Power Flow problem under incomplete uncertainty knowledge.
method Distributionally robust chance-constrained formulation with probability trimmings and optimal transport.
result Distributional robustness improves expected cost and system reliability.
New method for sequential probability assignment reduces regret using contextual Shtarkov sums.
problem Minimizing regret in sequential probability assignment with arbitrary hypothesis classes.
method Introducing contextual Shtarkov sum and contextual Normalized Maximum Likelihood (cNML) algorithm.
result The contextual Shtarkov sum characterizes minimax regret and provides a minimax optimal strategy.
A new approach to hedging using contextual bandit models outperforms traditional methods.
problem Effective replication of financial contracts in incomplete markets with low transaction costs.
method Viewing hedging as a contextual k k k -armed bandit problem, using reinforcement learning. result The contextual bandit model provides more accurate and sample-efficient hedging than Q Q Q -learning. A new model optimizes portfolios by learning stock return distributions conditioned on factors.
problem Optimizing portfolios with high-dimensional asset-specific factors.
method Conditional Diffusion Transformer architecture linking each asset's return to its factor vector.
result The model outperforms benchmarks in mean-variance and mean-CVaR optimization.
This paper improves route choice models by incorporating contextual factors.
problem Existing route choice models lack consideration of dynamic contextual conditions.
method Knowledge distillation from Stated Choice Experiments in Immersive Virtual Environment.
result High-fidelity route choice models with increased predictive power.
Develops new optimization techniques for decision-making under uncertainty.
problem Decision-making under uncertainty with complex cost functions and nested expectations.
method Introduces Multistage Conditional Compositional Optimization (MCCO) and develops multilevel Monte Carlo techniques.
result New optimization techniques reduce scenario complexity from exponential to polynomial growth.
Shrinkage covariance is a multifactor model with risk factors and principal components.
problem Out-of-sample instabilities in sample covariance matrices.
method Combines risk factors and principal components with a block-diagonal factor covariance matrix.
result Shrinkage is a regularization scheme with less out-of-sample instability.
Develops new methods for risk-aware decision-making in medical bandits.
problem Risk-averse decision-making in medical contexts with limited data.
method Safe, anytime-valid concentration bounds, risk-aware contextual bandits, nonparametric algorithms.
result Improved decision-making algorithms for postoperative patient follow-up.
Universal algorithm learns unknown distribution for various decision-making problems.
problem Various statistical measures in contextual sequential decision-making.
method Infinite-dimensional functional regression oracle for cumulative distribution functions.
result Utility regret rate bounded by polynomial decay of eigenvalue sequence.
This paper sets communication complexity bounds for distributed RL.
problem Establishing minimum communication requirements for distributed RL.
method Information-theoretic lower bounds and algorithm development.
result Developed algorithms achieving optimal risk up to logarithmic factors.
Paper studies CLO with partial feedback, improving decision-making in uncertain contexts.
problem Improving decision-making in contexts with uncertain cost coefficients using partial feedback.
method Unified class of offline learning algorithms for CLO with different types of feedback, using IERM framework.
result Fast-rate regret bound for IERM with partial feedback and misspecified model classes.
Develops a robust learning method for unknown context distributions.
problem Learning from data in different, unknown contexts.
method Focuses on excess risks, constructs distribution sets with statistical coverage.
result Shows robustness in worst-case scenarios without sacrificing nominal performance.
Develops active learning method for linear optimization with margin-based criterion.
problem Optimizing decisions in linear optimization problems with limited labeled data.
method Smart Predict-then-Optimize (SPO) loss and margin-based active learning algorithm.
result Algorithm achieves significantly fewer labels than naive supervised learning, especially for minimizing SPO loss.
A new DRL model for intraday trading incorporating positional context.
problem Neglecting positional context in existing DRL intraday trading strategies.
method Introducing positional features into the state space of a DRL model.
result Significant improvement in profitability and risk-adjusted metrics.
Contextualized ML learns context-dependent effects using deep learning.
problem Learning heterogeneous and context-dependent effects in data.
method Applying deep learning to the meta-relationship between contextual information and context-specific parametric models.
result Unified framework for cluster analysis and cohort modeling.
IDS improves reinforcement learning with contextual information.
problem Optimizing IDS for contextual reinforcement learning.
method Investigated contextual bandit problems and proposed a computationally-efficient IDS.
result Contextual IDS outperforms conditional IDS by considering future contexts.
RiskLabs uses LLMs to predict financial risks from multimodal data.
problem Financial risk prediction using AI techniques.
method Integrates multimodal financial data (textual, vocal, time series, news) into LLMs for prediction.
result Empirical results show effectiveness in forecasting market volatility and variance.
Observer learns optimal policy from learner's actions without rewards.
problem Learning optimal policy from non-rewarded actions of a non-stationary learner.
method Two-Phase Suffix Imitation framework.
result Observer achieves convergence rate of O ~ ( 1 / N ) \tilde O(1/\sqrt{N}) O ~ ( 1/ N ) . Improves off-policy evaluation weights for balanced state-action pair distribution.
problem Imbalance in importance sampling weights for off-policy evaluation of contextual bandits.
method Balanced Off-Policy Evaluation (B-OPE) method that minimizes imbalance to desired counterfactual distribution of state-action pairs.
result Experimental evidence shows B-OPE improves offline policy evaluation in both discrete and continuous action spaces.
SADCBO optimizes contextual variables by balancing relevance and cost.
problem Optimizing contextual variables with varying costs and unknown relevance.
method Adaptive selection of relevant contextual variables using sensitivity analysis and early stopping.
result Consistent improvement in optimization across various examples.
Anomaly detection method separates contextual from behavioral attributes.
problem Detect anomalies in data without labeled examples.
method Uses joint deep variational generative models.
result Robust to anomalous or novel contextual attributes.
New framework for conditional risk minimization using optimal transport.
problem High-stakes decisions with side information, especially economic conditions.
method Universal framework based on union-ball formulation in optimal transport.
result Offers interpretability, tractability, and scalability for various risk functionals.
The bundle approach and n-contextuality reveal quantum model contextuality.
problem Understanding contextuality in quantum models using topology.
method Using the bundle approach, we describe contextuality as the non-existence of global sections in the measure bundle. We introduce n-contextuality to explore model dependence on scenario topology.
result Quantum theory and GHZ models exhibit all levels of n-contextuality, showing contextuality is related to holonomy group non-triviality.
BSG learns dynamic network spillovers and uncertainty quantification.
problem Identifying indirect spillovers and systemic risk in dynamic networks.
method Bayesian Spillover Graphs using FEVD and Bayesian time series models.
result Significant performance gains over baselines in identifying source and sink nodes.
The study evaluates forecast risk-adjusted performance using various metrics.
problem Evaluating forecast reliability beyond accuracy.
method Risk-adjusted performance measures (Sharpe, Sortino, Omega ratios) and Edge Ratio.
result Machine learning models often offer attractive risk profiles but not necessarily higher reliability.
OSOM solves multi-armed and linear contextual bandits efficiently.
problem Simultaneously optimal algorithm for multi-armed and linear contextual bandits.
method Design of a single computationally efficient algorithm that adapts to both regimes.
result Simultaneously optimal regret rates in both simple multi-armed and linear contextual bandits.
Proposes a neural network for contextual regression.
problem Improving model efficiency and interpretability in regression with contextual features.
method Simple contextual neural network (SCtxtNN) that separates context identification from context-specific regression.
result SCtxtNN achieves lower excess mean squared error and more stable performance than feed-forward neural networks.
Paper tackles domain adaptation for contextual bandits with sub-linear regret.
problem Adapting contextual bandit algorithms across domains with distribution shift.
method Learn a bandit model for the target domain using feedback from the source domain.
result Sub-linear regret bound maintained across domains.
Transformers model contextual relations using probabilistic measures, revealing their expressive power.
problem Lack of clear understanding of Transformer's ability to model contextual relations.
method Introduced a measure-theoretic framework connecting softmax attention and entropy-regularized optimal transport.
result Transformer architectures can approximate arbitrary contextual relations, and the choice of normalization affects how these relations are represented.
Paper solves stochastic contextual linear bandits using linear bandit algorithms.
problem Stochastic contextual linear bandits with unknown context distribution.
method Establishes a reduction framework to convert to linear bandit problems.
result Achieves nearly optimal regret bound of O ( d T log T ) O(d\sqrt{T\log T}) O ( d T log T ) . Introduces R package for contextual bandit algorithms.
problem Lack of standardized tools for comparing bandit algorithms.
method Object-oriented R package for parallelized comparison of contextual and context-free bandit policies.
result Facilitates simulation and offline analysis of bandit algorithms.