Study of participating policies with guaranteed minimum interest rate and surrender option.
problem Analyzing the value and optimal surrender strategy of participating policies with minimum interest rate guarantee and surrender option.
method Probabilistic analysis using optimal stopping and free boundary theory.
result Identification of an optimal surrender strategy involving stop-loss and too-good-to-persist boundaries.
In the NIPS 2017 Learning to Run challenge, participants were tasked with building a controller for a musculoskeletal model to make it run as fast as possible through an obstacle course. Top participants were invited to describe their algorithms. In this work, we present eight solutions that used deep reinforcement lea…
Analyzes how economic policies affect wealth distribution in Bitcoin token economy.
problem Impact of economic policies on wealth distribution in token economies.
method Eliminated noise in wealth distribution data using macroeconomic and microeconomic time series. Causality analysis between BIPs and wealth distribution data.
result Proposed a structure for economic policy taxonomy in token economies.
This research improves interpretability in sequential explanations using mental models.
problem Improving interpretability in sequential explanations between two parties.
method A reinforcement learning framework that selects explanations based on the explainee's mental model.
result Mental model-based policies increase interpretability over random selection in multiple sequential explanations.
A mechanism to share risks and costs with guarantees against extreme outcomes.
problem Softening extreme individual burdens in risk sharing schemes.
method Formalizes Certified Allocation Problem; uses Conformal Risk Sharing with interpretable sharing policy and split conformal calibration.
result Reduces extreme obligations for high-risk agents while controlling harm to others.
Sequential decision making for lifetime maximization is a critical problem in many real-world applications, such as medical treatment and portfolio selection. In these applications, a "reneging" phenomenon, where participants may disengage from future interactions after observing an unsatisfiable outcome, is rather pre…
Framework optimizes battery storage for markets by separating long-term degradation from short-term market dynamics.
problem Intractable computation due to timescale mismatch between battery degradation and market dynamics.
method Approximate dynamic programming with value function approximation and pseudo-time encoding.
result Policy outperforms benchmarks in real-time market scenarios.
This paper introduces a novel framework for designing fair and sustainable unemployment benefits, grounded in cooperative game theory and real-time fiscal policy. The labor market is modeled as a coalitional game, where a random subset of participants is employed, generating stochastic economic output. To ensure fairne…
In the NeurIPS 2018 Artificial Intelligence for Prosthetics challenge, participants were tasked with building a controller for a musculoskeletal model with a goal of matching a given time-varying velocity vector. Top participants were invited to describe their algorithms. In this work, we describe the challenge and pre…
Federated Learning tackles limited user participation with a new risk-aware approach.
problem Limited availability of users in federated learning environments.
method Random Access Model (RAM) and Conditional Value-at-Risk (CVaR) to design a risk-aware federated learning algorithm.
result The proposed approach achieves significantly improved performance under various setups compared to standard federated learning.
The paper tackles personalized policy learning from diverse data sources in a federated setting.
problem Learning personalized decision policies from observational bandit feedback across multiple heterogeneous data sources.
method Introduces a novel regret analysis for distinguishing global and local regret, and presents a federated policy learning algorithm using local policies trained with doubly robust offline policy evaluation strategies.
result Establishes finite-sample upper bounds on global and local regret, characterizing them by source heterogeneity and distribution shift.
A small investor provides liquidity at the best bid and ask prices of a limit order market. For small spreads and frequent orders of other market participants, we explicitly determine the investor's optimal policy and welfare. In doing so, we allow for general dynamics of the mid price, the spread, and the order flow, …
Optimizes pension mix of PAYGO, EET, and individual savings.
problem Balancing PAYGO, EET, and individual savings in funded pension schemes.
method Solves a Nash equilibrium between pension participants and government, considering age-dependent preferences and optimal asset allocation.
result Identifies critical ages and optimal contribution rates for maximizing overall utility.
Paper proposes a framework for token economy simulation and wealth distribution.
problem Simulation and regulation of token economies.
method Formal analysis framework for tokenomics, defining mechanisms for wealth distribution and stability.
result Algorithmic regulatory controls for token economies to achieve desired wealth distribution.
This work models market regimes using CTMSTOU and simulates trading policies.
problem Defining and understanding market regimes in finance.
method Discrete event time multi-agent market simulation with CTMSTOU model.
result Illustrates the importance of regime-awareness in trading policies.
New estimator improves policy evaluation in resource allocation RCTs.
problem Difficulty in evaluating policies optimizing limited resource allocation through RCTs.
method Proposes a novel estimator involving retrospective reshuffling of participants across experimental arms.
result The new estimator provides more accurate policy evaluations than common methods.
Deep reinforcement learning (DRL) is a booming area of artificial intelligence. Many practical applications of DRL naturally involve more than one collaborative learners, making it important to study DRL in a multi-agent context. Previous research showed that effective learning in complex multi-agent systems demands fo…
We consider a dynamic pricing problem for repeated contextual second-price auctions with multiple strategic buyers who aim to maximize their long-term time discounted utility. The seller has limited information on buyers' overall demand curves which depends on a non-parametric market-noise distribution, and buyers may …
UBI model proves financial equilibrium exists.
problem Proving existence of financial equilibrium with UBI.
method Backward stochastic differential equation (BSDE) approach.
result Equilibrium exists in UBI model.
We investigate the combination of actor-critic reinforcement learning algorithms with uniform large-scale experience replay and propose solutions for two challenges: (a) efficient actor-critic learning with experience replay (b) stability of off-policy learning where agents learn from other agents behaviour. We employ …
AI-driven tax policies improve economic equality and productivity.
problem Lack of appropriate economic data and limited opportunity to experiment.
method Two-level deep reinforcement learning approach to learn dynamic tax policies from observational data.
result AI-driven tax policies improve the trade-off between equality and productivity by 16%.
Optimizes cryptocurrency exchanges' risk management by reducing positions based on leverage.
problem Managing risk in cryptocurrency futures exchanges during large price moves.
method Formulates ADL as an optimization problem to minimize risk of loss, using a water-filling rule to equalize leverage.
result The optimal ADL policy minimizes maximum leverage among participants, providing a transparent and implementable benchmark.
Paper constructs a CRRIX index to assess cryptocurrency market risks from regulatory changes.
problem Lack of indices quantifying regulatory risks in cryptocurrencies.
method CRRIX index based on news coverage frequency, using Latent Dirichlet Allocation and Hellinger distance.
result CRRIX successfully captures major policy-changing moments and synchronizes with market volatility.
Optimizes capital structure for life insurance companies with surplus participation.
problem Determining the optimal participation rate in life insurance contracts.
method Adapted Leland's dynamic capital structure model to life insurance context.
result Optimal participation rate is highly sensitive to contract duration and tax rate.
Paper proposes MMVFL for multi-class VFL with multiple participants.
problem Privacy-preserving multi-class VFL with multiple participants.
method Extends multi-view learning to enable label sharing among multiple VFL participants.
result MMVFL effectively shares label information among multiple VFL participants and matches multi-class classification performance.
The paper uncovers two key laws of market impact influenced by volume and participation rate.
problem Understanding the roles of volume and participation rate in market price response.
method Extending the no arbitrage approach to include sophisticated market participants, deriving price dynamics from order flow dynamics.
result Recovery of two square root laws governing market impact.
Study uses contextual bandits to optimize charity exposure in donation solicitation.
problem Optimizing charity exposure in donation solicitation using survey responses.
method Adaptive experiment design to balance cumulative regret minimization and simple regret minimization.
result Adaptive experimentation yields better policy learning outcomes than uniform randomization.
CADR estimator improves inference for contextual bandit data.
problem Valid inference on contextual bandit data.
method CADR estimator for policy value, addressing adaptive data collection challenges.
result CADR provides correct coverage of confidence intervals.
Local adaptation improves federated learning models.
problem Improving accuracy of federated learning models on non-iid data.
method Local adaptation techniques (fine-tuning, multi-task learning, knowledge distillation).
result Participants benefit from local adaptation, improving federated model accuracy.
Unified analysis of FL with arbitrary client participation.
problem Challenges of intermittent client availability and efficiency in FL.
method Introduces a generalized FedAvg and novel analysis capturing client participation.
result Unified convergence upper bounds for various participation patterns.
Paper proposes auction method for smart derivatives to avoid disputes.
problem Disputes over derivative liquidation processes in smart contracts.
method Defines an auction type resolution for smart derivatives.
result Proposes a beneficial method for smart derivatives participants.
Lapse-supported life insurance exacerbates adverse selection risks.
problem Lapse-supported life insurance increases adverse selection costs.
method Modeling 'Term to 100' contracts and analyzing three methods of managing lapse surplus.
result Adverse selection losses can be almost unlimited under certain conditions.
Study examines dependence of extreme electricity prices in Australian markets.
problem Understanding and managing risks of extreme price outcomes in Australian electricity markets.
method Examined extremal dependence using extremograms for 5-minute and 30-minute price data.
result Persistence and dependence of extreme prices are influenced by market structure and renewable energy share.
New approach tackles non-Markovian behavior in maternal health programs.
problem Improving adherence and engagement in maternal and child healthcare programs.
method Extending RMABs to non-Markovian settings, using time-series forecasting and TARI policy.
result Significant increase in engagement and content listened compared to existing methods.
Paper proposes a risk-averse approach to energy storage price arbitrage using conformal uncertainty quantification.
problem Inherent volatility and uncertainty of real-time electricity prices create financial risks for storage arbitrage.
method Two-layer prediction model with conformal uncertainty quantification for high coverage of real-time price uncertainty.
result The framework achieves good profit margins with minimal losses, demonstrating effectiveness in real-time market.
The task of dialog management is commonly decomposed into two sequential subtasks: dialog state tracking and dialog policy learning. In an end-to-end dialog system, the aim of dialog state tracking is to accurately estimate the true dialog state from noisy observations produced by the speech recognition and the natural…
FedAMD framework improves federated learning with partial client participation.
problem Data heterogeneity and inactive client updates in partial client participation.
method Anchor sampling divides clients into anchor and miner groups, using large and small batches respectively.
result FedAMD achieves faster convergence and improved model performance compared to state-of-the-art methods.
TT-DAC-PS: A deterministic actor-critic approach for optimal trade execution
problem Optimal execution of large stock sell programs
method Twin-Target Deterministic Actor-Critic with Policy Smoothing
result Reduces mean implementation shortfall percentage
Paper tackles unknown participation in FL, proposing FedAU for better performance.
problem Unknown participation statistics in federated learning impact performance.
method Adapting aggregation weights in FedAvg based on participation history.
result FedAU converges to optimal solution and has desirable properties.
Bayesian framework improves trading robustness against market shifts.
problem Insufficient robustness and overfitting in trading models.
method Bayesian Robust Framework integrating macro-conditioned GAN and adversarial learning.
result Framework outperforms state-of-the-art models in diverse financial instruments.
The purpose of this article is to introduce, analyze and compare two performance participation methods based on a portfolio consisting of two risky assets: Option-Based Performance Participation (OBPP) and Constant Proportion Performance Participation (CPPP). By generalizing the provided guarantee to a participation in…
Extended model ensures long-term survival of traders in limited stock market participation.
problem Limited stock market participation and survival of traders over long periods.
method Extended Basak and Cuoco (1998) model with different time-preference coefficients.
result Parameter restrictions ensure long-term survival of traders.
Data poisoning attacks can severely degrade FL models, especially targeting specific classes.
problem Data poisoning attacks against federated learning systems.
method Demonstrated targeted attacks on FL systems, analyzed attack longevity, and proposed a defense strategy.
result Data poisoning attacks can cause substantial drops in classification accuracy and recall with a small percentage of malicious participants.
In this paper, we present a new task that investigates how people interact with and make judgments about towers of blocks. In Experiment~1, participants in the lab solved a series of problems in which they had to re-configure three blocks from an initial to a final configuration. We recorded whether they used one hand …
Proves existence of equilibrium in limited participation economy.
problem Existence of an equilibrium in an economy with limited financial market access.
method Proves global existence of Radner equilibrium using BSDEs with unique solution.
result Proves existence of Radner equilibrium with limited participation.
Insurance companies often include very long-term guarantees in participating life insurance products, which can turn out to be very valuable. Under a guaranteed annuity options (G.A.O), the insurer guarantees to convert a policyholder's accumulated funds to a life annuity at a fixed rated when the policy matures. Both …
In the Pioneer 100 (P100) Wellness Project (Price and others, 2017), multiple types of data are collected on a single set of healthy participants at multiple timepoints in order to characterize and optimize wellness. One way to do this is to identify clusters, or subgroups, among the participants, and then to tailor pe…
Flexible device participation improves federated learning convergence.
problem Strict device participation limits federated learning reach.
method Analytical results and new aggregation scheme for flexible participation.
result Convergence improved with flexible device participation.