Discrete-time games reveal payoffs after both players stop, leading to new equilibrium strategies.
problem Non-zero-sum stopping games with delayed payoff revelation.
method Analyzes simultaneous and sequential stopping strategies, proving Nash equilibria in both cases.
result Existence of Nash equilibria in mixed and pure stopping strategies.
A new algorithm tackles delayed combinatorial semi-bandit with causal relations.
problem Optimizing decisions in a non-stationary environment with delayed and causally related rewards.
method Formalized as a non-stationary delayed combinatorial semi-bandit problem, the approach models causal relations with a directed graph in a stationary structural equation model. The agent learns these relations from delayed feedback to optimize decisions.
result Proved a regret bound for the proposed algorithm's performance.
New model for music streaming recommends songs based on past play history.
problem Nonstationary stochastic bandit model with delay-dependent rewards.
method Ranking policies approximating optimal policy with bounded regret.
result Algorithm with O ~ ( k T ) \widetilde{\mathcal{O}}\big(\!\sqrt{kT}\big) O ( k T ) regret and O ( k ln ln T ) \mathcal{O}\big(k\ln\ln T\big) O ( k ln ln T ) switches. Signature payoffs price complex derivatives accurately.
problem Pricing complex derivatives like options.
method Signature of price path for continuous payoffs.
result Signature payoffs can price various derivatives accurately.
Method constructs CFMMs matching desired payoffs.
problem Creating CFMMs with specific payoff functions.
method Uses convex analysis and Fenchel conjugacy.
result Every concave, nonnegative, nondecreasing, 1-homogeneous payoff has a corresponding convex CFMM.
Optimal payoff choice constrained by Bregman-Wasserstein divergence.
problem Maximizing utility under a deviation constraint from a benchmark.
method Solving the problem using Bregman-Wasserstein divergence with a convex function φ.
result Provided the optimal payoff choice in this setting.
We study the optimal timing of derivative purchases in incomplete markets. In our model, an investor attempts to maximize the spread between her model price and the offered market price through optimally timing her purchase. Both the investor and the market value the options by risk-neutral expectations but under diffe…
Optimal portfolio yields a digital option payoff.
problem Portfolio optimization under generalized dual theory of choice.
method Characterized optimal solution and derived it in closed form.
result Payoff is a digital option that yields in-the-money payoff in good market scenarios.
This paper modifies the Ait-Sahalia model to better describe interest rate behaviors.
problem Inadequate specifications of the original Ait-Sahalia model to explain various interest rate phenomena.
method Proposes a modified hybrid Poisson-jump Ait-Sahalia model and uses truncated EM techniques for numerical approximation.
result Validates the modified model using Monte Carlo simulations for bond and barrier option payoffs.
Study finds cheapest possible payoff under ambiguity, linking to maxmin expected utility.
problem Finding cost-efficient payoffs in uncertain market conditions.
method Developed a new concept of robust cost-efficient payoff and linked it to maxmin expected utility.
result Solutions to maxmin robust expected utility are robust cost-efficient.
The paper uncovers the impact of price and payoff autocorrelations in multi-period asset pricing models.
problem Hidden dependence of asset pricing models on price and payoff autocorrelations.
method Obtained approximations of the basic pricing equation describing various parameters.
result Valid results for other pricing models like ICAPM and APM.
Develops a stochastic approach to financial market delays.
problem Modeling delays in financial markets with multiple assets.
method Introduces a general stochastic framework for information and order execution delays.
result Delayed markets maintain fundamental asset pricing theorems and no asymptotic free lunch condition.
Paper tackles action delays in reinforcement learning, proposing a delay-aware framework.
problem Action delays degrade reinforcement learning performance in real-world systems.
method Formal definition of delay-aware MDP, transformation into standard MDP with augmented states, delay-aware model-based reinforcement learning framework.
result Proposed framework is more efficient in training and transferable between systems with various delay durations.
New method uses neural networks for better financial hedging.
problem Spanning multi-asset payoffs with vanilla options.
method One-hidden-layer feedforward neural networks for numerical solution.
result Better hedging results with neural networks compared to single-asset approaches.
Paper shows how to replicate payoffs without oracles in CFMMs.
problem Replicating payoffs without oracles in CFMMs.
method Using liquidity provider shares in CFMMs to match any monotonic payoff.
result Explicit method and formula for trading functions and earnings.
New algorithm tackles delayed feedback in Lipschitz bandits with sublinear regret.
problem Delayed feedback in Lipschitz bandits.
method Design of algorithms for bounded and unbounded stochastic delays.
result Sublinear regret guarantees for both bounded and unbounded delays.
New algorithms ensure fair selection in combinatorial semi-bandit with unrestricted delays.
problem Fair selection in stochastic combinatorial semi-bandit with delayed feedback.
method Introduced merit-based fairness constraints and new bandit algorithms for reward and fairness.
result Achieved sublinear expected reward and fairness regrets with dependence on delay distribution quantiles.
The study analyzes how wartime controls influenced zaibatsu stock prices in Japan.
problem How wartime economic controls affected zaibatsu stock prices in Japan.
method Developed a four-portfolio asset-pricing model and used a CAPM-AR(p)-SV event-study framework.
result Wartime economic controls influenced stock prices through financing wedges and zaibatsu affiliation.
Banker-OMD improves online learning with delayed feedback.
problem Handling delayed feedback in online learning.
method Generalized Online Mirror Descent (OMD) framework.
result Achieves nearly-optimal performance in three bandit scenarios.
New algorithm handles delayed feedback robustly, reducing regret without knowing delay bounds.
problem Bandits with variably delayed feedback, especially excessive delays.
method Implicit exploration scheme, adaptive skipping, drifted regret control.
result Can tolerate arbitrary excessive delays up to order T, reducing regret.
Nonparametric pricing and hedging of exotic derivatives using signature payoffs.
problem Pricing and hedging exotic derivatives accurately and efficiently.
method Introducing signature payoffs and using them to approximate and price exotic derivatives nonparametrically.
result Signature payoffs enable accurate and computationally tractable pricing and hedging of exotic derivatives.
Proposes a nonparametric model for predicting conversion rates with delayed feedback.
problem Predicting conversion rates with time delays and unknown distribution.
method Nonparametric delayed feedback model without assuming a specific distribution.
result The proposed model outperforms existing methods in conversion rate prediction.
Gradient descent with delayed updates converges faster with noise, even when delays are significant.
problem Analyzing convergence of gradient descent with delayed gradients and stochastic noise.
method Novel technique using generating functions for convergence analysis.
result Convergence bounds show that stochastic noise mitigates the negative effects of delays, improving performance.
New algorithm for multiarmed bandits with variable, unbounded delays achieves similar regret bounds.
problem Variable, unbounded delays in multiarmed bandits.
method Introduces a new algorithm that skips rounds with excessively large delays and uses a doubling scheme.
result Achieves the same regret bound as Exp3 with variable, unbounded delays.
Paper tackles delays in multi-agent reinforcement learning, improving performance.
problem Challenges in reinforcement learning due to delays in real-world systems.
method Proposes a novel framework for multi-agent reinforcement learning with delays, using Delay-Aware Markov Games and centralized-decentralized training.
result Demonstrates significant improvement in performance with delay-aware multi-agent reinforcement learning.
New algorithm adapts to unknown smoothness in contextual bandits.
problem Adapting to unknown smoothness in non-parametric multi-armed bandits.
method Develops a self-similarity condition-based policy to adapt to unknown smoothness.
result Matches known smoothness case's regret rate for differentiable and non-differentiable payoff functions.
Derives a Feynman-Kac formula for a fixed delay CIR model.
problem Modeling financial processes with fixed delay.
method Proves existence and uniqueness of a strong solution for a specific SDDE.
result Derives a Feynman-Kac type formula leading to an affine bond pricing formula.
New bandit problem with delayed, aggregated feedback analyzed.
problem Stochastic K K K -armed bandit problem with delayed, aggregated anonymous feedback. method Developed algorithm matching worst case regret of non-anonymous problem.
result Regret increase can be maintained in the harder delayed, aggregated anonymous feedback setting.
Study on synchronization in financial markets with time delays.
problem Understanding market dynamics and synchronization in financial systems with time delays.
method Examined a system of coupled non-linear delay-differential equations, linearized for small delays, and analyzed collective dynamics using bifurcation diagrams and numerical solutions.
result Demonstrated that limit cycles can be maintained in coupled N-asset models with appropriate parameterization, leading to market synchronization.
BayTiDe discovers time-delayed differential equations from noisy data.
problem Discovering time-delayed differential equations from data with large delays and noise.
method Bayesian inference with a sparsity-promoting prior.
result BayTiDe accurately identifies time-delayed differential equations with accuracy proportional to data resolution.
New algorithm tackles stochastic bandits with varying arm-dependent delays.
problem Applying existing algorithms to stochastic delayed bandit settings is restricted by strong assumptions on delay distributions.
method Proposes a simple UCB-based algorithm called PatientBandits that weakens assumptions on delay distributions.
result Provides bounds on regret and performance lower bounds for the PatientBandits algorithm.
Optimized options portfolio with a specific payoff function.
problem Optimizing an options portfolio with a fixed payoff function.
method Formulated as an integer linear programming problem, including an objective payoff function and constraints.
result Optimum solution for European call and put options on Taiwan Futures Exchange.
Model analyzes how delayed information impacts option pricing.
problem Effects of delayed information on option pricing.
method Binomial model, closed form formula for convex contingent claims, convergence analysis.
result Delayed information exaggerates the volatility smile.
TSMB handles time delays in multivariate time series data.
problem Varying time delays in multivariate time series data complicate predictions.
method Time Series Model Bootstrap (TSMB) framework for nonparametric time delay estimation.
result TSMB improves model performance in dynamic data environments.
Delayed-RNN approximates stacked and bidirectional RNNs.
problem Improving RNN expressiveness and representational capacity.
method Weight-constrained delayed-RNN, equivalent to stacked-RNNs, with partial acausality.
result Delayed-RNN can approximate stacked and bidirectional RNNs, outperforming them in some tasks.
The paper tackles sequential learning with Gaussian payoffs and side observations, providing lower bounds and algorithms.
problem Sequential learning with Gaussian payoffs and side information.
method Non-asymptotic lower bounds and algorithms for minimizing regret.
result Proved non-asymptotic lower bounds and provided algorithms achieving these bounds.
New algorithm reduces regret in delayed feedback generalised linear bandits.
problem Regret in delayed feedback generalised linear bandits.
method Adaptation of optimistic algorithm to delayed feedback.
result Achieves a regret bound independent of the horizon's delay penalty.
Adapts Exp3 to adversarial bandits with delays and data.
problem Adversarial multi-armed bandits with delayed feedback.
method Tuned Exp3 variants with step-size adaptation and implicit exploration.
result Optimal regret bounds of log ( K ) ( T K + D ) \sqrt{\log(K)(TK + D)} log ( K ) ( T K + D ) with high probability. Capacity-Constrained Online Convex Optimization with Delayed Feedback
problem Online learning with delayed feedback under a hard capacity constraint
method Reduction to a delayed and weighted OCO problem using a scheduler
result First regret guarantees for capacity-constrained OCO under convex and strongly convex losses
New algorithm tackles non-stationary delayed feedback in recommender systems.
problem Challenges in learning from delayed feedback in non-stationary environments.
method Developed a UCRL-based algorithm for non-stationary, delayed bandits with intermediate observations.
result Sublinear regret guarantees for the proposed algorithm in non-stationary delayed environments.
Develops a new method for robust risk measurement by averaging nearby payoffs.
problem Measuring risk under uncertainty with a focus on robustness.
method Averaging nearby payoffs weighted by a chosen metric.
result The method leads to a convex risk measure and provides stability under large neighborhoods.
PCTS optimizes noisy, delayed, multi-fidelity feedbacks in black-box optimization.
problem Optimizing unknown functions with noisy, delayed, and multi-fidelity feedbacks.
method ProCrastinated Tree Search (PCTS) with DUCB1 and DUCBV algorithms.
result PCTS achieves better regret bounds for delayed, noisy, and multi-fidelity feedbacks.
Study online learning with delays and capacity constraints, achieving optimal regret bounds.
problem Online learning with delays and capacity constraints.
method Novel scheduling and preemptive techniques, matching upper and lower bounds.
result Achieves optimal regret bounds across all capacity levels.
Study market delay effects on contingent claims pricing.
problem Delayed market information impacts contingent claims pricing.
method Analyzes Black-Scholes and binomial models with delay.
result Scaling limit of super-replication prices equals G-expectation.
Agent optimizes perpetual contract liquidation with transaction costs and risk.
problem Optimizing perpetual contract liquidation with transaction costs and risk.
method Solving stochastic control problem for optimal trading strategy.
result Closed-form expression and approximations for optimal strategy.
Multi-armed bandit problems are the most basic examples of sequential decision problems with an exploration-exploitation trade-off. This is the balance between staying with the option that gave highest payoffs in the past and exploring new options that might give higher payoffs in the future. Although the study of band…
New Async-SGD and Async-SGDI methods converge for non-convex problems with unbounded delays.
problem Improving convergence of asynchronous stochastic gradient descent with unbounded delays in non-convex learning.
method Developed Async-SGD and Async-SGDI methods for non-convex optimization with unbounded gradient delays, proving convergence rates and establishing a unifying sufficient condition.
result Proved o ( 1 / k ) o(1/\sqrt{k}) o ( 1/ k ) convergence rate for Async-SGD and o ( 1 / k ) o(1/k) o ( 1/ k ) for Async-SGDI. Online learning with delayed feedback has received increasing attention recently due to its several applications in distributed, web-based learning problems. In this paper we provide a systematic study of the topic, and analyze the effect of delay on the regret of online learning algorithms. Somewhat surprisingly, it t…