The paper models delayed event occurrences in insurance data.
problem Systematic underestimation of event occurrence due to observation delays.
method Modeling the time between event occurrence and observation, considering event day and calendar effects.
result A granular model for event observation delay heterogeneity.
Neural Laplace Control tackles offline RL for continuous-time delayed systems with irregular observations.
problem Offline reinforcement learning problems involving continuous-time environments with delays and irregular observations.
method Combines a Neural Laplace dynamics model with a model predictive control (MPC) planner.
result Achieves near expert policy performance on continuous-time delayed environments.
New algorithm tackles non-stationary delayed feedback in recommender systems.
problem Challenges in learning from delayed feedback in non-stationary environments.
method Developed a UCRL-based algorithm for non-stationary, delayed bandits with intermediate observations.
result Sublinear regret guarantees for the proposed algorithm in non-stationary delayed environments.
New bandit problem with delayed, aggregated feedback analyzed.
problem Stochastic K K K -armed bandit problem with delayed, aggregated anonymous feedback. method Developed algorithm matching worst case regret of non-anonymous problem.
result Regret increase can be maintained in the harder delayed, aggregated anonymous feedback setting.
Proposes neural delay differential equations for stable system identification with partially observed states.
problem Learning stable models for systems with partial or delayed observations.
method Augments states with history, uses neural delay differential equations, and ensures stability through time delay analysis.
result The approach ensures stability of learned models for partially observed systems.
New algorithm for multiarmed bandits with variable, unbounded delays achieves similar regret bounds.
problem Variable, unbounded delays in multiarmed bandits.
method Introduces a new algorithm that skips rounds with excessively large delays and uses a doubling scheme.
result Achieves the same regret bound as Exp3 with variable, unbounded delays.
New algorithms for linear bandits with delayed feedback, achieving optimal regret bounds.
problem Delayed and partially observable feedback in stochastic linear bandits.
method Formalized as stochastic delayed linear bandit, proposed two algorithms t O T F L i n U C B { t OTFLinUCB} tO T F L in U C B and t O T F L i n T S { t OTFLinTS} tO T F L in T S . result Proved optimal i l d e O ( d T ) ilde O(\smash{d\sqrt{T}}) i l d e O ( d T ) bounds on the regret for t O T F L i n U C B { t OTFLinUCB} tO T F L in U C B . Capacity-Constrained Online Convex Optimization with Delayed Feedback
problem Online learning with delayed feedback under a hard capacity constraint
method Reduction to a delayed and weighted OCO problem using a scheduler
result First regret guarantees for capacity-constrained OCO under convex and strongly convex losses
New algorithm handles delayed feedback robustly, reducing regret without knowing delay bounds.
problem Bandits with variably delayed feedback, especially excessive delays.
method Implicit exploration scheme, adaptive skipping, drifted regret control.
result Can tolerate arbitrary excessive delays up to order T, reducing regret.
Optimal strategy for reinforcement learning with delayed observations.
problem Delayed state observation in reinforcement learning.
method Combines augmentation method and upper confidence bound approach.
result Minimax optimal regret bound of i l d e O ( H D max S A K ) ilde{\mathcal{O}}(H \sqrt{D_{\max} SAK}) i l d e O ( H D m a x S A K ) . Optimal policy for multi-hypothesis testing with controlled sensing to minimize delay and error.
problem Minimizing delay in multi-hypothesis testing with controlled sensing.
method Designing a policy to control the delay while ensuring error probability constraint.
result Policy achieves information-theoretic lower bound on expected delay asymptotically.
Study uses randomized allocation for delayed rewards in multi-armed bandits.
problem Delayed rewards in contextual multi-armed bandits.
method Randomized allocation with nonparametric estimation.
result Strongly consistent strategy for delayed rewards.
Study explores strategies for randomized allocation in delayed rewards bandits.
problem Understanding the exploration-exploitation tradeoff in randomized strategies with delayed rewards.
method Examines two strategies: updating exploration sequence at every time point vs. updating only when a new reward is observed.
result The strategy updating only when a new reward is observed leads to strong consistency in allocation for a wider scope of situations.
Algorithm improves RL by discovering delayed causal relations.
problem Improving data-efficiency and interpretability in RL.
method Predicts observations with Markov assumption, introduces hidden variables to explain past events.
result Significantly improves RL performance on simulated and real tasks.
Develops a stochastic approach to financial market delays.
problem Modeling delays in financial markets with multiple assets.
method Introduces a general stochastic framework for information and order execution delays.
result Delayed markets maintain fundamental asset pricing theorems and no asymptotic free lunch condition.
Adapts Exp3 to adversarial bandits with delays and data.
problem Adversarial multi-armed bandits with delayed feedback.
method Tuned Exp3 variants with step-size adaptation and implicit exploration.
result Optimal regret bounds of log ( K ) ( T K + D ) \sqrt{\log(K)(TK + D)} log ( K ) ( T K + D ) with high probability. New algorithm tackles stochastic bandits with varying arm-dependent delays.
problem Applying existing algorithms to stochastic delayed bandit settings is restricted by strong assumptions on delay distributions.
method Proposes a simple UCB-based algorithm called PatientBandits that weakens assumptions on delay distributions.
result Provides bounds on regret and performance lower bounds for the PatientBandits algorithm.
New algorithm tackles delayed feedback in Lipschitz bandits with sublinear regret.
problem Delayed feedback in Lipschitz bandits.
method Design of algorithms for bounded and unbounded stochastic delays.
result Sublinear regret guarantees for both bounded and unbounded delays.
Study improves risk evaluation timing with right-censored reporting delays.
problem Improving risk evaluation under short observation windows due to administrative censoring.
method Jointly models parametric hazards for event and reporting processes, uses Monte Carlo expectation-maximization algorithm, and proposes transfer-learning procedure.
result Improves accuracy of timely risk evaluation under administrative censoring.
A controller learns to control a nonlinear plant with unknown model and partial observation using continuous deep Q-learning.
problem Designing a controller for a nonlinear plant with unknown model and partial sensor observation under network delays.
method Continuous deep Q-learning applied to an extended state including past control inputs and outputs.
result The controller can learn a robust control policy to network delays with partial sensor observation.
Enhances forecasting of complex systems using FKMD.
problem Forecasting high-dimensional dynamical systems with unknown features.
method Featurized Koopman Mode Decomposition (FKMD) using delay embedding and learned Mahalanobis distance.
result Improves prediction accuracy for various complex systems.
Dual learning algorithm addresses delayed conversions in CVR prediction.
problem Challenges in predicting conversion rate due to delayed feedback.
method Proposes two unbiased estimators and a dual learning algorithm.
result Demonstrates practical value of the proposed approach through empirical evaluations.
Adapts two algorithms for online learning with delayed rewards.
problem Online learning with delayed rewards in generalized linear contextual bandits.
method Modifies upper confidence bounds and Thompson sampling algorithms for delayed rewards.
result Both algorithms can be made robust to delays, improving their performance.
BayTiDe discovers time-delayed differential equations from noisy data.
problem Discovering time-delayed differential equations from data with large delays and noise.
method Bayesian inference with a sparsity-promoting prior.
result BayTiDe accurately identifies time-delayed differential equations with accuracy proportional to data resolution.
Measures time-delay embedding for noisy, sparse data.
problem Applying Takens' embedding theorem to real-world, noisy data.
method Formulated a measure-theoretic generalization of the embedding theorem, using optimal transport.
result Reconstructed full state of dynamical systems from time-lagged partial observations robust to noise and sparsity.
Paper tackles delays in multi-agent reinforcement learning, improving performance.
problem Challenges in reinforcement learning due to delays in real-world systems.
method Proposes a novel framework for multi-agent reinforcement learning with delays, using Delay-Aware Markov Games and centralized-decentralized training.
result Demonstrates significant improvement in performance with delay-aware multi-agent reinforcement learning.
Time-delayed embeddings avoid self-intersections for high enough delay.
problem Analyzing self-intersections in time-delayed embeddings.
method Study of time-delayed coordinate maps for diffeomorphisms on compact manifolds.
result For high enough delay, time-delayed embeddings avoid self-intersections almost everywhere.
Study utility indifference pricing with delayed investment information in a Bachelier model.
problem Investment decisions based on delayed information in a Bachelier model.
method Developed discrete-time duality and used techniques from [7] to compute scaling limits.
result Utility indifference prices scaling limit for vanishing delay with quadratic penalty.
Paper uses queue theory to model financial signals with relativistic delay.
problem Relativistic delay in financial trading signals.
method Modified M/M/G queue theory.
result Describes propagation of trading signals with finite velocity.
Proposes a nonparametric model for predicting conversion rates with delayed feedback.
problem Predicting conversion rates with time delays and unknown distribution.
method Nonparametric delayed feedback model without assuming a specific distribution.
result The proposed model outperforms existing methods in conversion rate prediction.
Helps visually impaired users make better decisions by adjusting their observations.
problem Systematic biases in users' perception and processing of visual information.
method Synthesizes new observations based on true observations to correct user biases.
result Significant improvement in task performance for users with various biases.
Efficiently estimates conversion probability in online display advertising.
problem Estimating conversion probability with delays and large data sets.
method Compromise estimator combining logistic regression and joint model.
result Computational efficiency with less bias than previous methods.
Paper proposes GrokTransfer to eliminate delayed generalization in neural networks.
problem Delayed generalization in neural networks, compromising predictability and efficiency.
method Trains a smaller, weaker model to reach a nontrivial test performance, then uses its learned input embedding to initialize the stronger model.
result GrokTransfer enables the target model to generalize directly without delay, across various tasks.
DASA speeds up SA with delayed agents, achieving N-fold speedup.
problem Speeding up Stochastic Approximation with asynchronous delays.
method DASA: Delay-Adaptive Multi-Agent Stochastic Approximation algorithm.
result First algorithm with convergence rate dependent on mixing time and average delay.
Paper tackles dueling bandits with delayed feedback, revealing preference bias.
problem Real-world dueling bandit applications often face delays in feedback.
method Introduces biased dueling bandit problem with stochastic delayed feedback, presents two algorithms.
result Two algorithms achieve optimal regret bounds for dueling bandit problems with delay.
EviTrack improves sequential prediction in delayed disambiguation scenarios.
problem Challenges in sequential prediction with delayed disambiguation where early observations are ambiguous.
method EviTrack operates over latent trajectories, applying evidence- and likelihood-ratio-based selection to delay commitment until supported by data.
result EviTrack outperforms sampling-based baselines in a controlled synthetic benchmark, achieving faster post-disambiguation recovery.
New algorithm for multi-armed bandits with delayed, partially observed rewards.
problem Sequential decision-making with delayed feedback.
method Proposed multi-armed bandits with generalized temporally-partitioned rewards, introducing β-spread property.
result Upper bound on performance of TP-UCB-FR-G algorithm improves state of the art.
New algorithm optimizes noisy function evaluations with delayed feedback.
problem Optimizing unknown functions from noisy, delayed feedback.
method Kernel bandit problem with stochastically delayed feedback.
result Proposes an algorithm with improved regret bound for non-smooth kernels.
New RL algorithm handles delayed feedback with posterior sampling.
problem Challenges of delayed feedback in reinforcement learning with linear function approximation.
method Posterior sampling with delayed feedback for value-based RL.
result Achieves optimal regret guarantee with improved computational efficiency.
A system estimates delayed context for online scoring using convex optimization.
problem Estimating agent scores with delayed context information.
method Online convex game between agent and system; leveraging correlation function.
result Error in score estimate is small if online convex game has low regret.
Improved online convex optimization with delayed feedback using curvature.
problem Online convex optimization with curved losses and delayed feedback.
method Variant of follow-the-regularized-leader and Online Newton Step algorithm with adaptive learning rate.
result Regret bounds of order min { σ max ln T , d t o t } \min\{σ_{\max}\ln T, \sqrt{d_{\mathrm{tot}}}\} min { σ m a x ln T , d tot } for exp-concave losses. Study non-oblivious adversarial bandits with delayed feedback and propose algorithms with improved regret bounds.
problem Adversarial bandit problem with delayed, composite anonymous feedback.
method Propose wrapper algorithm for non-oblivious delay setting, achieving o ( T ) o(T) o ( T ) policy regret. result Achieve o ( T ) o(T) o ( T ) policy regret for many adversarial bandit problems with bounded memory loss sequences. Improved algorithm for bandits with delayed feedback, combining adversarial and stochastic performance.
problem Adversarial and stochastic multiarmed bandits with delayed feedback.
method Modified Zimmert and Seldin's algorithm with near-optimal regret guarantees.
result Near-optimal regret guarantees in both adversarial and stochastic settings.
New method disentangles latent variables in nonstationary data.
problem Disentangling latent variables in nonstationary sequential data.
method NCTRL framework exploiting Markov assumption and temporal structure.
result Independent latent components can be recovered from nonlinear mixture without auxiliary variables.
The paper defines and analyzes Poissonian occupation times for negative Lévy processes.
problem Analyzing the time spent below zero for Lévy processes with interruptions.
method Introduces Poissonian occupation times for spectrally negative Lévy processes.
result Extends results on continuous observation to interrupted observation.
New algorithm for decentralized online convex optimization with unknown delays.
problem Decentralized online convex optimization with unknown, time-varying feedback delays.
method Proposes a novel algorithm that incorporates adaptive learning rates and gossip-based delay estimation.
result Achieves improved regret bounds of O(N √d tot + N √T (1-σ^2)^(1/4)) and O(N δmax ln T α) for different settings.
Proposes a method to train classifiers with delayed feedback using a time window.
problem Training classifiers with delayed feedback that can be biased due to delayed user actions.
method Uses a time window to select samples for training, constructs unbiased empirical risk from all samples.
result Improves classifier performance by using all samples with a time window assumption.
This paper tackles delayed feedback in continuous training for CTR prediction, improving model performance by 3%.
problem Delayed feedback in CTR prediction leads to inferior performance and user experience.
method Comparing 5 loss functions and models in offline and online settings.
result Proposed methods outperform previous state-of-the-art by 3% relative cross entropy (RCE).