New approach to portfolio optimization shows entropy regularization is ineffective.
problem Entropy regularization in mean-variance portfolio optimization under drift uncertainty.
method Combining Bayesian filtering and stochastic policy optimization.
result Entropy regularization does not accelerate learning about unknown drift.
This paper identifies drift Lipschitz budget K as key to diffusion policy expressivity and statistical trade-offs.
problem Understanding and maximizing the expressivity of diffusion policies while managing statistical limitations.
method Identifying drift Lipschitz budget K as central, quantifying expressivity and statistical behavior, proving lower bounds, and providing practical implementation guidelines.
result Balancing expressivity and statistical complexity yields a finite-sample performance gap, with rates depending on sample size and drift type.
AMUSE uses reinforcement learning to predict optimal model updates.
problem Concept drift weakens model performance over time.
method Reinforcement learning in a simulated environment.
result AMUSE proactively recommends updates based on performance improvements.
This paper tackles robust policy learning under concept drifts, improving upon existing methods.
problem Tackles robust policy learning under concept drifts, improving upon existing methods.
method Develops a doubly-robust estimator and a learning algorithm to maximize policy value within a given policy class.
result The proposed algorithm achieves sub-optimality gap of the order κ ( Π ) n − 1 / 2 κ(Π)n^{-1/2} κ ( Π ) n − 1/2 , demonstrating substantial improvement over existing benchmarks. Bayesian investor learns unknown asset drift, trades mean-variance optimal portfolio, but policy is robust to observation model distortion.
problem Bayesian portfolio selection with observation model distortion
method Robust Bayesian portfolio selection
result Robust policy and its price are closed form, with price of robustness half the variance of the non-robust investor's loss.
Bayesian Markowitz portfolio problem shows entropy regularization is ineffective.
problem Entropy regularization in Bayesian Markowitz portfolio optimization.
method Combines continuous-time Bayesian filtering with stochastic policy optimization.
result Entropy regularization does not accelerate learning of unknown drift.
Paper achieves ε − 2 ε^{-2} ε − 2 sample complexity for actor-critic methods with minimal assumptions.
problem Achieving ε − 2 ε^{-2} ε − 2 sample complexity for actor-critic methods under minimal assumptions. method Single-loop, single-timescale implementation; coupled Lyapunov drift framework.
result First i l d e O ( ε − 2 ) ilde{\mathcal{O}}(ε^{-2}) i l d e O ( ε − 2 ) sample complexity guarantee for finding an ε ε ε -optimal policy. ESPD improves learning efficiency in sparse reward reinforcement learning.
problem Sparse reward reinforcement learning challenges.
method Evolutionary Stochastic Policy Distillation (ESPD) based on drifted random walk insight.
result High learning efficiency demonstrated in MuJoCo robotics control suite experiments.
Study improves survival analysis for credit risk by accounting for data drift.
problem Survival analysis in credit risk assumes a stationary data-generating process, but real-world data drift affects model performance.
method Proposes a dynamic joint modelling framework integrating longitudinal behavioural markers and hazard formulations, combined with drift-adaptive techniques.
result Proposed model outperforms classical survival models and drift-adaptive learners in various data drift scenarios.
AI governance lagging in finance despite widespread use.
problem Lack of operational governance frameworks for AI in finance.
method Proposes a four-layer framework with computable instantiations.
result Demonstrates the effectiveness of the proposed framework through a case study.
New budget quantifies drift in closed-loop learning, improving reproducibility.
problem Characterizing statistical learning under distributional drift in closed-loop settings.
method Introduces an intrinsic drift budget C T C_T C T quantifying cumulative information-geometric motion of the data distribution. result Proves a drift-feedback bound of order T − 1 / 2 + C T / T T^{-1/2}+C_T/T T − 1/2 + C T / T for prequential reproducibility, up to controlled second-order remainder terms. Variational Proximal Policy Optimization improves reinforcement learning from human feedback.
problem Policy mode collapse and brittle exploration loops in reinforcement learning.
method Particle-based variational inference framework with Mixture-of-Experts architecture.
result Significant improvements in complex reasoning benchmarks.
A dual-learner strategy tracks concept drift in nonstationary data streams.
problem Learning from nonstationary data with abrupt or gradual changes.
method Alternating learners framework with long- and short-memory models.
result Effective tracking and prediction of concept drift in streaming data.
This study develops a dynamic inverse optimization framework to recover hidden, time-varying preferences from observed allocation trajectories.
problem The gap between classical optimization theory and real-world practice, especially in the presence of drift and shocks.
method Dynamic inverse optimization framework using a drift-aware estimator grounded in convex analysis and online learning theory.
result Sharp static and dynamic regret bounds for the framework, demonstrating its responsiveness to gradual drift and sudden shocks.
Central bank optimizes bailout cash injection to limit defaults.
problem Optimizing cash injection to limit defaults in a system of mutual obligations.
method Proved convergence and solved a drift controlled Stefan problem using mean field control and policy gradient methods.
result Optimal strategies involve subsidizing banks with equity values in a time-dependent region.
Develops PromptShift-CRC for drift-aware conformal risk control in foundation models under prompt and domain shift.
problem Fixed calibration risk in foundation models due to prompt and domain shift.
method Embeds prompts and responses, measures drift, gives more weight to recent examples, and updates risk online.
result Develops method to control risk up to terms for distribution mismatch and weighted quantile uncertainty.
Study helps start-ups choose best activities to succeed through milestones.
problem Navigating multiple milestones in multi-activity start-ups.
method Stochastic control model with multiple controls, optimal policy analysis.
result Optimal strategies depend on riskiness and cost-effectiveness.
Investing in declining tech boosts profits, study finds.
problem Optimal decision-making in declining profit streams.
method Modeling profit stream as Brownian motion with negative drift, analyzing thresholds for investment and exit.
result Investment threshold decreases in volatility when profit boost is large.
BCPO optimizes offline RL policies by converting uncertainty into conservative bounds.
problem Offline RL's fragility under distribution shifts and model errors.
method Bayesian approach with credible lower bounds and KL regularization.
result BCPO yields an uncertainty-calibrated policy that avoids exploiting model errors.
An algorithm for efficient experimentation in a dynamic environment with personalized preferences and context drifts.
problem Efficiently recommending decisions to users with personalized preferences in a context where the environment is changing over time.
method Dri-MED, inspired from the linear version of the MED strategy, adapted to handle non-stationary heteroskedastic noise.
result The instance-dependent regret scales as $ ilde{\mathcal O}\left(\fracκ{ ildeΔ}d^2(\log(T)
ight)$ , with i l d e Δ ildeΔ i l d e Δ being the constraint-aware sub-optimality gap. A new correction term improves sample efficiency in deep reinforcement learning.
problem Momentum accumulation in TD learning leads to doubly stale gradients.
method Proposed a correction term to address the issue of doubly stale gradients.
result Improves sample efficiency in policy evaluation.
SFPO optimizes LLM reasoning by repositioning before updating, improving stability and efficiency.
problem Noisy gradients from low-quality rollouts cause instability and inefficient exploration in on-policy RL algorithms.
method Decomposes each step into three stages: a short fast trajectory, repositioning, and slow correction, preserving the objective and rollout process unchanged.
result SFPO consistently improves stability, reduces rollouts, and accelerates convergence, outperforming GRPO on math reasoning benchmarks.
Wealth redistribution through Fokker-Planck equation controls preserves Gini coefficient.
problem Preserving Gini coefficient through proportional wealth tax.
method Formulating optimal redistribution as a control problem for Fokker-Planck equation.
result Progressive taxes redistribute within policy-relevant timescales.
Study optimal stock order placement in a diffusive market.
problem Optimal placement of a small order in a diffusive limit order book.
method Characterization of optimal limit order placement policy, analysis of behavior under different market conditions, and a simple method to approximate critical time and optimal order placement.
result Existence of a critical time t0 such that for t > t0, optimal placement differs from the best bid and second best bid.
This review covers learning under concept drift, including detection, understanding, and adaptation.
problem Unforeseeable changes in data distribution over time impact machine learning performance.
method Reviews and analyzes methodologies and techniques for concept drift detection, understanding, and adaptation.
result Establishes a framework for learning under concept drift with three main components.
In this paper we consider an energy storage optimization problem in finite time in a model with partial information that allows for a changing economic environment. The state process consists of the storage level controlled by the storage manager and the energy price process, which is a diffusion process the drift of w…
Lyapunov-based analysis shows polynomial sample complexity for WCMDPs and RBs.
problem Learning in WCMDPs and RBs under a generative model.
method Lyapunov-based analysis framework.
result Near-optimal policies can be learned with polynomial complexity.
New framework for detecting data drift in continuous time.
problem Drift in data distribution over time.
method Probability theoretical framework for continuous time drift.
result New efficient drift detection method and decomposition of data.
Identifies features most relevant to concept drift in data.
problem Identifying features most relevant to concept drift.
method Distinguishing between drift inducing and faithfully drifting features; deriving minimal subsets of features to characterize drift.
result Derives a detection algorithm for concept drift.
New framework detects adversarial concept drift in streaming data.
problem Adversarial concept drift in dynamic environments.
method Predict-Detect streaming framework for unsupervised drift detection and recovery.
result Framework detects adversarial drift with <6% labeled data, improving active learning for imbalanced data.
New method detects when models influence their own drift in real-time data streams.
problem Models can induce concept drift in real-time data streams.
method CheckerBoard Performative Drift Detection (CB-PDD)
result CB-PDD effectively detects performative drift in real-time data streams.
This research identifies flaws in drift detection methods and creates adversarial data streams to exploit them.
problem The challenge of detecting data distribution changes (drift) in real-time systems.
method Developed adversarial data streams to show weaknesses in existing drift detection schemes.
result Demonstrated that common drift detection methods can be fooled by adversarial data streams.
Paper proposes a semi-supervised method for detecting concept drift in streaming environments.
problem Detecting concept drift in streaming environments with limited labeled data.
method Utilizes density estimation of posterior probabilities in partially labeled streaming data.
result Demonstrates superior concept drift detection in streaming environments with limited labeled data.
We quantify forgetting in post-training models, distinguishing mass and drift.
problem Understanding and preventing forgetting in post-training generative models.
method Developed theoretical results under a two-mode mixture abstraction, formalizing mass and drift forgetting.
result Forgetting can be precisely quantified based on divergence direction, geometric overlap, and training regime.
A new drift detection method based on autoregressive models.
problem Concept drift in real-world data leads to decreased model performance.
method Autoregressive based drift detection method (ADDM).
result ADDM outperforms state-of-the-art drift detection methods.
New methods for anytime-valid off-policy inference in contextual bandits.
problem Estimating properties of hypothetical policies in adaptive experiments.
method Modern martingale techniques for comprehensive OPE inference.
result Valid anytime inference for off-policy mean reward values and entire reward distributions.
Adaptive sampling detects local concept drift with limited labels.
problem Detecting local concept drift in dynamic environments with scarce labels.
method Combines residual-based exploration and exploitation with EWMA monitoring.
result Superior performance in label efficiency and drift detection accuracy.
Algorithm detects concept drift and adapts models in streaming data.
problem Concept drift in streaming data renders models inaccurate.
method Adaptive learning algorithm that detects drifts and reacts to them.
result Risk competitive to an algorithm with perfect drift knowledge.
This work optimizes RL algorithms using entropy regularisation for continuous-time LQ problems.
problem Designing RL algorithms to balance exploration and exploitation in noisy environments.
method Entropy regularisation in two formulations: exploratory control and proximal policy update.
result Regret of O ( N ) \mathcal{O}(\sqrt{N}) O ( N ) for both learning algorithms over N N N episodes. The paper studies how expert opinions improve stock return predictions in a market with a hidden drift.
problem Improving stock return predictions in a market with a hidden Gaussian drift.
method Uses Kalman filter techniques to estimate the hidden drift from noisy expert opinions and stock returns.
result The Kalman filter estimates of the drift converge to the hidden drift as the frequency of expert opinions increases.
Classifies polynomial growth solutions to drift-harmonic equations on asymptotically paraboloidal manifolds.
problem Classifying polynomial growth solutions to drift-harmonic equations on specific types of manifolds.
method Inductive argument that alternates between constructing and asymptotically controlling drift-harmonic functions.
result All drift-harmonic functions with polynomial growth asymptotically separate variables and dimensions of spaces are computed.
New method detects drift in high-dimensional data.
problem Understanding and localizing concept drift in learning systems.
method Conformal predictions for drift localization.
result Our approach outperforms existing methods on image datasets.
Detects drifts in data for classification tasks using constrained embeddings.
problem Drifts in data affect model performance; unsupervised methods ignore label information.
method Task-sensitive semi-supervised drift detection with constrained low-dimensional embedding.
result Successfully detects real drifts affecting classification performance.
RL agent learns to manage inventory and price in dealer market simulations.
problem Managing inventory and price in a dealer market with RL.
method Multi-agent simulation, reinforcement learning, different reward formulations.
result RL agent learns competitor's pricing and manages inventory effectively.
PDD detects concept drift using explainable AI, improving model performance in dynamic environments.
problem Detecting and adapting to concept drift in predictive models.
method Profile Drift Detection (PDD) using Partial Dependence Profiles (PDPs).
result PDD outperforms existing methods in detecting concept drift and maintaining high predictive performance.
This paper studies concept drift detectors for financial time series.
problem Improving accuracy on financial time series with concept drifts.
method Three simple concept drift detectors tailored to financial time series.
result Two of the detectors are as effective as state-of-the-art detectors.
Paper proposes a framework to detect adversarial concept drifts under poisoning attacks.
problem Adversarial concept drift in data streams.
method Augmented Restricted Boltzmann Machine with improved gradient computation and energy function.
result High robustness and efficacy of the proposed drift detection framework in adversarial scenarios.
Kernel-Gradient Drifting improves generative modeling for non-Euclidean data.
problem Challenges in generative modeling for non-Euclidean data.
method Replaces Euclidean displacement with kernel-induced directions, exposing score-based structure.
result Kernel-gradient drifting enables state-of-the-art one-step generation for non-Euclidean data.