New Q-Learning algorithm reduces switching cost in MDPs.
problem Reducing adaptivity in real-world applications.
method Q-Learning with UCB2 exploration, quantified by local switching cost.
result Achieves sublinear regret with low switching cost.
New algorithm reduces switching costs in multinomial logit bandit problems.
problem Minimizing switching costs in multinomial logit bandit problems.
method Proposed AT-DUCB and FH-DUCB algorithms with low assortment switching costs.
result AT-DUCB and FH-DUCB algorithms achieve almost optimal minimax regret with low switching costs.
Adaptive Bayesian Optimization for resource-constrained experiments with switching costs.
problem Sequential experimental design with varying costs for changing design variables.
method Adapted batch algorithms to sequential problem, proposing cost-aware and cost-ignorant methods.
result Cost-aware algorithm outperforms tuned process-constrained algorithms in all settings considered.
We study the power of different types of adaptive (nonoblivious) adversaries in the setting of prediction with expert advice, under both full-information and bandit feedback. We measure the player's performance using a new notion of regret, also known as policy regret, which better captures the adversary's adaptiveness…
Optimal switching regret for all segmentations in online convex optimisation.
problem Non-stationary online convex optimisation problems.
method Developed an efficient algorithm to achieve optimal switching regret on every possible segmentation.
result Achieved asymptotically optimal switching regret on every possible segmentation simultaneously.
Solves risk-aware optimal switching problems in discrete time.
problem Non-Markovian optimal switching problems with risk awareness and general filtration.
method Solves reflected backward stochastic difference equations.
result Existence and uniqueness of solutions for the problems.
New method forecasts time series with changing variances.
problem Real-world processes with changing variances cannot be captured by classical models.
method State-space model with Markov switching variances, using online learning and expert aggregation.
result Proposed method outperforms traditional expert aggregation and is robust to misspecification.
New algorithm reduces switching costs in RL beyond linear MDPs.
problem Costly policy switching in reinforcement learning.
method ELEANOR-LowSwitching algorithm for linear Bellman-complete MDPs.
result Achieves near-optimal regret with logarithmic switching cost.
New RL algorithm reduces policy switching cost to loglog(T) with similar regret.
problem Low policy switching cost in real-life RL applications.
method Stage-wise exploration and adaptive policy elimination.
result Regret of O ( H S A log log T ) O(HSA \log\log T) O ( H S A log log T ) with O ( H S A log log T ) O(HSA \log\log T) O ( H S A log log T ) switching cost. Adaptive framework improves NB accuracy by fusing two index categories.
problem Challenges in attribute weighted NB, especially fusion of two indexes.
method Proposes ATFNB framework using switching factor to fuse two index categories.
result ATFNB outperforms basic NB and state-of-the-art models.
A new policy switching technique improves offline RL performance.
problem Challenges in adapting off-policy algorithms to different datasets and tasks.
method Combines off-policy RL and BC, using epistemic uncertainty for policy switching.
result Outperforms individual algorithms and state-of-the-art methods on benchmarks.
SAHMM-VAE separates sources adaptively using hidden Markov priors.
problem Unsupervised blind source separation.
method Source-wise adaptive Hidden Markov prior variational autoencoder.
result Different latent dimensions align with different source-specific temporal organizations.
Adaptive Heston model calibration using PCRLB and switching filters.
problem Estimating volatility in stochastic volatility models like Heston.
method Bayesian filtering (EKF, UKF, PF) with PCRLB for parameter estimation.
result Adaptive estimation of Heston model parameters improves volatility estimation.
Deep switch networks generate discrete data and language.
problem Generating high-dimensional discrete data and natural language.
method Adaptive switches model conditional distributions of discrete random variables. Maximum-likelihood objective function training with stochastic gradient descent.
result Stable and interpretable training of deep networks without backpropagation.
Efficient RL algorithms for linear function approximation with limited adaptivity constraints.
problem Limited adaptivity in reinforcement learning with linear function approximation.
method Proposed two efficient online RL algorithms for episodic linear Markov decision processes under batch learning and rare policy switch models.
result Achieved efficient regret bounds for both batch learning and rare policy switch models, with substantial reduction in adaptivity.
Researchers adaptively analyze market regimes to reveal investor behavior shifts.
problem Market relationships shift across different regimes, affecting investor behavior.
method Combining Kalman filtering, Markov-switching, and asymmetric response estimation.
result Foreign investors' predictive power increases during crises, while individual investors react more strongly to positive shocks.
New method tracks significant arm switches to improve bandit algorithms.
problem Adaptive procedures for bandits with unknown changes in reward distribution.
method Proposes a new notion of significant shift to count severe changes.
result Achieves faster rates than previous methods, especially when few changes are severe.
PCGS-TF uses a Transformer to adaptively control expert switching in non-stationary environments.
problem Static regret is insufficient for strictly online prediction in non-stationary settings.
method Policy-Controlled Generalized Share (PCGS) with a Transformer as an update controller.
result PCGS-TF achieves the lowest dynamic regret in non-stationary families and expert pools.
ADVISOR dynamically balances imitation and reinforcement learning to overcome the imitation gap.
problem The gap between imitation learning and reinforcement learning when teaching agents have privileged information.
method Adaptive Insubordination (ADVISOR) dynamically weights imitation and reward-based reinforcement learning losses.
result On-the-fly switching with ADVISOR outperforms pure imitation, pure reinforcement learning, and their combinations.
Markovian RNN adapts to nonstationary data using HMM for better time series prediction.
problem Nonstationary sequential data in real-life applications.
method Markovian RNN with HMM for regime switching and end-to-end optimization.
result Significant performance gains over vanilla RNN and Markov Switching ARIMA.
Proposes AWS method for precise speech enhancement using DNN.
problem T-F resolution problem in fixed-resolution short-time frequency transforms.
method Incorporates trainable adaptive window switching into speech enhancement procedure.
result Achieved higher signal-to-distortion ratio than conventional methods.
Algorithm for bandits with switching costs achieves optimal regret bounds.
problem Optimal regret bounds for stochastic and adversarial bandits with switching costs.
method Adaptation of Tsallis-INF algorithm with no prior knowledge of regime or time horizon.
result Achieves minimax optimal regret bounds in various settings.
Adaptive beamforming collapses in highly non-stationary environments, but the Universal Switching Beamformer resolves this by dynamically adjusting memory length.
problem Adaptive beamforming performance degrades in highly non-stationary environments.
method Integrating sequential prediction into the beamforming architecture.
result The USB achieves agility and precision in tracking highly non-stationary scenes.
New algorithm minimizes cumulative loss in dynamic linear bandits without prior knowledge of comparator switches.
problem Minimizing cumulative loss in dynamic linear bandits with unknown number of switches.
method Combining several bandit algorithms to adapt to unknown number of switches without prior knowledge.
result First algorithm achieving optimal regret guarantee of O ( d ( 1 + S T ) T ) \mathcal{O}\big(\sqrt{d(1+S_T) T}\big) O ( d ( 1 + S T ) T ) up to poly-logarithmic terms. Efficient algorithms for online convex optimization with limited switching decisions.
problem Online convex optimization with limited switching decisions.
method Presented computationally efficient algorithms for both general and strongly convex losses.
result Regret bounds of O ( T / S ) O(T/S) O ( T / S ) for general convex losses and O ~ ( T / S 2 ) \widetilde O(T/S^2) O ( T / S 2 ) for strongly convex losses. The study uses State Switching Markov Autoregressive models to identify and predict market regimes.
problem Adapting to abrupt changes in financial markets and identifying stable investment strategies.
method State Switching Markov Autoregressive models and Wyckoff Price Regimes.
result A dynamically adaptive trading system that outperforms traditional alphas.
HireVAE adapts to market regimes for online stock prediction.
problem Building an online and adaptive factor model for stock prediction.
method HireVAE uses a hierarchical latent space to estimate latent factors from historical market information.
result HireVAE outperforms previous methods in active returns across benchmarks.
Framework models multiscale dynamics with Bayesian learning for regime changes.
problem Analyzing complex interactions between fast and slow processes.
method Hierarchical state-space modeling with Sequential Monte Carlo.
result Bayesian approach accurately tracks state transitions and identifies switching dynamics.
Adaptive activity monitoring framework for wearable sensors.
problem Efficiently monitor human activities with low power consumption.
method Switching Gaussian process model with block circulant embedding and FFT for inference.
result Optimized trade-off between sensor power consumption and prediction performance.
A switchable deep beamformer enables versatile image processing.
problem Training and storing separate beamformers for each application.
method Switchable deep beamformer using Adaptive Instance Normalization (AdaIN) layers.
result Single network can produce various image processing outputs.
Two algorithms achieve optimal regret with limited adaptivity in multinomial logistic bandits.
problem Achieving optimal regret with limited adaptivity in multinomial logistic bandits.
method Presented two algorithms, B-MNL-CB and RS-MNL, for batched and rarely-switching paradigms.
result Achieved i l d e O ( T ) ilde{O}(\sqrt{T}) i l d e O ( T ) regret with limited adaptivity. Study on adaptivity constraints in linear contextual bandits with optimal design.
problem Impact of adaptivity constraints on linear contextual bandits.
method Two models of limited adaptivity: batch learning and rare policy switches. Proposed distributional optimal design.
result Achieves minimax-optimal regret with optimal number of policy switches and batches.
Adaptive TD learning reduces bias in policy evaluation by switching between TD and MC methods.
problem Achieving accurate policy evaluation with Temporal Difference (TD) learning in the presence of state-specific uncertainty.
method Adaptive switching between TD and Monte Carlo (MC) methods, using learned confidence intervals to detect and mitigate bias.
result The proposed adaptive algorithm outperforms existing methods in policy evaluation tasks.
New method identifies nonstationary causal structures in time series data.
problem Identifying causal relationships in time series data that change over time.
method High-order Markov Switching Models for regime-dependent causal discovery.
result Scalable approach for estimating high-order regime-dependent causal structures.
Enhances image quality to improve test-time adaptation accuracy.
problem Reducing accuracy loss due to distribution shift in deep networks.
method Integrates image enhancement with TTA methods to reduce prediction uncertainty.
result TECA method increases accuracy of TTA methods without hyperparameters.
Robots' agility in changing terrain helps financial models adapt to market shifts.
problem Challenges in financial market forecasting due to regime switching.
method Adapts pretrained LLMs using intrinsic market rewards and reinforcement learning.
result Significantly improved accuracy in adapting to market regime shifts.
We link disjoint longitudinal data for rare disease patients using latent representations and mixed-effects regression.
problem Analyzing treatment switches in rare diseases with limited data and changing measurement instruments.
method We embed item values into a shared latent space using variational autoencoders and apply mixed-effects regression to quantify treatment effects.
result Our approach allows for statistical inference and quantifies the impact of treatment switches in spinal muscular atrophy.
In the present paper, we studied a Dynamic Stochastic Block Model (DSBM) under the assumptions that the connection probabilities, as functions of time, are smooth and that at most s s s nodes can switch their class memberships between two consecutive time points. We estimate the edge probability tensor by a kernel-type p…
New method solves complex financial option pricing with varying time steps.
problem Pricing American options with varying time steps and regime switching.
method Explicit Runge-Kutta-Fehlberg scheme with fourth-order compact finite difference in space and high order analytical approximation.
result The method provides better performance in terms of computational speed and accuracy.
MPM uses machine learning to switch between two portfolio strategies for better risk management.
problem Adaptive portfolio strategy selection for improved risk management.
method XGBoost learns to switch between HRP and NRP strategies.
result MPM outperforms both HRP and NRP in risk-reward profile and interpretability.
Efficiently optimize GPs by reusing candidate solutions multiple times.
problem High computational cost of Gaussian process optimization due to unique historical points.
method Sticking to a candidate solution for multiple evaluation steps and limiting switches.
result Improved efficiency and practicality of Gaussian process optimization algorithms.
Efficient algorithms for contextual slate bandits with limited adaptivity.
problem Contextual slate bandit problem with limited adaptivity.
method Proposed B-SlateGLinCB and RS-SlateGLinCB algorithms for batched and rarely-switching settings.
result Achieved regret bounds of O(Nd^(3/2)√T) and O(Nd√T) under diversity assumption.
Investors in Target Date Funds are automatically switched from high risk to low risk assets as their retirements approach. Such funds have become very popular, but our analysis brings into question the rationale for them. Based on both a model with parameters fitted to historical returns and on bootstrap resampling, we…
PEAR dynamically reconfigures agent roles to prevent persistent biases in multi-agent debates.
problem Persistent positional biases and sensitivity to role assignments in fixed topologies.
method Dynamic reconfiguration of agent roles and sparse topologies based on evolving agent states.
result Significantly improves average accuracy over debate baselines across multiple reasoning benchmarks.
New DR-IC estimator reduces bias and variance in OPE.
problem Estimating value of a target policy using logged data from a different policy.
method DR-IC estimator that combines parametric reward model and context-based switching rule.
result DR-IC estimator outperforms state-of-the-art OPE algorithms.
New polynomial invariants derived from birack and switch structures.
problem Polynomial invariants of braids.
method Switch structures, birack colorings, quiver-valued invariants.
result New polynomial invariants of braids.
In this paper, we study optimal switching problems under ambiguity. To characterize the optimal switching under ambiguity in the finite horizon, we use multidimensional reflected backward stochastic differential equations (multidimensional RBSDEs) and show that a value function of the optimal switching under ambiguity …
Survey reviews code-switched speech and language processing.
problem Processing code-switched text and speech for multilingual communities.
method Reviews computational approaches and lists available resources.
result Essential for building intelligent agents that interact in multilingual settings.