Estimates RL data for dynamic treatment effects using GMM.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We present and prove properties of a new offline policy evaluator for an exploration learning setting which is superior to previous evaluators. In particular, it simultaneously and correctly incorporates techniques from importance weighting, doubly robust evaluation, and nonstationary policy evaluation approaches. In a…
A new method optimizes in nonstationary environments with many arms efficiently.
A contextual bandit method evaluates and improves inventory control policies.
Flexible nonstationary Gaussian process with neural network parameters.
New algorithm balances personalization and statistical validity in MRTs.
ConvNets improve nonstationary covariance estimation for large-scale spatial data.
SyMPLER improves time series forecasting in nonstationary environments with explainable models.
Adaptive estimation for nonstationary time series reduces computational cost.
Bootstrap method for Markov chains in reinforcement learning.
The condition for stationary increments, not scaling, detemines long time pair autocorrelations. An incorrect assumption of stationary increments generates spurious stylized facts, fat tails and a Hurst exponent H_s=1/2, when the increments are nonstationary, as they are in FX markets. The nonstationarity arises from s…
Rational bubbles form in nonstationary models of real assets.
Recent work in distance metric learning has focused on learning transformations of data that best align with provided sets of pairwise similarity and dissimilarity constraints. The learned transformations lead to improved retrieval, classification, and clustering algorithms due to the better adapted distance or similar…
DORIS algorithm achieves no-regret learning in Markov games with adversarial opponents.
Recent work in distance metric learning has focused on learning transformations of data that best align with specified pairwise similarity and dissimilarity constraints, often supplied by a human observer. The learned transformations lead to improved retrieval, classification, and clustering algorithms due to the bette…
DTS improves robustness of bandit algorithms in nonstationary environments.
Exponential inequalities are main tools in machine learning theory. To prove exponential inequalities for non i.i.d random variables allows to extend many learning techniques to these variables. Indeed, much work has been done both on inequalities and learning theory for time series, in the past 15 years. However, for …
Motivated by the many real-world applications of reinforcement learning (RL) that require safe-policy iterations, we consider the problem of off-policy evaluation (OPE) -- the problem of evaluating a new policy using the historical data obtained by different behavior policies -- under the model of nonstationary episodi…
A key challenge in leveraging data augmentation for neural network training is choosing an effective augmentation policy from a large search space of candidate operations. Properly chosen augmentation policies can lead to significant generalization improvements; however, state-of-the-art approaches such as AutoAugment …
The method of cointegration in regression analysis is based on an assumption of stationary increments. Stationary increments with fixed time lag are called integration I(d). A class of regression models where cointegration works was identified by Granger and yields the ergodic behavior required for equilibrium expectat…
We consider the problem of off-policy evaluation in Markov decision processes. Off-policy evaluation is the task of evaluating the expected return of one policy with data generated by a different, behavior policy. Importance sampling is a technique for off-policy evaluation that re-weights off-policy returns to account…
Motivated by recommendation problems in music streaming platforms, we propose a nonstationary stochastic bandit model in which the expected reward of an arm depends on the number of rounds that have passed since the arm was last pulled. After proving that finding an optimal policy is NP-hard even when all model paramet…
We consider a continuous-time model for inventory management with Markov modulated non-stationary demands. We introduce active learning by assuming that the state of the world is unobserved and must be inferred by the manager. We also assume that demands are observed only when they are completely met. We first derive t…
Unified formulation bridges adversarial and nonstationary bandits.
Bayesian optimization has become a fundamental global optimization algorithm in many problems where sample efficiency is of paramount importance. Recently, there has been proposed a large number of new applications in fields such as robotics, machine learning, experimental design, simulation, etc. In this paper, we foc…
Paper analyzes algorithms for nonstationary saddle-point optimization problems.
A Markov Decision Process (MDP) is a popular model for reinforcement learning. However, its commonly used assumption of stationary dynamics and rewards is too stringent and fails to hold in adversarial, nonstationary, or multi-agent problems. We study an episodic setting where the parameters of an MDP can differ across…
The paper explains why estimating a history-dependent policy can reduce MSE in reinforcement learning.
Adaptive estimation of alpha-Stable distribution and Hurst exponent for nonstationary time series.
A new approach learns to represent context for nonstationary bandits.
The analysis of nonstationary time series is of great importance in many scientific fields such as physics and neuroscience. In recent years, Gaussian process regression has attracted substantial attention as a robust and powerful method for analyzing time series. In this paper, we introduce a new framework for analyzi…
Model nonstationary spatial processes using normalizing flows.
New algorithm for nonstationary multi-armed bandits with optimal performance.
We introduce a new approach for comparing reinforcement learning policies, using Wasserstein distances (WDs) in a newly defined latent behavioral space. We show that by utilizing the dual formulation of the WD, we can learn score functions over policy behaviors that can in turn be used to lead policy optimization towar…
Brain-computer interfaces (BCIs) have enabled prosthetic device control by decoding motor movements from neural activities. Neural signals recorded from cortex exhibit nonstationary property due to abrupt noises and neuroplastic changes in brain activities during motor control. Current state-of-the-art neural signal de…
SORSCNs improve nonstationary data modeling by self-organizing and adjusting network parameters.
We propose a method to clean covariance matrices of nonstationary systems by using time-independent eigenvalues.
It is ubiquitous in natural and social sciences that two variables, recorded temporally or spatially in a complex system, are cross-correlated and possess multifractal features. We propose a new method called multifractal detrended cross-correlation analysis (MF-DXA) to investigate the multifractal behaviors in the pow…
Study minimax off-policy evaluation in multi-armed bandits with known and unknown behavior policies.
Develops methods to test nonstationarity and detect change points in RL.
New method identifies nonstationary causal structures in time series data.
New method for identifying causal relationships in financial time series data.
We propose a novel approach to train a multi-modal policy from mixed demonstrations without their behavior labels. We develop a method to discover the latent factors of variation in the demonstrations. Specifically, our method is based on the variational autoencoder with a categorical latent variable. The encoder infer…
Markovian RNN adapts to nonstationary data using HMM for better time series prediction.
AREBA algorithm improves learning from imbalanced, nonstationary data.
BerlinUCB learns from episodic rewards in nonstationary contexts.
A novel nonstationary permanental process relaxes kernel constraints and captures complex data patterns.
Most real world phenomena such as sunlight distribution under a forest canopy, minerals concentration, stock valuation, exhibit nonstationary dynamics i.e. phenomenon variation changes depending on the locality. Nonstationary dynamics pose both theoretical and practical challenges to statistical machine learning algori…