Proposes modifications to model-based forests for HTE estimation in observational data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We study asymptotic properties of some (essentially conditional least squares) parameter estimators for the subcritical Heston model based on discrete time observations derived from conditional least squares estimators of some modified parameters.
A method to reduce bias in model-based policy evaluation by shifting operators.
We prove strong consistency and asymptotic normality of least squares estimators for the subcritical Heston model based on continuous time observations. We also present some numerical illustrations of our results.
Model-based neural networks generalize better than ReLU networks for sparse recovery.
Model-free reinforcement learning methods such as the Proximal Policy Optimization algorithm (PPO) have successfully applied in complex decision-making problems such as Atari games. However, these methods suffer from high variances and high sample complexity. On the other hand, model-based reinforcement learning method…
This work learns visual representations for deformable objects using contrastive estimation.
The goal of reinforcement learning (RL) is to let an agent learn an optimal control policy in an unknown environment so that future expected rewards are maximized. The model-free RL approach directly learns the policy based on data samples. Although using many samples tends to improve the accuracy of policy learning, c…
This work improves uncertainty estimates for LISTA estimators.
This study tackles adversarial corruption in model-based reinforcement learning.
We study asymptotic properties of maximum likelihood estimators for Heston models based on continuous time observations of the log-price process. We distinguish three cases: subcritical (also called ergodic), critical and supercritical. In the subcritical case, asymptotic normality is proved for all the parameters, whi…
Traditional model-based reinforcement learning approaches learn a model of the environment dynamics without explicitly considering how it will be used by the agent. In the presence of misspecified model classes, this can lead to poor estimates, as some relevant available information is ignored. In this paper, we introd…
A new method, Count-MORL, improves offline reinforcement learning by using state-action frequency.
New method designs experiments robustly for nonlinear estimation, improving parameter knowledge.
Simple model-based reinforcement learning outperforms model-free methods in complex tasks.
Unified framework for model-based RL with sample complexity guarantees.
New bounds assess policy evaluation under unobserved confounders, showing model-based methods are more effective.
We propose a penalized likelihood method to jointly estimate multiple precision matrices for use in quadratic discriminant analysis and model based clustering. A ridge penalty and a ridge fusion penalty are used to introduce shrinkage and promote similarity between precision matrix estimates. Block-wise coordinate desc…
Model-based reinforcement learning algorithms tend to achieve higher sample efficiency than model-free methods. However, due to the inevitable errors of learned models, model-based methods struggle to achieve the same asymptotic performance as model-free methods. In this paper, We propose a Policy Optimization method w…
RL agent outperforms model-based approach in detecting price manipulation.
Model based predictions of future trajectories of a dynamical system often suffer from inaccuracies, forcing model based control algorithms to re-plan often, thus being computationally expensive, suboptimal and not reliable. In this work, we propose a model agnostic method for estimating the uncertainty of a model?s pr…
An evolutionary algorithm (EA) is developed as an alternative to the EM algorithm for parameter estimation in model-based clustering. This EA facilitates a different search of the fitness landscape, i.e., the likelihood surface, utilizing both crossover and mutation. Furthermore, this EA represents an efficient approac…
Estimates of predictive uncertainty are important for accurate model-based planning and reinforcement learning. However, predictive uncertainties---especially ones derived from modern deep learning systems---can be inaccurate and impose a bottleneck on performance. This paper explores which uncertainties are needed for…
Optimistic RL algorithms are simplified for deep RL with competitive performance.
We examine the impact of learning Lipschitz continuous models in the context of model-based reinforcement learning. We provide a novel bound on multi-step prediction error of Lipschitz models where we quantify the error using the Wasserstein metric. We go on to prove an error bound for the value-function estimate arisi…
New estimator improves off-policy evaluation for large action spaces.
We introduce "Search with Amortized Value Estimates" (SAVE), an approach for combining model-free Q-learning with model-based Monte-Carlo Tree Search (MCTS). In SAVE, a learned prior over state-action values is used to guide MCTS, which estimates an improved set of state-action values. The new Q-estimates are then used…
Designing effective model-based reinforcement learning algorithms is difficult because the ease of data generation must be weighed against the bias of model-generated data. In this paper, we study the role of model usage in policy optimization both theoretically and empirically. We first formulate and analyze a model-b…
New model learning objective improves continuous control tasks.
Improved model-based estimation through tempered Bayes filter.
New insights into bias-variance tradeoff for data-driven optimization under local misspecification.
MoMA improves model-based RL by using unrestricted policy classes.
Forest-based methods estimate heterogeneous treatment effects, blending strengths for better performance.
MAGE optimizes policies using action gradients from model-based learning.
A new algorithm STE for model-based RL improves learning rates.
Study evaluates new models using human feedback from another model.
Model-based learning algorithms have been shown to use experience efficiently when learning to solve Markov Decision Processes (MDPs) with finite state and action spaces. However, their high computational cost due to repeatedly solving an internal model inhibits their use in large-scale problems. We propose a method ba…
In this paper we study how to learn stochastic, multimodal transition dynamics in reinforcement learning (RL) tasks. We focus on evaluating transition function estimation, while we defer planning over this model to future work. Stochasticity is a fundamental property of many task environments. However, discriminative f…
This paper tackles distribution shift in model-based offline RL, proposing a shifts-aware reward method.
New method reduces variance in subpopulation model performance estimates.
Recent model-free reinforcement learning algorithms have proposed incorporating learned dynamics models as a source of additional data with the intention of reducing sample complexity. Such methods hold the promise of incorporating imagined data coupled with a notion of model uncertainty to accelerate the learning of c…
The effectiveness of model-based versus model-free methods is a long-standing question in reinforcement learning (RL). Motivated by recent empirical success of RL on continuous control tasks, we study the sample complexity of popular model-based and model-free algorithms on the Linear Quadratic Regulator (LQR). We show…
Model-based reinforcement learning is an appealing framework for creating agents that learn, plan, and act in sequential environments. Model-based algorithms typically involve learning a transition model that takes a state and an action and outputs the next state---a one-step model. This model can be composed with itse…
SMRL uses score matching for efficient RL with exponential family models.
In this text, we establish the risk model based on AR(1) series and propose the basic model which has a dependent structure under intensity of claim number. Considering some properties of the risk model, we take advantage of newton iteration method to figure out the adjustment coefficient and estimate the exponential u…
Paper proposes a new framework to improve policy optimization by aligning real and simulated data distributions.
Non-negative matrix factorization (NMF) is a technique for finding latent representations of data. The method has been applied to corpora to construct topic models. However, NMF has likelihood assumptions which are often violated by real document corpora. We present a double parametric bootstrap test for evaluating the…
This paper improves Bayesian inference for predictive models with limited data.