New method uses entropy to improve policy gradient exploration.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New method reduces state distribution mismatch in off-policy RL.
This work improves policy optimization by maximizing entropy of state distribution, leading to better exploration.
In this letter we borrow from the inference techniques developed for unbounded state-cardinality (nonparametric) variants of the HMM and use them to develop a tuning-parameter free, black-box inference procedure for Explicit-state-duration hidden Markov models (EDHMM). EDHMMs are HMMs that have latent states consisting…
The method approximates stationary distributions of Markov models by truncating irrelevant states.
New method estimates state-action stationary distribution for better off-policy policy evaluation.
CFIL uses coupled flows to model state distributions for imitation learning.
Study optimal offline RL with uncertainty sets and distribution shifts.
State-constrained offline RL expands RL's learning scope.
This paper tackles belief-state selection in simulators with latent states.
Quantum models learn unitary actions on entangled states from product states.
In this article we present an alternative model for the distribution of household incomes in the United States. We provide arguments from two differing perspectives which both yield the proposed income distribution curve, and then fit this curve to empirical data on household income distribution obtained from the Unite…
We present the data on wealth and income distributions in the United Kingdom, as well as on the income distributions in the individual states of the USA. In all of these data, we find that the great majority of population is described by an exponential distribution, whereas the high-end tail follows a power law. The di…
The paper analyzes variational autoencoders for state space models with risk bounds.
Proposes a new model for time series that considers smooth transitions between states.
New algorithm for ML models in gradually adapting data settings.
The problem of state estimation for unobservable distribution systems is considered. A deep learning approach to Bayesian state estimation is proposed for real-time applications. The proposed technique consists of distribution learning of stochastic power injection, a Monte Carlo technique for the training of a deep ne…
A new method learns state and proposal dynamics in state-space models using neural networks.
Learn true model from metastable samples of discrete distributions.
We study the problem of off-policy policy optimization in Markov decision processes, and develop a novel off-policy policy gradient method. Prior off-policy policy gradient approaches have generally ignored the mismatch between the distribution of states visited under the behavior policy used to collect data, and what …
The distribution of money is analysed in connection with the Boltzmann distribution of energy in the degenerate states of molecules. Plots of the population density of income distribution for various countries are well reproduced by a Gamma function, confirming the validity of the statistical distribution at equilibriu…
DeepGG generates graph distributions for drug discovery and molecular design.
Improves out-of-distribution detection in neural networks.
MPE framework proves universal approximation for quantum data distribution.
Due to the insufficient measurements in the distribution system state estimation (DSSE), full observability and redundant measurements are difficult to achieve without using the pseudo measurements. The matrix completion state estimation (MCSE) combines the matrix completion and power system model to estimate voltage b…
IHT improves sparse distribution learning.
State-space models are successfully used in many areas of science, engineering and economics to model time series and dynamical systems. We present a fully Bayesian approach to inference \emph{and learning} (i.e. state estimation and system identification) in nonlinear nonparametric state-space models. We place a Gauss…
New approach models how explanations shift with distribution changes.
In this work, we build on recent advances in distributional reinforcement learning to give a generally applicable, flexible, and state-of-the-art distributional variant of DQN. We achieve this by using quantile regression to approximate the full quantile function for the state-action return distribution. By reparameter…
Improved state estimation in nonlinear models using amortized backward variational inference.
Paper develops a new state estimation method for nonlinear systems.
RegFlow models future states with flexible probability distributions.
Causal Imitation Learning handles noisy measurements and distribution shifts.
Paper presents a fast method for estimating hidden states in Bayesian models.
A three-state model based on the Potts model is proposed to simulate financial markets. The three states are assigned to "buy", "sell" and "inactive" states. The model shows the main stylized facts observed in the financial market: fat-tailed distributions of returns and long time correlations in the absolute returns. …
A nonparametric approach for policy learning for POMDPs is proposed. The approach represents distributions over the states, observations, and actions as embeddings in feature spaces, which are reproducing kernel Hilbert spaces. Distributions over states given the observations are obtained by applying the kernel Bayes' …
IDAC improves reinforcement learning efficiency by modeling implicit distributions.
Studying general quantum many-body systems is one of the major challenges in modern physics because it requires an amount of computational resources that scales exponentially with the size of the system.Simulating the evolution of a state, or even storing its description, rapidly becomes intractable for exact classical…
A critical and challenging problem in reinforcement learning is how to learn the state-action value function from the experience replay buffer and simultaneously keep sample efficiency and faster convergence to a high quality solution. In prior works, transitions are uniformly sampled at random from the replay buffer o…
Distributed strategic learning has been getting attention in recent years. As systems become distributed finding Nash equilibria in a distributed fashion is becoming more important for various applications. In this paper, we develop a distributed strategic learning framework for seeking Nash equilibria under stochastic…
State aggregation is a popular model reduction method rooted in optimal control. It reduces the complexity of engineering systems by mapping the system's states into a small number of meta-states. The choice of aggregation map often depends on the data analysts' knowledge and is largely ad hoc. In this paper, we propos…
New RL approach uses future state and action visitation measures for better exploration.
We consider a network scenario in which agents can evaluate each other according to a score graph that models some interactions. The goal is to design a distributed protocol, run by the agents, that allows them to learn their unknown state among a finite set of possible values. We propose a Bayesian framework in which …
Efficiently estimates online variational learning using importance sampling.
Exploration is critical to a reinforcement learning agent's performance in its given environment. Prior exploration methods are often based on using heuristic auxiliary predictions to guide policy behavior, lacking a mathematically-grounded objective with clear properties. In contrast, we recast exploration as a proble…
The paper introduces a new intrinsic reward method for exploration in reinforcement learning.
Schrödinger bridge solved with Weyl calculus for quadratic state cost.
Machine learning promises methods that generalize well from finite labeled data. However, the brittleness of existing neural net approaches is revealed by notable failures, such as the existence of adversarial examples that are misclassified despite being nearly identical to a training example, or the inability of recu…