State entropy regularization improves robustness in reinforcement learning, especially under structured perturbations.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Entropy regularization is used to get improved optimization performance in reinforcement learning tasks. A common form of regularization is to maximize policy entropy to avoid premature convergence and lead to more stochastic policies for exploration through action space. However, this does not ensure exploration in th…
Paper proposes a policy-search algorithm to learn entropy-maximizing exploration policies in reward-free environments.
The study bounds entanglement entropy for coherent states on Kähler manifolds.
The paper introduces a new intrinsic reward method for exploration in reinforcement learning.
A neural network method estimates entropy production from system trajectories.
New RL approach uses future state and action visitation measures for better exploration.
Study on quantum state entanglement using Kaehler manifolds.
Proposes integrating global and local entropy for more reliable LLMs.
Study Transformer layers under cross-entropy training using mean field control.
PEOC uses policy entropy to detect untrained states in RL.
Many recent models of trade dynamics use the simple idea of wealth exchanges among economic agents in order to obtain a stable or equilibrium distribution of wealth among the agents. In particular, a plain analogy compares the wealth in a society with the energy in a physical system, and the trade between agents to the…
Estimate relaxation times in nonextensive systems using gradient flow for Tsallis entropy maximization.
We adapt tools from information theory to analyze how an observer comes to synchronize with the hidden states of a finitary, stationary stochastic process. We show that synchronization is determined by both the process's internal organization and by an observer's model of it. We analyze these components using the conve…
Paper proves Jeffrey's update rule minimizes relative entropy.
This paper improves reinforcement learning policies in a scalable way.
The policy gradient theorem is defined based on an objective with respect to the initial distribution over states. In the discounted case, this results in policies that are optimal for one distribution over initial states, but may not be uniformly optimal for others, no matter where the agent starts from. Furthermore, …
This study proposes hidden state curiosity to enhance RL models' resilience against noise.
We consider the problem of learning from demonstrated trajectories with inverse reinforcement learning (IRL). Motivated by a limitation of the classical maximum entropy model in capturing the structure of the network of states, we propose an IRL model based on a generalized version of the causal entropy maximization pr…
Tree-AMP simplifies inference in complex tree-structured models.
Enhances RL by controlling policy stochasticity through trajectory entropy constraints.
In 1870s, L. Boltzmann proved the famous -theorem for the Boltzmann equation in the kinetic theory of gas and gave the statistical interpretation of the thermodynamic entropy. In 2002, G. Perelman introduced the notion of -entropy and proved the -entropy formula for the Ricci flow. This plays a crucial role in…
In his 2011 work, Maas has shown that the law of any time-reversible continuous-time Markov chain with finite state space evolves like a gradient flow of the relative entropy with respect to its stationary distribution. In this work we show the converse to the above by showing that if the relative law of a Markov chain…
We study a new ensemble of random correlation matrices related to multivariate Student (or more generally elliptic) random variables. We establish the exact density of states of empirical correlation matrices that generalizes the Marcenko-Pastur result. The comparison between the theoretical density of states in the St…
We propose a general framework for entropy-regularized average-reward reinforcement learning in Markov decision processes (MDPs). Our approach is based on extending the linear-programming formulation of policy optimization in MDPs to accommodate convex regularization functions. Our key result is showing that using the …
This thesis applies entropy as a model independent measure to address three research questions concerning financial time series. In the first study we apply transfer entropy to drawdowns and drawups in foreign exchange rates, to study their correlation and cross correlation. When applied to daily and hourly EUR/USD and…
JES optimizes expensive functions by considering joint entropy over input and output spaces.
Two hitherto disconnected threads of research, diverse exploration (DE) and maximum entropy RL have addressed a wide range of problems facing reinforcement learning algorithms via ostensibly distinct mechanisms. In this work, we identify a connection between these two approaches. First, a discriminator-based diversity …
In recent years, deep reinforcement learning has been shown to be adept at solving sequential decision processes with high-dimensional state spaces such as in the Atari games. Many reinforcement learning problems, however, involve high-dimensional discrete action spaces as well as high-dimensional state spaces. This pa…
Entropy-regularized NPG converges linearly with linear function approximation.
Optimal control in latent factor models uses Tsallis entropy for exploration.
Study shows depth improves generalization in deep learning models.
Many policy gradient methods are variants of Actor-Critic (AC), where a value function (critic) is learned to facilitate updating the parameterized policy (actor). The update to the actor involves a log-likelihood update weighted by the action-values, with the addition of entropy regularization for soft variants. In th…
The aim of this paper is to state and prove polynomial analogues of the classical Manning inequality relating the topological entropy of a geodesic flow with the growth rate of the volume of balls in the universal covering. To this aim we use two numerical conjugacy invariants, the {\em strong polynomial entropy $h_{po…
Entropy minimization has been widely used in unsupervised domain adaptation (UDA). However, existing works reveal that entropy minimization only may result into collapsed trivial solutions. In this paper, we propose to avoid trivial solutions by further introducing diversity maximization. In order to achieve the possib…
Differentiable PF via entropy-regularized OT for better inference.
AR-DAE approximates entropy gradient for machine learning models.
Paper introduces a new measure combining entropy and Gini index.
Coupled entropy corrects flaws in Tsallis entropy for complex systems.
The financial market entropy is modeled using open quantum systems.
Accounting for the non-normality of asset returns remains challenging in robust portfolio optimization. In this article, we tackle this problem by assessing the risk of the portfolio through the "amount of randomness" conveyed by its returns. We achieve this by using an objective function that relies on the exponential…
We study the problem of identifying the causal relationship between two discrete random variables from observational data. We recently proposed a novel framework called entropic causality that works in a very general functional model but makes the assumption that the unobserved exogenous variable has small entropy in t…
The article presents a new entropy model for assessing stock market interest.
The well known maximum-entropy principle due to Jaynes, which states that given mean parameters, the maximum entropy distribution matching them is in an exponential family, has been very popular in machine learning due to its "Occam's razor" interpretation. Unfortunately, calculating the potentials in the maximum-entro…
New algorithms reduce complexity for learning in MDPs with entropy regularization.
Trust-region methods have yielded state-of-the-art results in policy search. A common approach is to use KL-divergence to bound the region of trust resulting in a natural gradient policy update. We show that the natural gradient and trust region optimization are equivalent if we use the natural parameterization of a st…
We present a new method of generating mixture models for data with categorical attributes. The keys to this approach are an entropy-based density metric in categorical space and annealing of high-entropy/low-density components from an initial state with many components. Pruning of low-density components using the entro…
In this paper, we present a new class of Markov decision processes (MDPs), called Tsallis MDPs, with Tsallis entropy maximization, which generalizes existing maximum entropy reinforcement learning (RL). A Tsallis MDP provides a unified framework for the original RL problem and RL with various types of entropy, includin…