Paper proposes a policy-search algorithm to learn entropy-maximizing exploration policies in reward-free environments.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Accounting for the non-normality of asset returns remains challenging in robust portfolio optimization. In this article, we tackle this problem by assessing the risk of the portfolio through the "amount of randomness" conveyed by its returns. We achieve this by using an objective function that relies on the exponential…
This paper finds a unique partition of a sample space for estimating continuous distributions.
In this paper, we suggest a framework to make use of mutual information as a regularization criterion to train Auto-Encoders (AEs). In the proposed framework, AEs are regularized by minimization of the mutual information between input and encoding variables of AEs during the training phase. In order to estimate the ent…
Information theory provides principled ways to analyze different inference and learning problems such as hypothesis testing, clustering, dimensionality reduction, classification, among others. However, the use of information theoretic quantities as test statistics, that is, as quantities obtained from empirical data, p…
Most data is multi-dimensional. Discovering whether any subset of dimensions, or subspaces, of such data is significantly correlated is a core task in data mining. To do so, we require a measure that quantifies how correlated a subspace is. For practical use, such a measure should be universal in the sense that it capt…
If pricing kernels are assumed non-negative then the inverse problem of finding the pricing kernel is well-posed. The constrained least squares method provides a consistent estimate of the pricing kernel. When the data are limited, a new method is suggested: relaxed maximization of the relative entropy. This estimator …
Causal discovery is a fundamental problem in statistics and has wide applications in different fields. Transfer Entropy (TE) is a important notion defined for measuring causality, which is essentially conditional Mutual Information (MI). Copula Entropy (CE) is a theory on measurement of statistical independence and is …
Mixture distributions arise in many parametric and non-parametric settings -- for example, in Gaussian mixture models and in non-parametric estimation. It is often necessary to compute the entropy of a mixture, but, in most cases, this quantity has no closed-form expression, making some form of approximation necessary.…
In this paper, we propose a general framework to learn a robust large-margin binary classifier when corrupt measurements, called anomalies, caused by sensor failure might be present in the training set. The goal is to minimize the generalization error of the classifier on non-corrupted measurements while controlling th…
In this paper, we propose a general framework to learn a robust large-margin binary classifier when corrupt measurements, called anomalies, caused by sensor failure might be present in the training set. The goal is to minimize the generalization error of the classifier on non-corrupted measurements while controlling th…
New method speeds up lead-lag detection between asynchronous time series.
State entropy regularization improves robustness in reinforcement learning, especially under structured perturbations.
Entropy regularization is used to get improved optimization performance in reinforcement learning tasks. A common form of regularization is to maximize policy entropy to avoid premature convergence and lead to more stochastic policies for exploration through action space. However, this does not ensure exploration in th…
Locally private mechanisms' output divergence bounds derived.
Method recovers causal diffusion mechanisms from steady-state data without parametric assumptions.
The study bounds entanglement entropy for coherent states on Kähler manifolds.
The paper introduces a new intrinsic reward method for exploration in reinforcement learning.
Simpler GNNs with low-rank non-parametric aggregators perform well on graph benchmarks.
Given two views of data, we consider the problem of finding the features of one view which can be most faithfully inferred from the other. We find that these are also the most correlated variables in the sense of deep canonical correlation analysis (DCCA). Moreover, we show that these variables can be used to construct…
Information transfer between time series is calculated by using the asymmetric information-theoretic measure known as transfer entropy. Geweke's autoregressive formulation of Granger causality is used to find linear transfer entropy, and Schreiber's general, non-parametric, information-theoretic formulation is used to …
A neural network method estimates entropy production from system trajectories.
This paper solves the intractability barrier in non-parametric information geometry by introducing a novel framework.
New RL approach uses future state and action visitation measures for better exploration.
Study evaluates policies in partially observable environments without full model specification.
Study on quantum state entanglement using Kaehler manifolds.
A new statistical model uses Orlicz-Sobolev spaces with Gaussian weight.
MAXENT method outperforms ML in sparse data with specific prior correlations.
Proposes integrating global and local entropy for more reliable LLMs.
The paper improves prediction intervals for non-parametric regression using histograms.
Estimates dependent parameters using Markovian dependence with shrinkage.
The paper reviews methods for estimating individual treatment effects using non-parametric regression models.
Study Transformer layers under cross-entropy training using mean field control.
Understanding and measuring model risk is important to financial practitioners. However, there lacks a non-parametric approach to model risk quantification in a dynamic setting and with path-dependent losses. We propose a complete theory generalizing the relative-entropic approach by Glasserman and Xu to the dynamic ca…
PEOC uses policy entropy to detect untrained states in RL.
Many recent models of trade dynamics use the simple idea of wealth exchanges among economic agents in order to obtain a stable or equilibrium distribution of wealth among the agents. In particular, a plain analogy compares the wealth in a society with the energy in a physical system, and the trade between agents to the…
A graph-based method for two-sample testing across connected nodes.
Estimate relaxation times in nonextensive systems using gradient flow for Tsallis entropy maximization.
This work develops a non-parametric test for relational independence in non-i.i.d. data.
Forest Fire Clustering discovers cell types from single-cell data.
New GoF test improves change point detection in multivariate time series.
We adapt tools from information theory to analyze how an observer comes to synchronize with the hidden states of a finitary, stationary stochastic process. We show that synchronization is determined by both the process's internal organization and by an observer's model of it. We analyze these components using the conve…
Paper proves Jeffrey's update rule minimizes relative entropy.
KQT-EWMA monitors multivariate data streams online with flexible and practical change detection.
This paper improves reinforcement learning policies in a scalable way.
A new framework based on the theory of copulas is proposed to address semi- supervised domain adaptation problems. The presented method factorizes any multivariate density into a product of marginal distributions and bivariate cop- ula functions. Therefore, changes in each of these factors can be detected and corrected…
Ensembles of classification and regression trees remain popular machine learning methods because they define flexible non-parametric models that predict well and are computationally efficient both during training and testing. During induction of decision trees one aims to find predicates that are maximally informative …
The policy gradient theorem is defined based on an objective with respect to the initial distribution over states. In the discounted case, this results in policies that are optimal for one distribution over initial states, but may not be uniformly optimal for others, no matter where the agent starts from. Furthermore, …