Accurately determining dependency structure is critical to discovering a system's causal organization. We recently showed that the transfer entropy fails in a key aspect of this---measuring information flow---due to its conflation of dyadic and polyadic relationships. We extend this observation to demonstrate that this…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Information-theoretic Bayesian regret bounds of Russo and Van Roy capture the dependence of regret on prior uncertainty. However, this dependence is through entropy, which can become arbitrarily large as the number of actions increases. We establish new bounds that depend instead on a notion of rate-distortion. Among o…
Information theory provides ideas for conceptualising information and measuring relationships between objects. It has found wide application in the sciences, but economics and finance have made surprisingly little use of it. We show that time series data can usefully be studied as information -- by noting the relations…
Improved price bounds for multi-asset derivatives using market option data.
The paper proposes methods to extract and analyze individual variable information from complex dependencies.
We discuss the connection between information and copula theories by showing that a copula can be employed to decompose the information content of a multivariate distribution into marginal and dependence components, with the latter quantified by the mutual information. We define the information excess as a measure of d…
We investigated financial market data to determine which factors affect information flow between stocks. Two factors, the time dependency and the degree of efficiency, were considered in the analysis of Korean, the Japanese, the Taiwanese, the Canadian, and US market data. We found that the frequency of the significant…
We simplify information measure computation using learned features.
MIC consistently estimates dependence in large datasets.
New bounds show limitations of sample-wise information-theoretic generalization.
DOS improves language model generation by considering inter-token dependencies.
In this work, we develop a novel regularizer to improve the learning of long-range dependency of sequence data. Applied on language modelling, our regularizer expresses the inductive bias that sequence variables should have high mutual information even though the model might not see abundant observations for complex lo…
Optimizes black-box functions with varying costs across multiple sources.
Recent works investigated the generalization properties in deep neural networks (DNNs) by studying the Information Bottleneck in DNNs. However, the mea- surement of the mutual information (MI) is often inaccurate due to the density estimation. To address this issue, we propose to measure the dependency instead of MI be…
In this work, we improve upon the stepwise analysis of noisy iterative learning algorithms initiated by Pensia, Jog, and Loh (2018) and recently extended by Bu, Zou, and Veeravalli (2019). Our main contributions are significantly improved mutual information bounds for Stochastic Gradient Langevin Dynamics via data-depe…
Investigates optimal portfolio strategies in markets with latent side information.
Mutual information minimum spanning trees are used to explore nonlinear dependencies on Brazilian equity network in the periods from June/01/2015 to January/26/2016, in which Brazil was under the government of President Dilma Rousseff, and from January/27/2016 to September/08/2016 which includes the government transiti…
Understanding how different information sources together transmit information is crucial in many domains. For example, understanding the neural code requires characterizing how different neurons contribute unique, redundant, or synergistic pieces of information about sensory or behavioral variables. Williams and Beer (…
In this paper we extend the theory of option pricing to take into account and explain the empirical evidence for asset prices such as non-Gaussian returns, long-range dependence, volatility clustering, non-Gaussian copula dependence, as well as theoretical issues such as asymmetric information and the presence of limit…
Study nearest-neighbor radii under dependent sampling, finding they remain informative.
We study 'meta-dependence' in conditional independence tests across different empirical distributions.
InfoAtlas speeds up MI estimation for real-time data analysis.
This paper improves dependency networks using information geometry.
We extend manifold capacity to nonlinear neural representations with contextual information.
New research on limits of transfer learning, proving key selection and dependence requirements.
Bounding the generalization error of learning algorithms has a long history, which yet falls short in explaining various generalization successes including those of deep learning. Two important difficulties are (i) exploiting the dependencies between the hypotheses, (ii) exploiting the dependence between the algorithm'…
Paper analyzes trade-offs between fairness, privacy, and accuracy using Chernoff Information.
The paper argues that normalized mutual information is biased in clustering and community detection.
We consider a sequential learning problem with Gaussian payoffs and side information: after selecting an action , the learner receives information about the payoff of every action in the form of Gaussian observations whose mean is the same as the mean payoff, but the variance depends on the pair (and may…
A new deep learning framework captures multi-scale spatio-temporal dependencies.
Study improves sampling from non-log-concave distributions using Fisher information.
The paper analyzes cryptocurrency trading networks using pairwise and high-order dependencies.
Measures dependence between two systems using Bayesian model comparison.
The Mutual Information (MI) is an often used measure of dependency between two random variables utilized in information theory, statistics and machine learning. Recently several MI estimators have been proposed that can achieve parametric MSE convergence rate. However, most of the previously proposed estimators have th…
In many cases, feature selection is often more complicated than identifying a single subset of input variables that would together explain the output. There may be interactions that depend on contextual information, i.e., variables that reveal to be relevant only in some specific circumstances. In this setting, the con…
A new method extracts events and their arguments efficiently from text.
Dropout technique is analyzed using information geometry.
The paper analyzes convergence of Langevin dynamics with time-dependent metrics.
Paper tackles robust classification under class-dependent domain shift.
We introduce a new framework for unsupervised learning of representations based on a novel hierarchical decomposition of information. Intuitively, data is passed through a series of progressively fine-grained sieves. Each layer of the sieve recovers a single latent factor that is maximally informative about multivariat…
New algorithms improve privacy in bandit problems with partial information.
The Partial Information Decomposition (PID) [arXiv:1004.2515] provides a theoretical framework to characterize and quantify the structure of multivariate information sharing. A new method (Idep) has recently been proposed for computing a two-predictor PID over discrete spaces. [arXiv:1709.06653] A lattice of maximum en…
New methods estimate point-wise dependency from neural MI models.
This work improves independence tests for high-dimensional data.
The vast majority of optimization and online learning algorithms today require some prior information about the data (often in the form of bounds on gradients or on the optimal parameter value). When this information is not available, these algorithms require laborious manual tuning of various hyperparameters, motivati…
Information concentration of probability measures have important implications in learning theory. Recently, it is discovered that the information content of a log-concave distribution concentrates around their differential entropy, albeit with an unpleasant dependence on the ambient dimension. In this work, we prove th…
The identification of relevant features, i.e., the driving variables that determine a process or the properties of a system, is an essential part of the analysis of data sets with a large number of variables. A mathematical rigorous approach to quantifying the relevance of these features is mutual information. Mutual i…
The paper decomposes probabilistic scores into reliability, uncertainty, and information loss.