Study on optimal information acquisition in Kyle model with entropy cost.
problem Optimal information acquisition in Kyle model with entropy cost.
method Continuous signals are optimal, and any signal with a logit posterior distribution yields the same ex-ante value.
result Posterior expected payoff becomes normally distributed as information acquisition cost increases.
Paper benchmarks mutual info estimators on diverse distributions.
problem Evaluating mutual information estimators on complex, real-world distributions.
method Constructs a diverse family of known-ground truth distributions, proposes a benchmark platform.
result Highlights differences in classical and neural estimators' performance across various conditions.
New estimator improves mutual information estimation.
problem Estimating mutual information in data science and machine learning.
method Proposes a new estimator that uses a preliminary estimate of the data distribution.
result A preliminary estimate helps in estimating mutual information more accurately.
Inference for normal and Monte Carlo distributions using minimum relative entropy.
problem Inference from partial information on expectations and covariances.
method Minimum relative entropy sub-manifolds, analytical formulas, Monte Carlo simulations.
result Improved numerical implementation for inference from partial information.
Concrete distribution properties examined on simplex.
problem Properties of Concrete distribution on simplex.
method Reflection and location-scale transformation of uniform distribution; explicit parameterization to Poincaré half-space.
result Fisher information and information metric are hyperbolic space; Fisher-Rao geodesic distance computed.
In-N-Out improves model robustness to out-of-distribution data.
problem Learning robust models with few in-distribution labeled examples.
method Pre-training with auxiliary information and self-training with pseudolabels.
result In-N-Out outperforms auxiliary inputs or outputs alone on both in-distribution and OOD error.
The paper analyzes extreme risk measures with limited distributional information.
problem Investigating risk measures under partial knowledge of distribution moments and shape.
method Employing probability inequalities and modified Schwarz inequality to derive bounds on distortion risk measures.
result Unified framework for calculating best- and worst-case scenarios of distortion risk measures.
The study explains stock return distributions using reaction functions.
problem Stock return distributions often deviate from normal distributions.
method Assumes normal event/information effects, financial over/underreaction, proposes reaction function model.
result Financial markets often underreact to minor events, overreact to significant ones, and react stronger to positive events.
Study connects covariance cleaning theory to information theory for heavy-tailed distributions.
problem Optimizing covariance matrices for heavy-tailed distributions using information theory.
method Minimizing Frobenius norm and information loss between true and estimated covariance matrices.
result Asymptotic regime of large matrices minimizes information loss for Student's t distributions.
Paper describes profiles of multivariate normal distributions and novel estimators for mutual information.
problem Estimating mutual information for complex distributions.
method Analytical description of profiles, introduction of Bend and Mix Models, Monte Carlo estimation.
result Bend and Mix Models accurately estimate mutual information profiles and provide Bayesian estimates.
Paper introduces a geometric approach to model similar probability distributions.
problem Incorporating similar probability distributions into graphical models.
method Information geometric approach to model similar distributions.
result Allows reinterpretation of existing models.
New method corrects active learning for distribution shifts and outliers.
problem Conventional active learning methods fail to account for test-time distribution.
method JEPIG, a hybrid of BALD and EPIG, maximizes expected predictive information gain.
result JEPIG outperforms conventional methods in active learning with distribution shifts.
In this communication, we describe some interrelations between generalized q-entropies and a generalized version of Fisher information. In information theory, the de Bruijn identity links the Fisher information and the derivative of the entropy. We show that this identity can be extended to generalized versions of en…
Study a market with uncertain informed traders, finding price impact depends on both asset value and informed trader count distribution.
problem Uncertain participation of informed traders in a market with limit orders.
method Characterized equilibrium by a fixed point integral equation, analyzed large order asymptotics, solved numerically.
result Equilibrium price impact depends on both asset value and distribution of informed traders, not just expected number of informed traders.
Word embeddings provide point representations of words containing useful semantic information. We introduce multimodal word distributions formed from Gaussian mixtures, for multiple word meanings, entailment, and rich uncertainty information. To learn these distributions, we propose an energy-based max-margin objective…
Study OOD generalization in meta-reinforcement learning using information theory.
problem Understanding how meta-reinforcement learning handles distribution shifts.
method Information-theoretic analysis of Markov Decision Processes and gradient-based algorithms.
result Established fine-grained generalization bounds for meta-reinforcement learning.
Augmented bridge matching preserves coupling information between distributions.
problem Preserving the original empirical pairing in flow and bridge matching processes.
method Augmenting the velocity field with initial sample point information.
result Simple modification recovers coupling information without losing Markovian property.
Efficiently estimates distributed mean with side information, near-optimal and universal.
problem Distributed mean estimation with side information in communication constrained settings.
method Wyner-Ziv estimators for communication and computation efficiency.
result Near-optimal and universal recovery guarantees for distributed optimization and compression.
Improves robustness of information bottleneck framework with sparsity-inducing prior.
problem Fixed-dimensional priors restrict flexibility and restrict robustness.
method Sparsity-inducing spike-slab categorical prior that learns dimension distribution per data point.
result Improves accuracy and robustness compared to traditional priors and other methods.
We introduce the Mutual Information Machine (MIM), a probabilistic auto-encoder for learning joint distributions over observations and latent variables. MIM reflects three design principles: 1) low divergence, to encourage the encoder and decoder to learn consistent factorizations of the same underlying distribution; 2…
The paper uses geometric methods to classify medical data histograms.
problem Classifying medical data histograms for disease diagnosis.
method Information geometry of beta distributions for comparing and classifying histograms.
result Geometric tools, particularly negatively curved Fisher information, enable unique mean calculation and K-means classification.
The conditional mutual information I(X;Y|Z) measures the average information that X and Y contain about each other given Z. This is an important primitive in many learning problems including conditional independence testing, graphical model inference, causal strength estimation and time-series problems. In several appl…
Paper analyzes trade-offs between fairness, privacy, and accuracy using Chernoff Information.
problem The relationship between fairness and privacy in machine learning.
method Utilizes Chernoff Information to characterize trade-offs, proposes Chernoff Difference and Noisy Chernoff Difference, develops CINE for neural estimation.
result Shows three distinct behaviors of Noisy Chernoff Difference based on data distribution.
We introduce the Mutual Information Machine (MIM), a novel formulation of representation learning, using a joint distribution over the observations and latent state in an encoder/decoder framework. Our key principles are symmetry and mutual information, where symmetry encourages the encoder and decoder to learn differe…
This letter introduces an abstract learning problem called the "set embedding": The objective is to map sets into probability distributions so as to lose less information. We relate set union and intersection operations with corresponding interpolations of probability distributions. We also demonstrate a preliminary so…
In this paper, we describe the "implicit autoencoder" (IAE), a generative autoencoder in which both the generative path and the recognition path are parametrized by implicit distributions. We use two generative adversarial networks to define the reconstruction and the regularization cost functions of the implicit autoe…
FIRE method improves model performance in federated learning by penalizing fragmentation-induced covariate shifts.
problem Performance degradation in federated learning due to data fragmentation and covariate shift.
method FIRE method accumulates fragmentation-induced covariate shift divergences via approximate Fisher information and uses it as a per-fragment loss penalty.
result FIRE outperforms importance weighting and federated learning benchmarks by up to 5.3% on shifted validation sets.
Improved mean estimation for symmetric distributions with finite-sample guarantees.
problem Estimating the mean of a symmetric distribution from samples.
method Using Fisher information rate for finite-sample guarantees.
result Finite-sample convergence close to subgaussian with variance 1/(n * I_r), where I_r is r-smoothed Fisher information.
As Computer Vision moves from a passive analysis of pixels to active analysis of semantics, the breadth of information algorithms need to reason over has expanded significantly. One of the key challenges in this vein is the ability to identify the information required to make a decision, and select an action that will …
New method uses KL-divergence to create non-informative priors for multivariate Gaussian.
problem Handling hyperparameters for non-informative limits in multivariate Gaussian conjugate priors.
method Using scaled KL-divergence between multivariate Gaussians to construct Wishart and normal-Wishart conjugate priors.
result Forming non-informative priors without violating Wishart shape parameter restrictions.
We propose an unsupervised object matching method for relational data, which finds matchings between objects in different relational datasets without correspondence information. For example, the proposed method matches documents in different languages in multi-lingual document-word networks without dictionaries nor ali…
Paper derives best- and worst-case GlueVaR measures with incomplete data.
problem Risk measurement with limited information and shape constraints.
method Unified framework based on partial distribution information and shape properties.
result Characterization of extremal GlueVaR distributions with convex envelopes.
New approach combines invariance and information bottleneck for OOD generalization.
problem OOD generalization failures in classification tasks.
method Revisit linear regression tasks, prove information bottleneck constraint necessary, propose combined approach.
result Combined invariance and information bottleneck approach improves OOD generalization.
We derive asset pricing formula for markets with incomplete information and subjective views.
problem Asset pricing in markets with informational imperfections and subjective investor beliefs.
method Closed-form market equilibrium formula based on Merton's model, non-linear system of equations, conditional posterior distribution.
result Derivation of market reference model for excess returns under random shadow-costs.
Study explores geometric structure and prior for beta-logistic distribution.
problem Understanding the geometric structure and prior distributions of the beta-logistic distribution.
method Exploring dual geometric structure and uncovering α-parallel prior. result The beta-logistic distribution admits an α-parallel prior for any real number α. We show that gamma distributions provide models for departures from randomness since every neighbourhood of an exponential distribution contains a neighbourhood of gamma distributions, using an information theoretic metric topology. We derive also the information geometry of the 3-manifold of McKay bivariate gamma dist…
We discuss the connection between information and copula theories by showing that a copula can be employed to decompose the information content of a multivariate distribution into marginal and dependence components, with the latter quantified by the mutual information. We define the information excess as a measure of d…
Bayesian method predicts asset returns for better portfolio optimization.
problem Uncertainty in financial markets makes traditional portfolio optimization methods unreliable.
method Bayesian predictive synthesis (BPS) combined with dynamic linear models.
result Predicted distribution information improves portfolio performance.
Bayesian Tensor Network combines prior and data likelihood for efficient prediction and parameter estimation.
problem Overfitting and poor performance in Tensor Network models.
method Introduce prior distribution, use Laplace approximation for posterior predictive distribution, and propose stable initialization for parameter estimation.
result Reduces overfitting and improves performance of Tensor Network models.
Paper explores Elliptical Wishart distributions in signal processing and machine learning.
problem Estimating parameters of Elliptical Wishart distributions.
method Proposes fixed point and Riemannian optimization algorithms for maximum likelihood estimation.
result Characterizes existence, uniqueness, and convergence of the MLE.
Precise estimation of uncertainty in predictions for AI systems is a critical factor in ensuring trust and safety. Deep neural networks trained with a conventional method are prone to over-confident predictions. In contrast to Bayesian neural networks that learn approximate distributions on weights to infer prediction …
Bayesian PROCOVA uses AI to adjust for covariates in RCTs.
problem Unbiased and precise treatment effect inferences from RCTs.
method Generative AI constructs digital twins for covariate adjustment, using an additive mixture prior.
result Efficiency gains in smaller RCTs compared to frequentist methods.
Feature selection is one of the most fundamental problems in machine learning. An extensive body of work on information-theoretic feature selection exists which is based on maximizing mutual information between subsets of features and class labels. Practical methods are forced to rely on approximations due to the diffi…
New divergence measures improve KL approximation.
problem Improving KL divergence approximation without AC condition.
method Introduced α-geodesical skew divergence. result Properties of α-geodesical skew divergence studied. Machine learning is used to compute achievable information rates (AIRs) for a simplified fiber channel. The approach jointly optimizes the input distribution (constellation shaping) and the auxiliary channel distribution to compute AIRs without explicit channel knowledge in an end-to-end fashion.
The paper calculates bounds for risk metrics and entropies under partial information constraints.
problem Analyzing risk metrics and entropies for unimodal, symmetric distributions with limited information.
method Develops lower and upper bounds for worst-case distortion riskmetrics and weighted entropy for unimodal, symmetric distributions with known mean and variance.
result Sharp upper bounds for distortion riskmetrics and weighted entropy for symmetric distributions.
We describe Information Forests, an approach to classification that generalizes Random Forests by replacing the splitting criterion of non-leaf nodes from a discriminative one -- based on the entropy of the label distribution -- to a generative one -- based on maximizing the information divergence between the class-con…
Paper formulates mutual information optimal control for discrete-time systems.
problem Optimal control of discrete-time linear systems with mutual information.
method Formulates MIOCP as an extension of MEOCP, derives optimal policy and prior, proposes alternating minimization algorithm.
result Proposes an alternating minimization algorithm for MIOCP.