GANs can be used to extract Fisher vectors for unsupervised feature learning.
problem Unsupervised feature extraction for classification and similarity tasks.
method Derive Fisher Information and Fisher Vectors from GANs' density model.
result GAN-induced Fisher Vectors perform competitively in unsupervised feature extraction.
Learning an encoding of feature vectors in terms of an over-complete dictionary or a information geometric (Fisher vectors) construct is wide-spread in statistical signal processing and computer vision. In content based information retrieval using deep-learning classifiers, such encodings are learnt on the flattened la…
TopoFisher learns topological summaries by maximizing Fisher information, improving parameter efficiency and inference quality.
problem Simulation-based inference misses key information in low-order statistics, especially for non-Gaussian fields.
method TopoFisher uses a differentiable persistent-homology pipeline that learns topological summaries by maximizing local Gaussian Fisher information.
result TopoFisher recovers much of the available information and outperforms fixed topological vectorizations in weak gravitational lensing.
Paper defines Fisher co-metric on cotangent bundle and clarifies its relation to variance.
problem Defines Fisher co-metric on cotangent bundle and clarifies its relation to variance.
method Defines Fisher co-metric directly from Fisher metric without going through tangent bundle, using a natural correspondence between cotangent vectors and random variables.
result Clarifies the relation between Fisher co-metric and variance/covariance, trivializing the Cramér-Rao inequality.
MULTIFIT tests independence between two random vectors using multiscale Fisher's test.
problem Detecting local dependence between two random vectors.
method MULTIFIT uses a resampling-free approach to test independence.
result MULTIFIT can easily handle large sample sizes and interpret dependency nature.
New discriminant analysis using GDS projection improves face recognition.
problem Improving face recognition accuracy with limited data.
method GDS projection onto generalized difference subspace, simplified Fisher criterion, normalization.
result GDS projection and gFDA are equivalent, inheriting FDA's discriminant ability.
Market strategies minimize Fisher information to minimize risk.
problem Applying minimum Fisher information principle to market dynamics.
method Analytical extension to quantum harmonic oscillator eigenstates and Gibbs distribution.
result Minimizing Fisher information reduces information and risk.
Concrete distribution properties examined on simplex.
problem Properties of Concrete distribution on simplex.
method Reflection and location-scale transformation of uniform distribution; explicit parameterization to Poincaré half-space.
result Fisher information and information metric are hyperbolic space; Fisher-Rao geodesic distance computed.
The family N of n-variate normal distributions is parameterized by the cone of positive definite symmetric n×n-matrices and the n-dimensional real vector space. Equipped with the Fisher information metric, N becomes a Riemannian manifold. As such, it is diffeomorphic, but not isometr…
Improved location estimation for high-dimensional data with finite sample size.
problem Estimating the shift in high-dimensional data with limited samples.
method Smoothed estimators and bounds on subgamma vectors.
result Convergence to Cramér-Rao bound for finite sample sizes.
Batch Normalization (BN) is essential to effectively train state-of-the-art deep Convolutional Neural Networks (CNN). It normalizes inputs to the layers during training using the statistics of each mini-batch. In this work, we study BN from the viewpoint of Fisher kernels. We show that assuming samples within a mini-ba…
In many structured prediction problems, complex relationships between variables are compactly defined using graphical structures. The most prevalent graphical prediction methods---probabilistic graphical models and large margin methods---have their own distinct strengths but also possess significant drawbacks. Conditio…
Paper improves matrix-valued data classification using nonparametric LDA.
problem Classification of matrix-valued data in neuroimaging and signal processing.
method Nonparametric LDA based on NPMLE for vectorized and scaled matrices.
result Improves classification performance across various data structures.
A new method improves few-shot learning by combining ProtoNet with LFD.
problem Few-shot learning struggles with high variance support sets.
method Combines ProtoNet with Local Fisher Discriminant Analysis.
result Superior classification accuracy on miniImageNet and tieredImageNet.
A framework monitors and diagnoses concept drift in supervised learning models.
problem Changes in predictive relationships over time render models suboptimal.
method Score vector monitoring using exponentially weighted moving average.
result Score-based approach detects concept drift more effectively than error-based methods.
Study on Gaussian-width complexity on statistical manifolds and its applications in learning and recovery.
problem Understanding the geometry of statistical manifolds and its implications for learning and recovery.
method Analysis of Fisher width and inverse-Fisher width, proving their complementary roles and establishing a relation between them.
result Established a sharp relation between Fisher width and inverse-Fisher width, showing they cannot reduce relative to Euclidean scale.
Fisher auto-encoders use Fisher divergence for more robust generative modeling.
problem Model uncertainty in generative models.
method Minimizing Fisher divergence between true and modeled joint distributions.
result Fisher auto-encoders can more accurately quantify model uncertainty.
A new geometric concept, the dead direction, bridges singular learning theory and information geometry.
problem The gap between singular learning theory and information geometry.
method Introducing the dead direction, a unit vector along degenerating Fisher metric, and showing its KL order can be recovered.
result The KL order of the dead direction can be recovered as the decay rate of the directional Fisher curvature, providing a handle on singular geometry.
Natural gradient descent, which preconditions a gradient descent update with the Fisher information matrix of the underlying statistical model, is a way to capture partial second-order information. Several highly visible works have advocated an approximation known as the empirical Fisher, drawing connections between ap…
Adaptive classifier optimizes high-dimensional data with spiked covariance structure.
problem Classification of high-dimensional data with spiked covariance structure.
method Adaptive classifier that whitens data, screens features, and applies Fisher linear discriminant.
result The classifier is Bayes optimal under certain conditions and performs well on real and synthetic data.
A new histogram layer improves texture analysis performance.
problem Extracting features for texture analysis from local spatial regions.
method Directly computes local spatial distribution of features during backpropagation.
result Improves performance on three material/texture datasets.
In this communication, we describe some interrelations between generalized q-entropies and a generalized version of Fisher information. In information theory, the de Bruijn identity links the Fisher information and the derivative of the entropy. We show that this identity can be extended to generalized versions of en…
Generative Adversarial Networks (GANs) are powerful models for learning complex distributions. Stable training of GANs has been addressed in many recent works which explore different metrics between distributions. In this paper we introduce Fisher GAN which fits within the Integral Probability Metrics (IPM) framework f…
It is well known that in a supervised classification setting when the number of features is smaller than the number of observations, Fisher's linear discriminant rule is asymptotically Bayes. However, there are numerous modern applications where classification is needed in the high-dimensional setting. Naive implementa…
The study examines Fisher-Riemann geodesics for nonparametric probability densities.
problem Understanding nonparametric probability densities using Fisher-Riemann geometry.
method Obtaining Fisher-Riemann geodesics as a limit of parametric cases with increasing parameters.
result The weak limit approach for nonparametric probability densities.
We propose a modified χβ-divergence, give some of its properties, and show that this leads to the definition of a generalized Fisher information. We give generalized Cramér-Rao inequalities, involving this Fisher information, an extension of the Fisher information matrix, and arbitrary norms and power of the estimat…
Blog post discusses various implementations of Fisher Information for EWC in continual learning.
problem Improving Elastic Weight Consolidation (EWC) results by optimizing Fisher Information computation.
method Empirically compares different implementations of Fisher Information for EWC.
result Many reported EWC results can be improved by changing Fisher Information computation methods.
Fisher width is a geometric measure of complexity on statistical manifolds.
problem Complexity measures on statistical manifolds
method Introducing Fisher width as a Fisher-geometric analogue of Gaussian width
result Fisher width retains key structural features of Gaussian width while capturing anisotropic geometric effects
The paper studies metrics on Lie groups SO(2) and SO(3) and their integrability.
problem Integrability of gradient systems on Lie groups via Fisher metrics.
method Analysis of Souriau-Fisher metrics and 2-cocycles on Lie groups SO(2) and SO(3).
result Cocycles can locally modify Fisher metrics on Lie group orbits.
Paper optimizes classification of distributions using Wasserstein metric.
problem Classifying instances represented by distributions on a vector space.
method Maximizing Fisher's ratio in the Wasserstein metric space through iterative algorithm.
result The method enhances classification performance and is robust to variations in distribution summaries.
This paper shows any Kähler metric can be a Fisher information metric.
problem Establishing a new characterization of Kähler and coKähler manifolds.
method Statistical approach using Fisher information and exponential families.
result Any Kähler metric is a Fisher information metric.
A new way to describe correlation matrices makes modeling easier.
problem Describing correlation matrices in a flexible and positive-definite way.
method Introduces a novel parametrization that allows unrestricted vectors for correlation matrices.
result The new parametrization ensures positive definiteness without additional constraints.
Proofs Fisher-Rao distance on Gaussian covariance manifold.
problem Proving Fisher-Rao distance on Gaussian covariance manifold.
method Basic Riemannian geometry.
result Proof of Fisher-Rao distance on covariance cone.
Survey on closed-form Fisher-Rao distance expressions.
problem Finding closed-form expressions for Fisher-Rao distance.
method Collect and present examples of closed-form expressions for Fisher-Rao distance of discrete and continuous distributions.
result Presentation of closed-form expressions for Fisher-Rao distance of various distributions.
This paper introduces a neural sampler for scalable sampling from complex distributions.
problem Efficiently sampling from high-dimensional un-normalized distributions.
method Neural implicit sampler trained with KL and Fisher divergence methods.
result The neural sampler generates large batches of samples with low computational costs.
New statistics are introduced that maintain the Fisher metric structure closely, akin to sufficient statistics.
problem Maintaining the Fisher metric structure in statistical models.
method Characterizing statistics that maintain the Fisher metric structure bi-Lipschitz equivalently.
result Characterized statistics that preserve the Fisher metric structure closely.
We consider three different approaches to define natural Riemannian metrics on polytopes of stochastic matrices. First, we define a natural class of stochastic maps between these polytopes and give a metric characterization of Chentsov type in terms of invariance with respect to these maps. Second, we consider the Fish…
Fisher score is one of the most widely used supervised feature selection methods. However, it selects each feature independently according to their scores under the Fisher criterion, which leads to a suboptimal subset of features. In this paper, we present a generalized Fisher score to jointly select features. It aims …
We introduce Fisher consistency in the sense of unbiasedness as a desirable property for estimators of class prior probabilities. Lack of Fisher consistency could be used as a criterion to dismiss estimators that are unlikely to deliver precise estimates in test datasets under prior probability and more general dataset…
In information theory, Fisher information and Shannon information (entropy) are respectively used to quantify the uncertainty associated with the distribution modeling and the uncertainty in specifying the outcome of given variables. These two quantities are complementary and are jointly applied to information behavior…
Study on geometry of Dirichlet distributions using Fisher-Rao metric.
problem Understanding the geometry of Dirichlet distributions.
method Analysis of Fisher-Rao metric on Dirichlet distribution parameter space.
result Geodesic completeness and negative sectional curvature of the space.
A deep neural network is a hierarchical nonlinear model transforming input signals to output signals. Its input-output relation is considered to be stochastic, being described for a given input by a parameterized conditional probability distribution of outputs. The space of parameters consisting of weights and biases i…
Paper formulates particle flow using variational inference and Fisher-Rao gradient flow.
problem Estimating posterior densities in probabilistic models.
method Variational formulation of particle flow, Fisher-Rao gradient flow, Gaussian and Gaussian mixture approximations.
result Gaussian and Gaussian mixture approximations of Fisher-Rao particle flow reduce to Exact Daum and Huang particle flow under linear Gaussian assumptions.
Paper discusses the Fisher metric and differentiability in statistical models.
problem Understanding the relationship between Fisher metric and differentiability in statistical models.
method Comparison of different concepts and models in Information Geometry, mathematical statistics, and measure theory.
result Discussion of various models and their differentiability properties.
Splat Regression Models use mixtures of bump functions to approximate complex data.
problem Approximating complex data with high interpretability and accuracy.
method Model outputs are mixtures of heterogeneous and anisotropic bump functions (splats) weighted by output vectors. Fitting splat models reduces to optimization over mixing measures using Wasserstein-Fisher-Rao gradient flows.
result Unified theoretical framework for Gaussian Splatting and flexible approach for diverse problems.
Paper identifies key function spaces for ReLU networks based on Fisher information.
problem Understanding the structure of Fisher information matrices in ReLU networks.
method Spectral decomposition of Fisher information matrices, focusing on the first three eigenspaces.
result The first three eigenspaces account for 97.7% of the trace of the Fisher information matrix, corresponding to spherical harmonic functions of order ≤2.
A new measure of model complexity based on Fisher Information.
problem Model complexity measurement in statistical models.
method Effective dimension defined by the number of cubes needed to cover the model space.
result The effective dimension is scale-dependent and measures model complexity.
The Fisher-Rao geometry is applied to elliptical distributions for optimization and classification.
problem Optimizing and classifying covariance matrices using geometric tools.
method Riemannian optimization and intrinsic Cramér-Rao bounds.
result Geometric tools enhance covariance matrix estimation and classification.