Fisher width is a geometric measure of complexity on statistical manifolds.
problem Complexity measures on statistical manifolds
method Introducing Fisher width as a Fisher-geometric analogue of Gaussian width
result Fisher width retains key structural features of Gaussian width while capturing anisotropic geometric effects
Deep networks can be biased to learn top eigenfunctions of the kernel outside the training set.
problem Spectral bias of deep networks in the kernel regime.
method Quantitative bounds on L2 difference between finite-width and infinite-width network trajectories. result Deep networks learn top eigenfunctions of the Neural Tangent Kernel over the entire input space, not just the training set.
FedFisher improves one-shot FL by using Fisher information.
problem Reducing communication rounds in federated learning.
method Bayesian perspective, Fisher information matrices, diagonal Fisher, K-FAC approximation.
result FedFisher achieves vanishingly small error in two-layer neural networks.
This study explains why approximate NGD works well in wide neural networks.
problem Understanding why NGD with approximate Fisher information converges fast in wide neural networks.
method Analyzing asymptotic training dynamics in function space via the neural tangent kernel.
result NGD with approximate Fisher information achieves the same fast convergence as exact NGD under specific conditions.
The Fisher information matrix (FIM) plays an essential role in statistics and machine learning as a Riemannian metric tensor or a component of the Hessian matrix of loss functions. Focusing on the FIM and its variants in deep neural networks (DNNs), we reveal their characteristic scale dependence on the network width, …
The Fisher information matrix (FIM) is a fundamental quantity to represent the characteristics of a stochastic model, including deep neural networks (DNNs). The present study reveals novel statistics of FIM that are universal among a wide class of DNNs. To this end, we use random weights and large width limits, which e…
Modified training direction reduces generalization error in neural networks.
problem Reducing generalization error in neural networks.
method Theoretical analysis of modified natural gradient descent in function space.
result Modifying training direction in function space reduces total generalization error.
The paper proves geometric and spectral alignment for deep neural networks.
problem Understanding the singular spectra of deep neural network layers.
method Proves deterministic quotient-geometric estimates for singular spectra of Frobenius-normalized layer factors.
result Exact power-law spectra form a trace-normalized Cartan orbit under Frobenius normalization.
Fisher auto-encoders use Fisher divergence for more robust generative modeling.
problem Model uncertainty in generative models.
method Minimizing Fisher divergence between true and modeled joint distributions.
result Fisher auto-encoders can more accurately quantify model uncertainty.
Natural gradient descent, which preconditions a gradient descent update with the Fisher information matrix of the underlying statistical model, is a way to capture partial second-order information. Several highly visible works have advocated an approximation known as the empirical Fisher, drawing connections between ap…
Paper defines Fisher co-metric on cotangent bundle and clarifies its relation to variance.
problem Defines Fisher co-metric on cotangent bundle and clarifies its relation to variance.
method Defines Fisher co-metric directly from Fisher metric without going through tangent bundle, using a natural correspondence between cotangent vectors and random variables.
result Clarifies the relation between Fisher co-metric and variance/covariance, trivializing the Cramér-Rao inequality.
In this communication, we describe some interrelations between generalized q-entropies and a generalized version of Fisher information. In information theory, the de Bruijn identity links the Fisher information and the derivative of the entropy. We show that this identity can be extended to generalized versions of en…
Market strategies minimize Fisher information to minimize risk.
problem Applying minimum Fisher information principle to market dynamics.
method Analytical extension to quantum harmonic oscillator eigenstates and Gibbs distribution.
result Minimizing Fisher information reduces information and risk.
Generative Adversarial Networks (GANs) are powerful models for learning complex distributions. Stable training of GANs has been addressed in many recent works which explore different metrics between distributions. In this paper we introduce Fisher GAN which fits within the Integral Probability Metrics (IPM) framework f…
The study examines Fisher-Riemann geodesics for nonparametric probability densities.
problem Understanding nonparametric probability densities using Fisher-Riemann geometry.
method Obtaining Fisher-Riemann geodesics as a limit of parametric cases with increasing parameters.
result The weak limit approach for nonparametric probability densities.
We propose a modified χβ-divergence, give some of its properties, and show that this leads to the definition of a generalized Fisher information. We give generalized Cramér-Rao inequalities, involving this Fisher information, an extension of the Fisher information matrix, and arbitrary norms and power of the estimat…
Blog post discusses various implementations of Fisher Information for EWC in continual learning.
problem Improving Elastic Weight Consolidation (EWC) results by optimizing Fisher Information computation.
method Empirically compares different implementations of Fisher Information for EWC.
result Many reported EWC results can be improved by changing Fisher Information computation methods.
The paper studies metrics on Lie groups SO(2) and SO(3) and their integrability.
problem Integrability of gradient systems on Lie groups via Fisher metrics.
method Analysis of Souriau-Fisher metrics and 2-cocycles on Lie groups SO(2) and SO(3).
result Cocycles can locally modify Fisher metrics on Lie group orbits.
This paper shows any Kähler metric can be a Fisher information metric.
problem Establishing a new characterization of Kähler and coKähler manifolds.
method Statistical approach using Fisher information and exponential families.
result Any Kähler metric is a Fisher information metric.
Proofs Fisher-Rao distance on Gaussian covariance manifold.
problem Proving Fisher-Rao distance on Gaussian covariance manifold.
method Basic Riemannian geometry.
result Proof of Fisher-Rao distance on covariance cone.
Survey on closed-form Fisher-Rao distance expressions.
problem Finding closed-form expressions for Fisher-Rao distance.
method Collect and present examples of closed-form expressions for Fisher-Rao distance of discrete and continuous distributions.
result Presentation of closed-form expressions for Fisher-Rao distance of various distributions.
New statistics are introduced that maintain the Fisher metric structure closely, akin to sufficient statistics.
problem Maintaining the Fisher metric structure in statistical models.
method Characterizing statistics that maintain the Fisher metric structure bi-Lipschitz equivalently.
result Characterized statistics that preserve the Fisher metric structure closely.
We consider three different approaches to define natural Riemannian metrics on polytopes of stochastic matrices. First, we define a natural class of stochastic maps between these polytopes and give a metric characterization of Chentsov type in terms of invariance with respect to these maps. Second, we consider the Fish…
Fisher score is one of the most widely used supervised feature selection methods. However, it selects each feature independently according to their scores under the Fisher criterion, which leads to a suboptimal subset of features. In this paper, we present a generalized Fisher score to jointly select features. It aims …
We introduce Fisher consistency in the sense of unbiasedness as a desirable property for estimators of class prior probabilities. Lack of Fisher consistency could be used as a criterion to dismiss estimators that are unlikely to deliver precise estimates in test datasets under prior probability and more general dataset…
In information theory, Fisher information and Shannon information (entropy) are respectively used to quantify the uncertainty associated with the distribution modeling and the uncertainty in specifying the outcome of given variables. These two quantities are complementary and are jointly applied to information behavior…
Study on geometry of Dirichlet distributions using Fisher-Rao metric.
problem Understanding the geometry of Dirichlet distributions.
method Analysis of Fisher-Rao metric on Dirichlet distribution parameter space.
result Geodesic completeness and negative sectional curvature of the space.
A deep neural network is a hierarchical nonlinear model transforming input signals to output signals. Its input-output relation is considered to be stochastic, being described for a given input by a parameterized conditional probability distribution of outputs. The space of parameters consisting of weights and biases i…
Paper formulates particle flow using variational inference and Fisher-Rao gradient flow.
problem Estimating posterior densities in probabilistic models.
method Variational formulation of particle flow, Fisher-Rao gradient flow, Gaussian and Gaussian mixture approximations.
result Gaussian and Gaussian mixture approximations of Fisher-Rao particle flow reduce to Exact Daum and Huang particle flow under linear Gaussian assumptions.
Paper discusses the Fisher metric and differentiability in statistical models.
problem Understanding the relationship between Fisher metric and differentiability in statistical models.
method Comparison of different concepts and models in Information Geometry, mathematical statistics, and measure theory.
result Discussion of various models and their differentiability properties.
TopoFisher learns topological summaries by maximizing Fisher information, improving parameter efficiency and inference quality.
problem Simulation-based inference misses key information in low-order statistics, especially for non-Gaussian fields.
method TopoFisher uses a differentiable persistent-homology pipeline that learns topological summaries by maximizing local Gaussian Fisher information.
result TopoFisher recovers much of the available information and outperforms fixed topological vectorizations in weak gravitational lensing.
Normalization methods play an important role in enhancing the performance of deep learning while their theoretical understandings have been limited. To theoretically elucidate the effectiveness of normalization, we quantify the geometry of the parameter space determined by the Fisher information matrix (FIM), which als…
Paper identifies key function spaces for ReLU networks based on Fisher information.
problem Understanding the structure of Fisher information matrices in ReLU networks.
method Spectral decomposition of Fisher information matrices, focusing on the first three eigenspaces.
result The first three eigenspaces account for 97.7% of the trace of the Fisher information matrix, corresponding to spherical harmonic functions of order ≤2.
The Fisher-Rao geometry is applied to elliptical distributions for optimization and classification.
problem Optimizing and classifying covariance matrices using geometric tools.
method Riemannian optimization and intrinsic Cramér-Rao bounds.
result Geometric tools enhance covariance matrix estimation and classification.
The study examines Fisher information matrices and neural tangent kernels for simple ReLU networks with random weights.
problem Understanding the relationship between Fisher information matrices and neural tangent kernels for 2-layer ReLU networks.
method Analyzes Fisher information matrices and neural tangent kernels for 2-layer ReLU networks with random hidden weights, focusing on spectral decomposition and eigenfunctions.
result Obtained an approximation formula for functions represented by 2-layer neural networks.
New tensor framework connects Fisher information, hypergraphs, and multi-observable correlations.
problem Missing structure in pairwise Fisher graphs for multi-observable radiation patterns.
method Higher-order Fisher tensors and natural exponential-family coordinates.
result Exact triality of Fisher tensors, cumulants, and hypergraphs.
New method for natural policy gradients converges linearly.
problem Improving natural policy gradient methods for better convergence.
method Fisher-Rao gradient flow applied to state-action distributions.
result Linear convergence rate with geometry-dependent factor.
Global gradient estimates for Fisher-KPP equation on Finsler metric measure spaces.
problem Establishing gradient estimates for the Finslerian Fisher-KPP equation.
method Global gradient estimates on compact and noncompact Finsler metric measure spaces using the traditional CD(K,N) condition and new comparison theorems. result Global gradient estimates for positive solutions of the Finslerian Fisher-KPP equation.
Paper explores Fisher-Rao gradient flows and their kernel approximations.
problem Understanding and analyzing approximations of Fisher-Rao gradient flows.
method Rigorous investigation of Fisher-Rao and Wasserstein type gradient flows, focusing on kernel approximations.
result Proves evolutionary Γ-convergence for kernel-approximated Fisher-Rao flows, providing theoretical guarantees.
Study improves sampling from non-log-concave distributions using Fisher information.
problem Sampling from non-log-concave distributions with high Fisher information guarantees.
method Proximal sampler with RGO implementation, leveraging log-concave sampling results.
result Improved complexity guarantee in relative Fisher information for non-log-concave sampling.
We present a novel synthesis of Fisher information and asset pricing theory that yields a practical method for reconstructing the probability density implicit in security prices. The Fisher information approach to these inverse problems transforms the search for a probability density into the solution of a differential…
Two Fisher information matrix estimators are analyzed for neural networks, focusing on their variances and trade-offs.
problem Estimating the Fisher information matrix in neural networks due to its high computational cost.
method Examined two popular diagonal Fisher information matrix estimators and their variances in neural networks for regression and classification.
result The variances of the estimators depend on the non-linearity with respect to different parameter groups and should not be neglected.
We examine Generative Adversarial Networks (GANs) through the lens of deep Energy Based Models (EBMs), with the goal of exploiting the density model that follows from this formulation. In contrast to a traditional view where the discriminator learns a constant function when reaching convergence, here we show that it ca…
Fisher loss improves deep domain adaptation by learning discriminative within-class compact and between-class separable representations.
problem Improving deep domain adaptation performance by learning discriminative representations.
method Proposes a Fisher loss to learn discriminative representations that are within-class compact and between-class separable.
result Noticeable improvements in deep domain adaptation performance, e.g., 6.67% absolute improvement in mean accuracy on the Office-Home dataset.
We define the Wirtinger width of a knot. Then we prove the Wirtinger width of a knot equals its Gabai width. The algorithmic nature of the Wirtinger width leads to an efficient technique for establishing upper bounds on Gabai width. As an application, we use this technique to calculate the Gabai width of approximately …
This is a detailed tutorial paper which explains the Fisher discriminant Analysis (FDA) and kernel FDA. We start with projection and reconstruction. Then, one- and multi-dimensional FDA subspaces are covered. Scatters in two- and then multi-classes are explained in FDA. Then, we discuss on the rank of the scatters and …
We propose a Bayesian framework of Gaussian process in order to extend Fisher's discriminant to classify functional data such as spectra and images. The probability structure for our extended Fisher's discriminant is explicitly formulated, and we utilize the smoothness assumptions of functional data as prior probabilit…
Cosine schedule is optimal for discrete diffusion models.
problem Choosing the best discretization schedule for diffusion models.
method Optimized using Fisher-Rao geometry.
result Cosine schedule is Fisher-Rao optimal.