Paper identifies key function spaces for ReLU networks based on Fisher information.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
High-dimensional models become unstable when sample size falls below a critical level, leading to a phase transition.
This paper is a tutorial for eigenvalue and generalized eigenvalue problems. We first introduce eigenvalue problem, eigen-decomposition (spectral decomposition), and generalized eigenvalue problem. Then, we mention the optimization problems which yield to the eigenvalue and generalized eigenvalue problems. We also prov…
Many deep learning models are vulnerable to the adversarial attack, i.e., imperceptible but intentionally-designed perturbations to the input can cause incorrect output of the networks. In this paper, using information geometry, we provide a reasonable explanation for the vulnerability of deep learning models. By consi…
The study examines Fisher information matrices and neural tangent kernels for simple ReLU networks with random weights.
The Fisher information matrix (FIM) is a fundamental quantity to represent the characteristics of a stochastic model, including deep neural networks (DNNs). The present study reveals novel statistics of FIM that are universal among a wide class of DNNs. To this end, we use random weights and large width limits, which e…
The paper studies eigenvalues of graph Laplacians on data clouds and proves central limit theorems.
The paper analyzes how quantization affects the Fisher Information Matrix's dominant eigenvalue.
We propose a method to learn causal response representations through direct effect analysis.
We propose a scheme for defending against adversarial attacks by suppressing the largest eigenvalue of the Fisher information matrix (FIM). Our starting point is one explanation on the rationale of adversarial examples. Based on the idea of the difference between a benign sample and its adversarial example is measured …
Noise can affect the overparametrization of QNNs, enabling new directions but also suppressing sensitivity.
New distances for comparing multivariate normal distributions.
The Fisher information matrix (FIM) plays an essential role in statistics and machine learning as a Riemannian metric tensor or a component of the Hessian matrix of loss functions. Focusing on the FIM and its variants in deep neural networks (DNNs), we reveal their characteristic scale dependence on the network width, …
We introduce RSE to measure robustness in estimation problems.
Study on GEPs with generative priors, showing optimal statistical rates and proposing an iterative algorithm.
EigenVI uses orthogonal function expansions for efficient variational inference.
In this paper, we consider the sparse eigenvalue problem wherein the goal is to obtain a sparse solution to the generalized eigenvalue problem. We achieve this by constraining the cardinality of the solution to the generalized eigenvalue problem and obtain sparse principal component analysis (PCA), sparse canonical cor…
Combining insights from machine learning and quantum Monte Carlo, the stochastic reconfiguration method with neural network Ansatz states is a promising new direction for high-precision ground state estimation of quantum many-body problems. Even though this method works well in practice, little is known about the learn…
Data structure affects deep learning performance, study finds.
Study on Gaussian-width complexity on statistical manifolds and its applications in learning and recovery.
Fisher auto-encoders use Fisher divergence for more robust generative modeling.
Natural gradient descent, which preconditions a gradient descent update with the Fisher information matrix of the underlying statistical model, is a way to capture partial second-order information. Several highly visible works have advocated an approximation known as the empirical Fisher, drawing connections between ap…
Paper defines Fisher co-metric on cotangent bundle and clarifies its relation to variance.
In this communication, we describe some interrelations between generalized -entropies and a generalized version of Fisher information. In information theory, the de Bruijn identity links the Fisher information and the derivative of the entropy. We show that this identity can be extended to generalized versions of en…
Market strategies minimize Fisher information to minimize risk.
The paper proves geometric and spectral alignment for deep neural networks.
Generative Adversarial Networks (GANs) are powerful models for learning complex distributions. Stable training of GANs has been addressed in many recent works which explore different metrics between distributions. In this paper we introduce Fisher GAN which fits within the Integral Probability Metrics (IPM) framework f…
Dead-Direction Signatures (DDS) provide a cheap, closed-form spectral reading of a network's singular complexity.
The study examines Fisher-Riemann geodesics for nonparametric probability densities.
We propose a modified -divergence, give some of its properties, and show that this leads to the definition of a generalized Fisher information. We give generalized Cramér-Rao inequalities, involving this Fisher information, an extension of the Fisher information matrix, and arbitrary norms and power of the estimat…
Blog post discusses various implementations of Fisher Information for EWC in continual learning.
Adaptive classifier optimizes high-dimensional data with spiked covariance structure.
Fisher width is a geometric measure of complexity on statistical manifolds.
The paper studies metrics on Lie groups SO(2) and SO(3) and their integrability.
This paper shows any Kähler metric can be a Fisher information metric.
The Delta method is a classical procedure for quantifying epistemic uncertainty in statistical models, but its direct application to deep neural networks is prevented by the large number of parameters . We propose a low cost variant of the Delta method applicable to -regularized deep neural networks based on th…
Proofs Fisher-Rao distance on Gaussian covariance manifold.
Survey on closed-form Fisher-Rao distance expressions.
New statistics are introduced that maintain the Fisher metric structure closely, akin to sufficient statistics.
We consider three different approaches to define natural Riemannian metrics on polytopes of stochastic matrices. First, we define a natural class of stochastic maps between these polytopes and give a metric characterization of Chentsov type in terms of invariance with respect to these maps. Second, we consider the Fish…
Fisher score is one of the most widely used supervised feature selection methods. However, it selects each feature independently according to their scores under the Fisher criterion, which leads to a suboptimal subset of features. In this paper, we present a generalized Fisher score to jointly select features. It aims …
We introduce Fisher consistency in the sense of unbiasedness as a desirable property for estimators of class prior probabilities. Lack of Fisher consistency could be used as a criterion to dismiss estimators that are unlikely to deliver precise estimates in test datasets under prior probability and more general dataset…
Kernel discriminant analysis uses nonlinear embeddings to improve classification.
New method estimates covariance matrices without restrictive assumptions.
We present a new method which generalizes subspace learning based on eigenvalue and generalized eigenvalue problems. This method, Roweis Discriminant Analysis (RDA), is named after Sam Roweis to whom the field of subspace learning owes significantly. RDA is a family of infinite number of algorithms where Principal Comp…
In information theory, Fisher information and Shannon information (entropy) are respectively used to quantify the uncertainty associated with the distribution modeling and the uncertainty in specifying the outcome of given variables. These two quantities are complementary and are jointly applied to information behavior…
Study on geometry of Dirichlet distributions using Fisher-Rao metric.
A deep neural network is a hierarchical nonlinear model transforming input signals to output signals. Its input-output relation is considered to be stochastic, being described for a given input by a parameterized conditional probability distribution of outputs. The space of parameters consisting of weights and biases i…