Explains Fisher and Kernel Fisher Discriminant Analysis with examples and comparisons.
problem Classifying data with different features and dimensions.
method Projection and reconstruction, scatters analysis, PCA comparison, Fisher forest.
result Equivalence of Fisher and Linear Discriminant Analysis, effectiveness of Fisher forest.
Paper explores Fisher-Rao gradient flows and their kernel approximations.
problem Understanding and analyzing approximations of Fisher-Rao gradient flows.
method Rigorous investigation of Fisher-Rao and Wasserstein type gradient flows, focusing on kernel approximations.
result Proves evolutionary Γ-convergence for kernel-approximated Fisher-Rao flows, providing theoretical guarantees.
The study examines Fisher information matrices and neural tangent kernels for simple ReLU networks with random weights.
problem Understanding the relationship between Fisher information matrices and neural tangent kernels for 2-layer ReLU networks.
method Analyzes Fisher information matrices and neural tangent kernels for 2-layer ReLU networks with random hidden weights, focusing on spectral decomposition and eigenfunctions.
result Obtained an approximation formula for functions represented by 2-layer neural networks.
Discriminative model identifies readers and assesses comprehension from eye movements.
problem Inferring readers' identities and estimating their text comprehension from eye movements.
method Generative model of gaze patterns, Fisher-score representation, Fisher-SVM with Fisher kernel.
result SVM with Fisher kernel excels at identifying readers, but not comprehending text.
Generative models of eye gaze help identify viewers from images.
problem Identifying viewers from images based on their eye movements.
method Derived Fisher kernels from generative models of eye gaze to train a discriminative classifier.
result Performance of the classifier improves with better underlying generative models.
Kernel networks' stability edge linked to Fisher Information singularity.
problem Understanding the stability edge in high-capacity kernel Hopfield networks.
method Statistical manifold analysis and Riemannian geometry.
result The Ridge of Optimization corresponds to the Edge of Stability, revealing a dual equilibrium.
Discusses MultiFIT for multivariate dependence, comparing it to HSIC tests.
problem Comparing Multiscale Fisher's Independence Test (MultiFIT) to HSIC tests for multivariate dependence.
method Compares MultiFIT to HSIC tests, highlighting exact level control and performance limitations.
result Observes performance limitations of MultiFIT in terms of test power.
The paper proposes a novel MKL approach for OCC using ℓp-norm constraints.
problem Addressing the MKL problem for one-class classification.
method A min-max saddle point Lagrangian optimisation problem is formulated and solved efficiently.
result The proposed method outperforms baselines and other algorithms on various data sets.
Novel method interprets black box predictions using Fisher kernels and SBQ.
problem Interpreting black box models' predictions.
method Fisher kernels as feature embeddings, Sequential Bayesian Quadrature for selection.
result Method efficiently handles any subset of test predictions.
Algebraic topology methods have recently played an important role for statistical analysis with complicated geometric structured data such as shapes, linked twist maps, and material data. Among them, \textit{persistent homology} is a well-known tool to extract robust topological features, and outputs as \textit{persist…
Fisher score is one of the most widely used supervised feature selection methods. However, it selects each feature independently according to their scores under the Fisher criterion, which leads to a suboptimal subset of features. In this paper, we present a generalized Fisher score to jointly select features. It aims …
This work develops a particle system to approximate Fisher-Rao gradient flows in mean-field optimization.
problem Optimizing probability measures in neural network contexts.
method Constructing an interacting particle system approximating Fisher-Rao gradient flows.
result Propagation of chaos for the Fisher-Rao gradient flow in entropic mean-field optimization.
Paper identifies key function spaces for ReLU networks based on Fisher information.
problem Understanding the structure of Fisher information matrices in ReLU networks.
method Spectral decomposition of Fisher information matrices, focusing on the first three eigenspaces.
result The first three eigenspaces account for 97.7% of the trace of the Fisher information matrix, corresponding to spherical harmonic functions of order ≤2.
A new weighted FDA method improves face recognition accuracy.
problem Equal treatment of all class pairs in FDA leads to suboptimal performance.
method Cosine-weighted and automatically weighted FDA methods are proposed.
result Improved face recognition accuracy through weighted FDA.
Generative model gradients enhance MS/MS peptide identification.
problem Improving peptide identification from MS/MS spectra.
method Leverage log-likelihood gradients of generative models in a kernel-based classifier.
result Fisher kernel outperforms other methods on MS/MS datasets.
This study explains why approximate NGD works well in wide neural networks.
problem Understanding why NGD with approximate Fisher information converges fast in wide neural networks.
method Analyzing asymptotic training dynamics in function space via the neural tangent kernel.
result NGD with approximate Fisher information achieves the same fast convergence as exact NGD under specific conditions.
Proposes a novel node embedding framework for graphs using Fisher Information.
problem Lack of theoretical understanding of attention-based GNNs.
method Uses hierarchical kernels and Fisher Information to learn node embeddings.
result Proposed method outperforms existing GNNs on node classification benchmarks.
The paper analyzes rates for a modified gradient descent method using Stein variational gradients.
problem Improving the accuracy of gradient descent methods for complex target distributions.
method Derives finite-particle rates for regularized Stein variational gradient descent (R-SVGD).
result Establishes explicit non-asymptotic bounds for time-averaged empirical measures.
We propose to investigate test statistics for testing homogeneity in reproducing kernel Hilbert spaces. Asymptotic null distributions under null hypothesis are derived, and consistency against fixed and local alternatives is assessed. Finally, experimental evidence of the performance of the proposed approach on both ar…
Kernel methods linked to feature subspaces and maximal correlation kernels.
problem Understanding kernel methods and their relationship to feature extraction.
method Established a correspondence between feature subspaces and kernels, introduced maximal correlation kernels, and demonstrated their optimality.
result Kernel SVM on maximal correlation kernel achieves minimum prediction error.
Deep networks can be biased to learn top eigenfunctions of the kernel outside the training set.
problem Spectral bias of deep networks in the kernel regime.
method Quantitative bounds on L2 difference between finite-width and infinite-width network trajectories. result Deep networks learn top eigenfunctions of the Neural Tangent Kernel over the entire input space, not just the training set.
Explains eigenvalue and generalized eigenvalue problems with examples.
problem Eigenvalue and generalized eigenvalue problems.
method Introduction and examples from machine learning.
result Solutions to eigenvalue and generalized eigenvalue problems.
Geometric regularisation improves statistical models by avoiding degeneracy loci.
problem Non-identifiability, singular information, and moment indeterminacy in statistical models.
method Develops the geometric regularisation of distribution-kernel pairs (T,φ) using Whitney, Thom, and Mather theorems. result Finite-dimensional weak transversality theorem for generic kernels, avoiding degeneracy strata of high codimension.
Batch Normalization (BN) is essential to effectively train state-of-the-art deep Convolutional Neural Networks (CNN). It normalizes inputs to the layers during training using the statistics of each mini-batch. In this work, we study BN from the viewpoint of Fisher kernels. We show that assuming samples within a mini-ba…
Framework for accelerated gradient flows in Bayesian inverse problems.
problem Design efficient MCMC algorithms for Bayesian inverse problems.
method Nesterov's accelerated gradient flows in probability space, considering various information metrics.
result Proved convergence properties and proposed sampling-efficient algorithms for different metrics.
Study evaluates posterior covariance matrix W for frequentist evaluation of Bayesian estimators.
problem Evaluating variability of posterior estimates in Bayesian models.
method Use of Bayesian Infinitesimal Jackknife approximation and W-kernel.
result Principal space of W is central to frequentist evaluation of Bayesian models.
USD algorithm transports distributions with or without mass conservation.
problem Transporting distributions with different masses.
method Particle descent algorithm using Sobolev-Fisher discrepancy.
result USD converges to target distribution in MMD sense.
Training-free source selection for LLM families with shared vocabularies
problem Source selection for LLM families with shared vocabularies
method Fisher alignment at vocabulary scale
result Fisher alignment is a cosine between kernel mean embeddings in the joint activation-error space
Enhances SVGD with matrix-valued kernels for faster inference.
problem Efficient approximate inference in complex probability landscapes.
method Integrates geometric information through matrix-valued kernels in SVGD.
result Significant improvement in real-world Bayesian inference tasks.
Kernel-Gradient Drifting improves generative modeling for non-Euclidean data.
problem Challenges in generative modeling for non-Euclidean data.
method Replaces Euclidean displacement with kernel-induced directions, exposing score-based structure.
result Kernel-gradient drifting enables state-of-the-art one-step generation for non-Euclidean data.
Quantum machine learning tackles large datasets with randomized measurements.
problem Efficiently process large, high-dimensional datasets on quantum computers.
method Randomized measurements to scale linearly with dataset size and quadratic for post-processing.
result Substantial speed-up for noisy quantum computers, enabling image classification.
Kernel discriminant analysis uses nonlinear embeddings to improve classification.
problem Limited effectiveness of linear discriminant analysis in capturing nonlinear features.
method Study of nonlinear embeddings in kernel discriminant analysis using polynomial and Gaussian kernels, solving generalized eigenvalue problems.
result Polynomial and Gaussian discriminants capture class differences through population moments and randomized projections.
Modified training direction reduces generalization error in neural networks.
problem Reducing generalization error in neural networks.
method Theoretical analysis of modified natural gradient descent in function space.
result Modifying training direction in function space reduces total generalization error.
Unified framework for spectral methods, kernel learning, and manifold unfolding.
problem Tackles the unification and optimization of spectral dimensionality reduction methods.
method Unified spectral methods as kernel PCA, kernel learning by SDP, and detailed explanation of MVU variants.
result Unified understanding and optimization of manifold learning techniques.
Disputes the empirical Fisher approximation for natural gradient descent.
problem The empirical Fisher approximation fails to capture second-order information in general.
method Comparison of empirical Fisher and Fisher information matrices.
result The empirical Fisher does not generally approximate the Fisher or Hessian.
A new method generalizing subspace learning for improved classification.
problem Improving classification accuracy using subspace learning methods.
method Roweis Discriminant Analysis (RDA) which generalizes PCA, SPCA, and FDA.
result RDA and kernel RDA improve classification accuracy on benchmark datasets.
New sampling method uses gradient-free IPS with RKHS velocity field.
problem Efficient sampling from unnormalized target densities.
method Gradient-free interacting particle systems (IPS) with RKHS velocity field.
result IPS produce high-quality samples from various target distributions.
Stein transport improves Bayesian inference with faster convergence and reduced variance.
problem Efficiently approximating posterior distributions in Bayesian inference.
method A novel Bayesian inference method using Stein transport, which pushes particles along a curve of tempered distributions.
result Stein transport reaches posterior approximations faster and more accurately than Stein variational gradient descent (SVGD).
Develops an analytic theory for quantum imaginary time evolution.
problem Lack of a first-principle understanding of quantum imaginary time evolution.
method Interprets QITE as a form of VQA trained with QNGD and connects it to the geometric geodesic distance in the quantum Fisher information metric.
result QITE converges faster than vanilla gradient descent-based VQAs, though the advantage is suppressed by Hilbert space dimensionality.
Physics-informed neural networks improve by measuring effective dimensionality of constraints.
problem Task interference in physics-informed neural networks due to shared parameter space.
method Introduce effective dimensionality (deff) as an operator invariant to quantify constraints. result Effective dimensionality measures unconstrained parameter directions, independent of network architecture.
Study on Gaussian-width complexity on statistical manifolds and its applications in learning and recovery.
problem Understanding the geometry of statistical manifolds and its implications for learning and recovery.
method Analysis of Fisher width and inverse-Fisher width, proving their complementary roles and establishing a relation between them.
result Established a sharp relation between Fisher width and inverse-Fisher width, showing they cannot reduce relative to Euclidean scale.
Fisher auto-encoders use Fisher divergence for more robust generative modeling.
problem Model uncertainty in generative models.
method Minimizing Fisher divergence between true and modeled joint distributions.
result Fisher auto-encoders can more accurately quantify model uncertainty.
Paper defines Fisher co-metric on cotangent bundle and clarifies its relation to variance.
problem Defines Fisher co-metric on cotangent bundle and clarifies its relation to variance.
method Defines Fisher co-metric directly from Fisher metric without going through tangent bundle, using a natural correspondence between cotangent vectors and random variables.
result Clarifies the relation between Fisher co-metric and variance/covariance, trivializing the Cramér-Rao inequality.
In this communication, we describe some interrelations between generalized q-entropies and a generalized version of Fisher information. In information theory, the de Bruijn identity links the Fisher information and the derivative of the entropy. We show that this identity can be extended to generalized versions of en…
The Fisher information metric is an important foundation of information geometry, wherein it allows us to approximate the local geometry of a probability distribution. Recurrent neural networks such as the Sequence-to-Sequence (Seq2Seq) networks that have lately been used to yield state-of-the-art performance on speech…
Market strategies minimize Fisher information to minimize risk.
problem Applying minimum Fisher information principle to market dynamics.
method Analytical extension to quantum harmonic oscillator eigenstates and Gibbs distribution.
result Minimizing Fisher information reduces information and risk.
Paper analyzes SVGD algorithm for non-asymptotic convergence.
problem Optimizing a set of particles to approximate a target probability distribution.
method Finite time analysis of SVGD algorithm, providing descent lemma and convergence rates.
result SVGD algorithm decreases the objective at each iteration and converges to the target distribution.
Extends OC-KSR for multi-task one-class classification.
problem Improving one-class classification performance with shared information.
method Linear and non-linear structure learning mechanisms for multi-task one-class classification.
result Improved performance on multiple one-class problems.