A framework connects Taylor methods with Fisher-efficient NG.
problem Combining Taylor-based methods with Fisher-efficient NG.
method Constructs a theoretical framework linking Taylor approximation and NG.
result Mathematical justification for combining higher order methods with NG.
Generative Adversarial Networks (GANs) are powerful models for learning complex distributions. Stable training of GANs has been addressed in many recent works which explore different metrics between distributions. In this paper we introduce Fisher GAN which fits within the Integral Probability Metrics (IPM) framework f…
Paper proposes efficient method to calculate Fisher-Bingham distribution normalizing constant.
problem Efficiently calculating the normalizing constant of Fisher-Bingham distributions.
method Numerical integration with continuous Euler transform to Fourier-type integral representation.
result The method is fast and accurate, applicable to high-dimensional distributions.
TopoFisher learns topological summaries by maximizing Fisher information, improving parameter efficiency and inference quality.
problem Simulation-based inference misses key information in low-order statistics, especially for non-Gaussian fields.
method TopoFisher uses a differentiable persistent-homology pipeline that learns topological summaries by maximizing local Gaussian Fisher information.
result TopoFisher recovers much of the available information and outperforms fixed topological vectorizations in weak gravitational lensing.
Proposes a new variational method using Fisher divergence for efficient Bayesian inference.
problem Intractable posterior distributions in complex models.
method Minimizes Fisher divergence to approximate posterior distributions efficiently.
result Demonstrates superior performance over traditional methods in logistic regression.
FedFisher improves one-shot FL by using Fisher information.
problem Reducing communication rounds in federated learning.
method Bayesian perspective, Fisher information matrices, diagonal Fisher, K-FAC approximation.
result FedFisher achieves vanishingly small error in two-layer neural networks.
Automates debiasing for large language model evaluations through Fisher random walk.
problem Rigorous and scalable evaluation of large language models.
method Semiparametric efficient estimator using Fisher random walk for weighted residual balancing.
result Efficient estimation of contextual preference scores for large language models.
A new method improves stochastic gradient descent for faster and more efficient estimation.
problem Efficient and fast parametric estimation methods.
method Projected stochastic gradient descent corrected by Fisher scoring.
result The method is faster and more efficient than traditional methods.
New algorithm speeds up optimization of Gaussian process models.
problem Optimizing Gaussian process models with many parameters.
method Fisher scoring with Vecchia's approximation.
result Significantly faster optimization with improved accuracy.
We introduce RSE to measure robustness in estimation problems.
problem Estimating statistical models from observed data.
method Developed theory for spectral functions of measures to compute RSE.
result RSE reveals a reciprocal relationship with problem complexity.
Examines learning efficiency in neural networks and related models.
problem Analyzing efficiency in deep learning models with singular learning coefficients.
method Examined learning coefficients in neural networks and three-layer neural networks with ReLU units.
result Extended results to include Softmax function, providing a broader understanding of learning efficiency.
Directly estimates Fisher score for likelihood maximization.
problem Intractable likelihood functions with model simulations.
method Gradient-based optimization using local score matching and linear parameterization.
result Efficient approximation of Fisher score improves likelihood maximization.
Paper proposes efficient optimizers for large language models with fast convergence and low memory usage.
problem Designing efficient optimizers for large language models with low-memory requirements and fast convergence.
method Structured Fisher information matrix approximation and low-rank extension framework.
result New optimizers (RACS and Alice) achieve better convergence and lower memory usage than existing methods.
FLoE adapts LLMs by selectively deploying LoRA adapters based on layer importance and task requirements.
problem Uniform LoRA deployment across all layers leads to inefficient and redundant parameter allocation.
method FLoE uses Fisher information to dynamically identify task-critical layers and optimizes LoRA ranks.
result FLoE achieves significant efficiency-accuracy trade-offs, especially in resource-constrained environments.
Optimal preconditioning improves Langevin sampling efficiency.
problem Improving sampling efficiency in high-dimensional target distributions.
method Optimal preconditioning using Fisher information, applied to MALA.
result Adaptive MCMC scheme significantly outperforms other methods.
M-FISHER detects and adapts to streaming data shifts with statistical validity and stability.
problem Detecting and adapting to distributional shifts in streaming data.
method Constructs an exponential martingale from non-conformity scores and applies Ville's inequality for detection. Fisher-preconditioned updates for adaptation.
result Establishes M-FISHER as a principled approach for robust, anytime-valid detection and geometrically stable adaptation.
To answer the existence of optimal swimmer learning/teaching strategies, this work introduces a two-level clustering in order to analyze temporal dynamics of motor learning in breaststroke swimming. Each level have been performed through Sparse Fisher-EM, a unsupervised framework which can be applied efficiently on lar…
Estimates metric tensor on neuromanifolds using Fisher information and random methods.
problem Computing the metric tensor on high-dimensional neuromanifolds efficiently and accurately.
method Deterministic bounds and unbiased random estimators based on Hutchinson's trace method.
result An efficient random estimator with bounded standard deviation.
Faster gaze prediction with less parameters.
problem Overparameterized networks for gaze prediction.
method Fisher pruning combined with knowledge distillation.
result 10x speedup for fixation prediction.
The paper uses Fisher information to explain and defend against adversarial attacks.
problem Vulnerability of deep learning models to adversarial attacks.
method Proposes OSSA for adversarial attack and eigenvalues for detection.
result Fisher information reveals model vulnerability through eigenvalues.
Framework for accelerated gradient flows in Bayesian inverse problems.
problem Design efficient MCMC algorithms for Bayesian inverse problems.
method Nesterov's accelerated gradient flows in probability space, considering various information metrics.
result Proved convergence properties and proposed sampling-efficient algorithms for different metrics.
New methods improve Fisher Matrix approximations for neural networks at low cost.
problem High cost of solving Fisher Information Matrix (FIM) in neural networks.
method Direct minimization via Kronecker product singular value decomposition.
result Improved approximations to FIM provide more accurate and faster optimization.
FADE adapts machine learning models to evolving data efficiently.
problem Sequential covariate shift in dynamic environments.
method FADE uses Fisher information geometry for robust learning under SCS.
result FADE achieves up to 19% higher accuracy under severe shifts.
This paper solves the intractability barrier in non-parametric information geometry by introducing a novel framework.
problem The intractability barrier in non-parametric information geometry due to the Fisher-Rao metric being a functional.
method Introducing an Orthogonal Decomposition of the Tangent Space and deriving the Covariate Fisher Information Matrix (cFIM).
result Established a rigorous foundation for the G-entropy and provided fundamental limits of variance for semi-parametric estimators.
Disputes the empirical Fisher approximation for natural gradient descent.
problem The empirical Fisher approximation fails to capture second-order information in general.
method Comparison of empirical Fisher and Fisher information matrices.
result The empirical Fisher does not generally approximate the Fisher or Hessian.
A new machine learning model uses score matching to estimate probability densities efficiently.
problem Estimating probability density functions is challenging.
method Introduced a product Jacobi-Theta Boltzmann machine (pJTBM) and used score matching for efficient fitting.
result The pJTBM can fit probability densities more efficiently than the RTBM using score matching.
Study on Gaussian-width complexity on statistical manifolds and its applications in learning and recovery.
problem Understanding the geometry of statistical manifolds and its implications for learning and recovery.
method Analysis of Fisher width and inverse-Fisher width, proving their complementary roles and establishing a relation between them.
result Established a sharp relation between Fisher width and inverse-Fisher width, showing they cannot reduce relative to Euclidean scale.
Novel method interprets black box predictions using Fisher kernels and SBQ.
problem Interpreting black box models' predictions.
method Fisher kernels as feature embeddings, Sequential Bayesian Quadrature for selection.
result Method efficiently handles any subset of test predictions.
Estimates intrinsic dimensionality of biological datasets using Fisher separability.
problem High-dimensional biological datasets with complex structures.
method Fisher separability analysis to estimate intrinsic dimensionality.
result The method performs competitively with state-of-the-art measures and is robust to noise.
Fisher auto-encoders use Fisher divergence for more robust generative modeling.
problem Model uncertainty in generative models.
method Minimizing Fisher divergence between true and modeled joint distributions.
result Fisher auto-encoders can more accurately quantify model uncertainty.
This paper introduces a neural sampler for scalable sampling from complex distributions.
problem Efficiently sampling from high-dimensional un-normalized distributions.
method Neural implicit sampler trained with KL and Fisher divergence methods.
result The neural sampler generates large batches of samples with low computational costs.
Paper defines Fisher co-metric on cotangent bundle and clarifies its relation to variance.
problem Defines Fisher co-metric on cotangent bundle and clarifies its relation to variance.
method Defines Fisher co-metric directly from Fisher metric without going through tangent bundle, using a natural correspondence between cotangent vectors and random variables.
result Clarifies the relation between Fisher co-metric and variance/covariance, trivializing the Cramér-Rao inequality.
Second-order optimization methods such as natural gradient descent have the potential to speed up training of neural networks by correcting for the curvature of the loss function. Unfortunately, the exact natural gradient is impractical to compute for large models, and most approximations either require an expensive it…
Explains Fisher and Kernel Fisher Discriminant Analysis with examples and comparisons.
problem Classifying data with different features and dimensions.
method Projection and reconstruction, scatters analysis, PCA comparison, Fisher forest.
result Equivalence of Fisher and Linear Discriminant Analysis, effectiveness of Fisher forest.
In this communication, we describe some interrelations between generalized q-entropies and a generalized version of Fisher information. In information theory, the de Bruijn identity links the Fisher information and the derivative of the entropy. We show that this identity can be extended to generalized versions of en…
Unified framework for two types of matrix Lie group preconditioners in SGD.
problem Improving optimization efficiency in machine learning models.
method Unified framework for Newton and Fisher type preconditioners on matrix Lie groups.
result Efficient estimation of preconditioners on matrix Lie groups.
This paper analyzes Barlow Twins' representation efficiency using information-geometric methods.
problem Understanding and comparing the efficiency of self-supervised learning methods.
method Introduces an information-geometric framework to quantify representation efficiency and applies it to Barlow Twins.
result Proves that Barlow Twins achieves optimal representation efficiency (η=1).
Optimal ability estimation in adaptive testing with binary responses.
problem Estimating a continuous ability parameter from sequential binary responses.
method Adaptive selection of questions to maximize Fisher information, updating estimate using method-of-moments, and deciding accuracy with a test statistic.
result Fisher-tracking strategy achieves optimal performance in fixed-confidence and fixed-budget regimes.
This paper explores VAEs in Fisher-Shannon plane, revealing the relationship between Fisher information and Shannon entropy.
problem Understanding the relationship between Fisher information and Shannon entropy in VAEs.
method Investigation of VAEs in Fisher-Shannon plane, focusing on the trade-off between Fisher information and Shannon entropy.
result VAEs' representation learning and log-likelihood estimation are intrinsically related to Fisher information and Shannon entropy.
Market strategies minimize Fisher information to minimize risk.
problem Applying minimum Fisher information principle to market dynamics.
method Analytical extension to quantum harmonic oscillator eigenstates and Gibbs distribution.
result Minimizing Fisher information reduces information and risk.
Geometric Variational Inference improves efficiency in complex probability distributions.
problem Efficiently accessing information in non-linear and high-dimensional probability distributions.
method Geometric Variational Inference (geoVI) uses Riemannian geometry and the Fisher information metric to construct a coordinate transformation.
result geoVI provides a more efficient variational approximation by a normal distribution, demonstrated on various problems.
Paper presents a rank-1 approximation method for natural policy gradients in deep RL.
problem Computing natural gradients requires inverting the Fisher Information Matrix, which is computationally expensive.
method Develops a rank-1 approximation to the inverse Fisher Information Matrix for efficient natural policy optimization.
result The rank-1 approximation converges faster and has similar sample complexity to stochastic policy gradient methods.
We propose a modified χβ-divergence, give some of its properties, and show that this leads to the definition of a generalized Fisher information. We give generalized Cramér-Rao inequalities, involving this Fisher information, an extension of the Fisher information matrix, and arbitrary norms and power of the estimat…
The study examines Fisher-Riemann geodesics for nonparametric probability densities.
problem Understanding nonparametric probability densities using Fisher-Riemann geometry.
method Obtaining Fisher-Riemann geodesics as a limit of parametric cases with increasing parameters.
result The weak limit approach for nonparametric probability densities.
Proposes a novel node embedding framework for graphs using Fisher Information.
problem Lack of theoretical understanding of attention-based GNNs.
method Uses hierarchical kernels and Fisher Information to learn node embeddings.
result Proposed method outperforms existing GNNs on node classification benchmarks.
Blog post discusses various implementations of Fisher Information for EWC in continual learning.
problem Improving Elastic Weight Consolidation (EWC) results by optimizing Fisher Information computation.
method Empirically compares different implementations of Fisher Information for EWC.
result Many reported EWC results can be improved by changing Fisher Information computation methods.
Study shows splitting schemes can approximate WFR flows faster than the exact flow.
problem Improving sampling efficiency in Wasserstein-Fisher-Rao gradient flows.
method Investigates operator splitting techniques to numerically approximate WFR flows.
result A judicious choice of step size and operator ordering can lead to faster convergence of split schemes to the target distribution.
The paper proposes a novel MKL approach for OCC using ℓp-norm constraints.
problem Addressing the MKL problem for one-class classification.
method A min-max saddle point Lagrangian optimisation problem is formulated and solved efficiently.
result The proposed method outperforms baselines and other algorithms on various data sets.