Paper introduces new loss functions for Siamese networks using FDA.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Fisher loss improves deep domain adaptation by learning discriminative within-class compact and between-class separable representations.
New statistics are introduced that maintain the Fisher metric structure closely, akin to sufficient statistics.
In many structured prediction problems, complex relationships between variables are compactly defined using graphical structures. The most prevalent graphical prediction methods---probabilistic graphical models and large margin methods---have their own distinct strengths but also possess significant drawbacks. Conditio…
Optimal domain adaptation model using Fisher's Linear Discriminant.
Early training phase affects deep neural network optimization and generalization.
We consider composite loss functions for multiclass prediction comprising a proper (i.e., Fisher-consistent) loss over probability distributions and an inverse link function. We establish conditions for their (strong) convexity and explore the implications. We also show how the separation of concerns afforded by using …
Score matching fails to train VAEs robustly, revealing autoencoding loss insights.
The paper explores conditions for predicting optimization performance.
We analyze the variance of Fisher information estimators in deep learning models.
TopoFisher learns topological summaries by maximizing Fisher information, improving parameter efficiency and inference quality.
Proposes variational Gaussian approximations for solving the Kushner equation.
FIRE method improves model performance in federated learning by penalizing fragmentation-induced covariate shifts.
Optimal convex loss function improves regression coefficient estimation.
This paper proposes a new subspace learning method, named Quantized Fisher Discriminant Analysis (QFDA), which makes use of both machine learning and information theory. There is a lack of literature for combination of machine learning and information theory and this paper tries to tackle this gap. QFDA finds a subspac…
While enormous progress has been made to Variational Autoencoder (VAE) in recent years, similar to other deep networks, VAE with deep networks suffers from the problem of degeneration, which seriously weakens the correlation between the input and the corresponding latent codes, deviating from the goal of the representa…
Study detects boundaries in unlabeled noisy images without labels.
A new method improves uncertainty estimation in deep learning, especially for hard-to-label samples.
Survey of spectral, probabilistic, and deep metric learning methods.
We propose a robust adversarial prediction framework for general multiclass classification. Our method seeks predictive distributions that robustly optimize non-convex and non-continuous multiclass loss metrics against the worst-case conditional label distributions (the adversarial distributions) that (approximately) m…
Information geometry provides a geometric approach to families of statistical models. The key geometric structures are the Fisher quadratic form and the Amari-Chentsov tensor. In statistics, the notion of sufficient statistic expresses the criterion for passing from one model to another without loss of information. Thi…
Fisher auto-encoders use Fisher divergence for more robust generative modeling.
Natural gradient descent, which preconditions a gradient descent update with the Fisher information matrix of the underlying statistical model, is a way to capture partial second-order information. Several highly visible works have advocated an approximation known as the empirical Fisher, drawing connections between ap…
Paper defines Fisher co-metric on cotangent bundle and clarifies its relation to variance.
Second-order optimization methods such as natural gradient descent have the potential to speed up training of neural networks by correcting for the curvature of the loss function. Unfortunately, the exact natural gradient is impractical to compute for large models, and most approximations either require an expensive it…
In this paper we revisit the weighted likelihood bootstrap, a method that generates samples from an approximate Bayesian posterior of a parametric model. We show that the same method can be derived, without approximation, under a Bayesian nonparametric model with the parameter of interest defined as minimising an expec…
We unify f-divergences, Bregman divergences, surrogate loss bounds (regret bounds), proper scoring rules, matching losses, cost curves, ROC-curves and information. We do this by systematically studying integral and variational representations of these objects and in so doing identify their primitives which all are rela…
In this communication, we describe some interrelations between generalized -entropies and a generalized version of Fisher information. In information theory, the de Bruijn identity links the Fisher information and the derivative of the entropy. We show that this identity can be extended to generalized versions of en…
Overparametrization improves QNN trainability by reducing spurious local minima.
Market strategies minimize Fisher information to minimize risk.
We study online learning under logarithmic loss with regular parametric models. Hedayati and Bartlett (2012b) showed that a Bayesian prediction strategy with Jeffreys prior and sequential normalized maximum likelihood (SNML) coincide and are optimal if and only if the latter is exchangeable, and if and only if the opti…
Generative Adversarial Networks (GANs) are powerful models for learning complex distributions. Stable training of GANs has been addressed in many recent works which explore different metrics between distributions. In this paper we introduce Fisher GAN which fits within the Integral Probability Metrics (IPM) framework f…
The study examines Fisher-Riemann geodesics for nonparametric probability densities.
We propose a modified -divergence, give some of its properties, and show that this leads to the definition of a generalized Fisher information. We give generalized Cramér-Rao inequalities, involving this Fisher information, an extension of the Fisher information matrix, and arbitrary norms and power of the estimat…
Blog post discusses various implementations of Fisher Information for EWC in continual learning.
Fisher width is a geometric measure of complexity on statistical manifolds.
The paper studies metrics on Lie groups SO(2) and SO(3) and their integrability.
This paper shows any Kähler metric can be a Fisher information metric.
The speed at which one can minimize an expected loss using stochastic methods depends on two properties: the curvature of the loss and the variance of the gradients. While most previous works focus on one or the other of these properties, we explore how their interaction affects optimization speed. Further, as the ulti…
Proofs Fisher-Rao distance on Gaussian covariance manifold.
In a graph convolutional network, we assume that the graph is generated wrt some observation noise. During learning, we make small random perturbations of the graph and try to improve generalization. Based on quantum information geometry, can be characterized by the eigendecomposition of the graph Laplaci…
Survey on closed-form Fisher-Rao distance expressions.
Batch Normalization (BatchNorm) is an extremely useful component of modern neural network architectures, enabling optimization using higher learning rates and achieving faster convergence. In this paper, we use mean-field theory to analytically quantify the impact of BatchNorm on the geometry of the loss landscape for …
We consider three different approaches to define natural Riemannian metrics on polytopes of stochastic matrices. First, we define a natural class of stochastic maps between these polytopes and give a metric characterization of Chentsov type in terms of invariance with respect to these maps. Second, we consider the Fish…
Fisher score is one of the most widely used supervised feature selection methods. However, it selects each feature independently according to their scores under the Fisher criterion, which leads to a suboptimal subset of features. In this paper, we present a generalized Fisher score to jointly select features. It aims …
This paper improves multi-label ranking by reweighting univariate losses, enhancing consistency and performance.
We introduce Fisher consistency in the sense of unbiasedness as a desirable property for estimators of class prior probabilities. Lack of Fisher consistency could be used as a criterion to dismiss estimators that are unlikely to deliver precise estimates in test datasets under prior probability and more general dataset…
The Fisher information matrix (FIM) plays an essential role in statistics and machine learning as a Riemannian metric tensor or a component of the Hessian matrix of loss functions. Focusing on the FIM and its variants in deep neural networks (DNNs), we reveal their characteristic scale dependence on the network width, …