Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

6.3%12.5%18.8%25.0% · Apr 199419922001200920182026
48 results for preconditioned Fisher divergence

New method for fast, accurate Gaussian process regression with data guarantees.

problem Slow and unreliable GP inference methods for nonparametric regression.
method Developed a novel objective function (preconditioned Fisher divergence) for scalable approximate GP regression with finite-data guarantees.
result Minimizing the pF divergence provides pointwise mean and variance estimates with tight 2-Wasserstein distance bounds and comparable empirical performance to variational sparse GPs.

M-FISHER detects and adapts to streaming data shifts with statistical validity and stability.

problem Detecting and adapting to distributional shifts in streaming data.
method Constructs an exponential martingale from non-conformity scores and applies Ville's inequality for detection. Fisher-preconditioned updates for adaptation.
result Establishes M-FISHER as a principled approach for robust, anytime-valid detection and geometrically stable adaptation.

Disputes the empirical Fisher approximation for natural gradient descent.

problem The empirical Fisher approximation fails to capture second-order information in general.
method Comparison of empirical Fisher and Fisher information matrices.
result The empirical Fisher does not generally approximate the Fisher or Hessian.

Adaptively preconditions SGLD for faster convergence and better generalization.

problem Pathological curvature in deep network loss landscapes.
method Adaptive estimation of noise parameters to precondition isotropic gradient noise.
result Adaptively preconditioned SGLD achieves faster convergence and generalization equivalent of SGD.

One way to avoid overfitting in machine learning is to use model parameters distributed according to a Bayesian posterior given the data, rather than the maximum likelihood estimator. Stochastic gradient Langevin dynamics (SGLD) is one algorithm to approximate such Bayesian posteriors for large models and datasets. SGL…

2017-12-04abs ↗pdf ↗

Preconditioned NFs speed up sampling from complex posterior distributions in inverse problems.

problem Sampling from posterior distributions of inverse problems with expensive forward operators.
method Preconditioning a conditional normalizing flow (NF) to speed up training.
result Significant speed-ups achieved compared to training NFs from scratch.

FOP improves deep learning optimizers with minimal computational overhead.

problem Training deep learning models can be hindered by high correlations and different scaling in parameter space.
method FOP uses first-order information to learn a preconditioning matrix that improves convergence without the high computational cost of second-order methods.
result FOP improves performance of standard deep learning optimizers on visual classification and reinforcement learning tasks.

The paper explores geometry of probability measures and barycenter maps.

problem Understanding the space of probability measures and their barycenter.
method Information geometry, Fisher metric, dualistic structures, divergences, geodesics.
result Recent developments in the geometry of probability measures and barycenter.

The paper proposes a new method to approximate Wasserstein-Fisher-Rao flows using Monte Carlo techniques.

problem Sampling from probability distributions and minimizing Kullback-Leibler divergence.
method Sequential Monte Carlo approximations of Wasserstein-Fisher-Rao gradient flows.
result The proposed method outperforms other Monte Carlo algorithms in certain conditions.

High-dimensional models become unstable when sample size falls below a critical level, leading to a phase transition.

problem Instability in high-dimensional learning models when sample size is insufficient.
method Proved the necessity of a Fisher eigenvalue threshold for stability, introduced Fisher floor for verification.
result A sharp phase transition between reliable concentration and inevitable failure in high-dimensional learning.

Paper analyzes Langevin dynamics for multimodal Gaussian mixtures, controlling errors across dimensions.

problem Challenges in obtaining stable diffusion-based samplers in high- and infinite-dimensional settings.
method Study of preconditioned Annealed Langevin Dynamics (ALD) for Gaussian mixtures, focusing on Euler-Maruyama (EM) and exponential-integrator schemes.
result Proves dimension-uniform KL bounds for the exponential-integrator scheme, allowing arbitrarily small divergence with dimension.

This study provides an explicit expansion of KL divergence's gradient flow in Fisher-Rao geometry.

problem Sampling techniques struggle to traverse between modes in non-convex potential functions.
method Explicit expansion of KL divergence's gradient flow in Fisher-Rao geometry.
result The convergence rate to π is independent of the potential function.

Dual Space Preconditioning speeds up gradient descent in overparameterized models.

problem Improving convergence of gradient descent in overparameterized linear models.
method Introducing a novel preconditioner of the form ablaK abla K for convex KK and applying it to overparameterized linear models.
result The iterates of the preconditioned gradient descent converge to a solution W{W}_{\infty} satisfying XW=Y{X}{W}_{\infty} = {Y}.

Researchers study the geometric properties of a specific type of stable processes.

problem Understanding the information geometry of tempered stable processes.
method Derivation of α-divergence, Fisher information matrices, and α-connections.
result Obtained Fisher information matrices and α-connections for statistical manifolds.

This paper introduces a neural sampler for scalable sampling from complex distributions.

problem Efficiently sampling from high-dimensional un-normalized distributions.
method Neural implicit sampler trained with KL and Fisher divergence methods.
result The neural sampler generates large batches of samples with low computational costs.

Develops information geometry for Lévy processes in finance.

problem Understanding the statistical properties of Lévy processes for financial modeling.
method Deriving α\alpha-divergences from Lévy triplets, identifying Fisher information matrix and α\alpha-connection.
result Identifies statistical implications and differential-geometric structures of Lévy processes.

This work analyzes how preconditioning affects generalization in machine learning models.

problem The impact of preconditioning on the generalization of machine learning models.
method An asymptotic bias-variance decomposition of the generalization error for ridgeless regression under various preconditioners.
result The optimal preconditioner depends on label noise, model specification, and signal alignment, with NGD potentially better under certain conditions.

New particle-based VI algorithm expands function class and improves scalability.

problem Limited function class in particle-based VI algorithms restricts flexibility and scalability.
method Introduces a functional regularization term to expand the function class and proposes PFG algorithm.
result Proposed PFG algorithm has larger function class, improved scalability, better adaptation to ill-conditioned distributions, and provable convergence.

Study of generalized Csiszár divergences and their application to Cramér-Rao bounds.

problem Deriving lower bounds for estimator variance using generalized divergences.
method Applied Eguchi's theory to derive Fisher information metric and dual affine connections.
result More widely applicable Cramér-Rao inequality for escort distributions.

A new machine learning model uses score matching to estimate probability densities efficiently.

problem Estimating probability density functions is challenging.
method Introduced a product Jacobi-Theta Boltzmann machine (pJTBM) and used score matching for efficient fitting.
result The pJTBM can fit probability densities more efficiently than the RTBM using score matching.

A new optimization method reduces memory and compute requirements for deep learning.

problem Memory and compute constraints in second-order stochastic optimizers for deep learning.
method Proposes KrAD, a novel factorization to approximate inverse Fisher matrix without inversion, leading to KrADagrad.
result Improves performance over Shampoo for 32-bit precision and comparable/generalization on real datasets.

Study on convergence rates of degenerate SDEs using Fisher information and generalized Bochner's formula.

problem Analysis of dynamical behaviors of degenerate stochastic differential equations.
method Use of Fisher information as Lyapunov functional, generalized Gamma calculus, and generalized Bochner's formula.
result Derivation of convergence rate conditions and examples in specific sub-Riemannian structures.

The paper connects tempering and entropic mirror descent for sampling.

problem Sampling from a target distribution with known unnormalized density.
method Establishes the connection between tempering SMC and entropic mirror descent, deriving convergence rates and geometric insights.
result Tempering SMC iterates correspond to entropic mirror descent on the reverse KL divergence, providing new optimization perspectives.

We unify f-divergences, Bregman divergences, surrogate loss bounds (regret bounds), proper scoring rules, matching losses, cost curves, ROC-curves and information. We do this by systematically studying integral and variational representations of these objects and in so doing identify their primitives which all are rela…

2009-01-05abs ↗pdf ↗

Paper formulates particle flow using variational inference and Fisher-Rao gradient flow.

problem Estimating posterior densities in probabilistic models.
method Variational formulation of particle flow, Fisher-Rao gradient flow, Gaussian and Gaussian mixture approximations.
result Gaussian and Gaussian mixture approximations of Fisher-Rao particle flow reduce to Exact Daum and Huang particle flow under linear Gaussian assumptions.

Study improves sampling from non-log-concave distributions using Fisher information.

problem Sampling from non-log-concave distributions with high Fisher information guarantees.
method Proximal sampler with RGO implementation, leveraging log-concave sampling results.
result Improved complexity guarantee in relative Fisher information for non-log-concave sampling.

New method estimates covariance matrices without restrictive assumptions.

problem Estimating high-dimensional covariance matrices under restrictive assumptions.
method Distributionally robust covariance estimation problems with mild conditions.
result Robust estimators are efficient, consistent, and perform well.

Squared families are a new model class derived from linear transformations, offering convenient properties and universal approximation.

problem Developing a new class of probability models that are easier to handle and have useful properties.
method Introducing squared families as families of probability densities obtained by squaring a linear transformation of a statistic, and showing their properties and applications.
result Squared families have convenient properties and can approximate target densities well.

This paper solves the intractability barrier in non-parametric information geometry by introducing a novel framework.

problem The intractability barrier in non-parametric information geometry due to the Fisher-Rao metric being a functional.
method Introducing an Orthogonal Decomposition of the Tangent Space and deriving the Covariate Fisher Information Matrix (cFIM).
result Established a rigorous foundation for the G-entropy and provided fundamental limits of variance for semi-parametric estimators.

Unified framework improves neural network robustness against label noise and adversarial attacks.

problem High sensitivity of neural networks to data contamination, including label noises and adversarial perturbations.
method Unified minimum-divergence estimation problem, rSDNet framework.
result Improves robustness to label corruption and adversarial attacks while maintaining competitive accuracy on clean data.

Optimistic likelihoods improve classification accuracy by considering nearby distributions.

problem Evaluating likelihoods of nominal distributions estimated from data, which can be inaccurate.
method Use ambiguity sets and geodesic/standard convex optimization to compute optimistic likelihoods.
result Optimistic likelihoods lead to better classification performance.

Paper analyzes latent space geometry in generative models using Fisher information.

problem Understanding the structure of latent spaces in generative models.
method Reconstructs Fisher information metric from generated samples and posterior distribution.
result Reveals fractal structure and abrupt changes in Fisher metric at phase boundaries.