A new method uses natural gradients for efficient distribution optimization.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We show that the distribution of symmetry of a naturally reductive nilpotent Lie group coincides with the invariant distribution induced by the set of fixed vectors of the isotropy. This extends a known result on compact naturally reductive spaces. We also address the study of the quotient by the foliation of symmetry.
Study shows current image classification models lack robustness to real-world dataset shifts.
A new copula minimizes distance between distributions.
Hardness proven for neural networks with natural weights.
Study of symmetry distributions in Lorentzian naturally reductive nilmanifolds.
A new method generates natural-looking adversarial examples by bounding internal activation values.
In optimization, the natural gradient method is well-known for likelihood maximization. The method uses the Kullback-Leibler divergence, corresponding infinitesimally to the Fisher-Rao metric, which is pulled back to the parameter space of a family of probability distributions. This way, gradients with respect to the p…
New method for natural policy gradients converges linearly.
A method for converting NIW parameters for better estimation.
Two novel distributed VB algorithms improve Bayesian inference in sensor networks.
This paper presents Natural Evolution Strategies (NES), a recent family of algorithms that constitute a more principled approach to black-box optimization than established evolutionary algorithms. NES maintains a parameterized distribution on the set of solution candidates, and the natural gradient is used to update th…
This is a lecture note for the course DS-GA 3001 <Natural Language Understanding with Distributed Representation> at the Center for Data Science , New York University in Fall, 2015. As the name of the course suggests, this lecture note introduces readers to a neural network based approach to natural language understand…
Paper characterizes DLN distribution, its properties, and estimation methods.
Natural-gradient methods enable fast and simple algorithms for variational inference, but due to computational difficulties, their use is mostly limited to \emph{minimal} exponential-family (EF) approximations. In this paper, we extend their application to estimate \emph{structured} approximations such as mixtures of E…
Learning the distribution of natural images is one of the hardest and most important problems in machine learning. The problem remains open, because the enormous complexity of the structures in natural images spans all length scales. We break down the complexity of the problem and show that the hierarchy of structures …
This work proves intrinsic robustness bounds for natural image distributions.
Model uses LLMs to process numerical data guided by natural language descriptions.
Study evaluates how well question-answering models generalize to new data types.
Neural networks (NN) have achieved state-of-the-art performance in various applications. Unfortunately in applications where training data is insufficient, they are often prone to overfitting. One effective way to alleviate this problem is to exploit the Bayesian approach by using Bayesian neural networks (BNN). Anothe…
A method for diffusion on probability simplex for generative models.
Federated learning studies separate client data and distribution gaps.
Enhanced transformer converts whispered speech to natural speech.
Energy-based models can generate complex images by combining simpler concepts.
Mitigates gender bias amplification in model predictions.
This paper makes two contributions to Bayesian machine learning algorithms. Firstly, we propose stochastic natural gradient expectation propagation (SNEP), a novel alternative to expectation propagation (EP), a popular variational inference algorithm. SNEP is a black box variational algorithm, in that it does not requi…
A method for predicting credal sets in classification tasks using conformal prediction.
Geospatial analysis lacks methods like the word vector representations and pre-trained networks that significantly boost performance across a wide range of natural language and computer vision tasks. To fill this gap, we introduce Tile2Vec, an unsupervised representation learning algorithm that extends the distribution…
We show that the unique 7th order ODE having 10 contact symmetries appears naturally in the theory of generic 2-distributions in dimension five.
We present Natural Gradient Boosting (NGBoost), an algorithm for generic probabilistic prediction via gradient boosting. Typical regression models return a point estimate, conditional on covariates, but probabilistic regression models output a full probability distribution over the outcome space, conditional on the cov…
NatPN provides fast, accurate uncertainty estimation for exponential family distributions.
We solve the mean parametrization of von Mises-Fisher distribution.
Paper develops a gradient-like proposal for discrete distributions without requiring natural differentiability.
The paper covers the new model of wage distribution in typical group of people. The model provides the opportunity to reparameterize applicable income distribution model: Pareto, logarithmically normal, logarithmically logistic, Dagum etc. The model ensures the graduation of Gini index values by polynomial degree of wa…
We study the Wasserstein natural gradient in parametric statistical models with continuous sample spaces. Our approach is to pull back the -Wasserstein metric tensor in the probability density space to a parameter space, equipping the latter with a positive definite metric tensor, under which it becomes a Riemanni…
We propose a robust estimator to improve maximum likelihood in probabilistic models.
Using a relationship between the moments of the probability distribution of times between the two consecutive trades (intertrade time distribution) and the moments of the distribution of a daily number of trades we show, that the underlying point process generating times of the trades is an essentially non-markovian lo…
DDPM encoder matches optimal transport for natural images.
We describe a model for capturing the statistical structure of local amplitude and local spatial phase in natural images. The model is based on a recently developed, factorized third-order Boltzmann machine that was shown to be effective at capturing higher-order structure in images by modeling dependencies among squar…
We consider three different approaches to define natural Riemannian metrics on polytopes of stochastic matrices. First, we define a natural class of stochastic maps between these polytopes and give a metric characterization of Chentsov type in terms of invariance with respect to these maps. Second, we consider the Fish…
The study models and forecasts natural gas prices using skewed, heavy-tailed distributions.
ADT improves model robustness by learning adversarial distributions.
The distribution of money is analysed in connection with the Boltzmann distribution of energy in the degenerate states of molecules. Plots of the population density of income distribution for various countries are well reproduced by a Gamma function, confirming the validity of the statistical distribution at equilibriu…
We consider automorphisms of homogeneous parabolic geometries with a fixed point. Parabolic geometries carry the distinguished distributions and we study those automorphisms which enjoy natural actions on the distributions at the fixed points. We describe the sets of such automorphisms on homogeneous parabolic geometri…
When modeling a probability distribution with a Bayesian network, we are faced with the problem of how to handle continuous variables. Most previous work has either solved the problem by discretizing, or assumed that the data are generated by a single Gaussian. In this paper we abandon the normality assumption and inst…
Exponential family distributions are highly useful in machine learning since their calculation can be performed efficiently through natural parameters. The exponential family has recently been extended to the t-exponential family, which contains Student-t distributions as family members and thus allows us to handle noi…
A four-dimensional Walker geometry is a four-dimensional manifold M with a neutral metric g and a parallel distribution of totally null two-planes. This distribution has a natural characterization as a projective spinor field subject to a certain constraint. Spinors therefore provide a natural tool for studying Walker …
Paper improves online time series forecasting by combining natural gradient and robust t-distribution.