This paper solves the intractability barrier in non-parametric information geometry by introducing a novel framework.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper solves the normalizability crisis in sequential inference by introducing bounded information geometry.
We consider the problems of clustering, classification, and visualization of high-dimensional data when no straightforward Euclidean representation exists. Typically, these tasks are performed by first reducing the high-dimensional data to some lower dimensional Euclidean space, as many manifold learning methods have b…
Estimates non-parametric logistic model using case-control data and external summary info.
We propose a method for learning Markov network structures for continuous data without invoking any assumptions about the distribution of the variables. The method makes use of previous work on a non-parametric estimator for mutual information which is used to create a non-parametric test for multivariate conditional i…
The paper develops a theory for identifying the best arm in non-parametric multi-armed bandits with a fixed budget.
Estimates classifier errors without ground truth using algebraic geometry.
Study uses non-parametric method to analyze EU ETS price determinants.
DPPS uses DP priors for Bayesian non-parametric multi-arm bandits.
Proposes using DII to identify non-linear causal relationships in EU Allowances returns.
A tractable pseudo-metric for non-parametric distributions via SPD geometry.
In this paper, we suggest a framework to make use of mutual information as a regularization criterion to train Auto-Encoders (AEs). In the proposed framework, AEs are regularized by minimization of the mutual information between input and encoding variables of AEs during the training phase. In order to estimate the ent…
New algorithm learns nonlinear phenomena from noisy local measurements without data exchange.
Study models Indian stock market using hyperbolic geometry for market stability and volatility analysis.
This article reviews the Author-Topic Model and presents a new non-parametric extension based on the Hierarchical Dirichlet Process. The extension is especially suitable when no prior information about the number of components necessary is available. A blocked Gibbs sampler is described and focus put on staying as clos…
Non-parametric time series forecasting without assuming a specific distribution.
A graph-based method for two-sample testing across connected nodes.
This paper finds a unique partition of a sample space for estimating continuous distributions.
Unified data representation learning improves non-parametric two-sample testing.
Datasets with hundreds of variables and many missing values are commonplace. In this setting, it is both statistically and computationally challenging to detect true predictive relationships between variables and also to suppress false positives. This paper proposes an approach that combines probabilistic programming, …
iCOS method estimates risk-neutral densities and option prices without model assumptions.
Formulates mechanics for probability distributions on statistical manifold.
The method of covariate adjustment is often used for estimation of population average treatment effects in observational studies. Graphical rules for determining all valid covariate adjustment sets from an assumed causal graphical model are well known. Restricting attention to causal linear models, a recent article der…
Neural Networks trained with gradient descent are known to be susceptible to catastrophic forgetting caused by parameter shift during the training process. In the context of Neural Machine Translation (NMT) this results in poor performance on heterogeneous datasets and on sub-tasks like rare phrase translation. On the …
The Mutual Information (MI) is an often used measure of dependency between two random variables utilized in information theory, statistics and machine learning. Recently several MI estimators have been proposed that can achieve parametric MSE convergence rate. However, most of the previously proposed estimators have th…
Framework predicts Navier-Stokes solutions on 2D domains using graph neural networks.
Complex analytic sets' Lipschitz geometry at infinity characterized.
The Fisher information matrix (FIM) is a foundational concept in statistical signal processing. The FIM depends on the probability distribution, assumed to belong to a smooth parametric family. Traditional approaches to estimating the FIM require estimating the probability distribution function (PDF), or its parameters…
The identification of relevant features, i.e., the driving variables that determine a process or the properties of a system, is an essential part of the analysis of data sets with a large number of variables. A mathematical rigorous approach to quantifying the relevance of these features is mutual information. Mutual i…
In order to study the geometry of interest rates market dynamics, Malliavin, Mancino and Recchioni [A non-parametric calibration of the HJM geometry: an application of Itô calculus to financial statistics, {\it Japanese Journal of Mathematics}, 2, pp.55--77, 2007] introduced a scheme, which is based on the Fourier Seri…
Automatically counts microglial cells in rat spinal cord images, providing precise counts and uncertainty estimates.
Meta learning with information theory and Gaussian processes.
KSD Thinning uses KSD to thin MCMC samples efficiently.
We present a dual-view mixture model to cluster users based on their features and latent behavioral functions. Every component of the mixture model represents a probability density over a feature view for observed user attributes and a behavior view for latent behavioral functions that are indirectly observed through u…
Data collection at a massive scale is becoming ubiquitous in a wide variety of settings, from vast offline databases to streaming real-time information. Learning algorithms deployed in such contexts must rely on single-pass inference, where the data history is never revisited. In streaming contexts, learning must also …
Ensembles of classification and regression trees remain popular machine learning methods because they define flexible non-parametric models that predict well and are computationally efficient both during training and testing. During induction of decision trees one aims to find predicates that are maximally informative …
Proposes method for eliciting non-parametric joint priors using normalizing flows.
The Information bottleneck method is an unsupervised non-parametric data organization technique. Given a joint distribution P(A,B), this method constructs a new variable T that extracts partitions, or clusters, over the values of A that are informative about B. The information bottleneck has already been applied to doc…
This paper investigates the use of methods from partial differential equations and the Calculus of variations to study learning problems that are regularized using graph Laplacians. Graph Laplacians are a powerful, flexible method for capturing local and global geometry in many classes of learning problems, and the tec…
UMAP connects to Information Geometry principles.
Active learning recovers choice model from noisy data.
Symplectic and Poisson structures proved for information geometry's Frobenius manifold.
Information geometry offers new tools for statistical analysis.
Adversarially robust machine learning has received much recent attention. However, prior attacks and defenses for non-parametric classifiers have been developed in an ad-hoc or classifier-specific basis. In this work, we take a holistic look at adversarial examples for non-parametric classifiers, including nearest neig…
Information divergence functions play a critical role in statistics and information theory. In this paper we show that a non-parametric f-divergence measure can be used to provide improved bounds on the minimum binary classification probability of error for the case when the training and test data are drawn from the sa…
Parametric UMAP learns a mapping from data to embeddings.
Develops information geometry for Lévy processes in finance.
Survey of linking information geometry and optimal transport.