Gaussian processes (GPs) offer a flexible class of priors for nonparametric Bayesian regression, but popular GP posterior inference methods are typically prohibitively slow or lack desirable finite-data guarantees on quality. We develop an approach to scalable approximate GP regression with finite-data guarantees on th…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Improves Laplace approximation for Bayesian inference on Riemannian manifolds.
Many modern data analysis problems involve inferences from streaming data. However, streaming data is not easily amenable to the standard probabilistic modeling approaches, which assume that we condition on finite data. We develop population variational Bayes, a new approach for using Bayesian modeling to analyze strea…
Meta-learning improves Bayesian causal discovery by sampling from the posterior.
Estimates causal effects in Gaussian Linear SCMs with finite data.
Bayesian inference typically requires the computation of an approximation to the posterior distribution. An important requirement for an approximate Bayesian inference algorithm is to output high-accuracy posterior mean and uncertainty estimates. Classical Monte Carlo methods, particularly Markov Chain Monte Carlo, rem…
Variational Causal Networks approximate Bayesian inference over causal structures.
We consider 1-qubit mixed quantum state estimation by adaptively updating measurements according to previously obtained outcomes and measurement settings. Updates are determined by the average-variance-optimality (A-optimality) criterion, known in the classical theory of experimental design and applied here to quantum …
Computing accurate estimates of the Fourier transform of analog signals from discrete data points is important in many fields of science and engineering. The conventional approach of performing the discrete Fourier transform of the data implicitly assumes periodicity and bandlimitedness of the signal. In this paper, we…
XGES improves GES by favoring early edge deletion, outperforming GES in finite data settings.
Improved neural network regression uncertainty estimation.
Measuring mutual information from finite data is difficult. Recent work has considered variational methods maximizing a lower bound. In this paper, we prove that serious statistical limitations are inherent to any method of measuring mutual information. More specifically, we show that any distribution-free high-confide…
We introduce a new class of possibly noncompact n-dimensional manifolds without boundary associated to finite data which we call topological automata. This class is large enough to contain many interesting examples of open 2-dimensional and 3-dimensional manifolds of interest to low-dimensional topologists. Our main re…
PACC Discovery improves causal inference from limited data.
The Bellman error is a poor proxy for value function accuracy, even with all state-action pairs.
This work introduces significativity indices for agreement values between classifiers.
We identify linear models from nonlinear systems with initialization constraints.
Feature selection aims to select the smallest subset of features for a specified level of performance. The optimal achievable classification performance on a feature subset is summarized by its Receiver Operating Curve (ROC). When infinite data is available, the Neyman- Pearson (NP) design procedure provides the most e…
Finite resources limit false discovery rate control in structured hypothesis spaces.
We study the problem of distinguishing between two distributions on a metric space; i.e., given metric measure spaces and , we are interested in the problem of determining from finite data whether or not is . The key is to use pairwise distances between observat…
We propose a novel method for clustering data which is grounded in information-theoretic principles and requires no parametric assumptions. Previous attempts to use information theory to define clusters in an assumption-free way are based on maximizing mutual information between data and cluster labels. We demonstrate …
A typical problem in causal modeling is the instability of model structure learning, i.e., small changes in finite data can result in completely different optimal models. The present work introduces a novel causal modeling algorithm for longitudinal data, that is robust for finite samples based on recent advances in st…
Users of a personalised recommendation system face a dilemma: recommendations can be improved by learning from data, but only if the other users are willing to share their private information. Good personalised predictions are vitally important in precision medicine, but genomic information on which the predictions are…
The development of a metric for structural data is a long-term problem in pattern recognition and machine learning. In this paper, we develop a general metric for comparing nonlinear dynamical systems that is defined with Perron-Frobenius operators in reproducing kernel Hilbert spaces. Our metric includes the existing …
While adversarial training can improve robust accuracy (against an adversary), it sometimes hurts standard accuracy (when there is no adversary). Previous work has studied this tradeoff between standard and robust accuracy, but only in the setting where no predictor performs well on both objectives in the infinite data…
Optimal ridge regularization computed iteratively from generative parameters.
We use the language of uninformative Bayesian prior choice to study the selection of appropriately simple effective models. We advocate for the prior which maximizes the mutual information between parameters and predictions, learning as much as possible from limited data. When many parameters are poorly constrained by …
The question of how best to estimate a continuous probability density from finite data is an intriguing open problem at the interface of statistics and physics. Previous work has argued that this problem can be addressed in a natural way using methods from statistical field theory. Here I describe new results that allo…
New method uses Gaussian processes for solving linear PDEs with boundary conditions.
VR methods improve SGD for faster machine learning.
Our goal in this paper is to develop an effective estimator of fractal dimension. We survey existing ideas in dimension estimation, with a focus on the currently popular method of Grassberger and Procaccia for the estimation of correlation dimension. There are two major difficulties in estimation based on this method. …
Study reveals how high-dimensional models are vulnerable to consistent adversarial attacks.
Margin-based classifiers have been popular in both machine learning and statistics for classification problems. Since a large number of classifiers are available, one natural question is which type of classifiers should be used given a particular classification task. We answer this question by investigating the asympto…
Gradient descent converges geometrically to optimal self-attention parameters.
New method reduces over-parametrization in neural networks, ensuring sparsity and finite network size.
FP-BMA improves generalization by encouraging flat posteriors in Bayesian Model Averaging.
New research shows CPE only occurs when Bayesian posterior underfits.
Differential privacy of Gaussian process posterior sampling
Theoretical framework for M-posteriors connects Bayesian and frequentist statistics.
New method improves generative model performance by fully conditioning variational posteriors.
PVI seeks a posterior that makes predictions closer to true data, not approximating the Bayesian posterior.
Bayesian neural networks (BNNs) hold great promise as a flexible and principled solution to deal with uncertainty when learning from finite data. Among approaches to realize probabilistic inference in deep neural networks, variational Bayes (VB) is theoretically grounded, generally applicable, and computationally effic…
Optimized -posteriors reduce KL divergence from true posterior in parametric misspecification.
This work explores how overparametrization and priors affect Bayesian neural network posteriors.
Paper analyzes mistake and generalization of MNIC classifiers.
New priors can update posteriors without re-estimating likelihoods.
Bayesian learning made scalable with posteriors library.
The representation of the approximate posterior is a critical aspect of effective variational autoencoders (VAEs). Poor choices for the approximate posterior have a detrimental impact on the generative performance of VAEs due to the mismatch with the true posterior. We extend the class of posterior models that may be l…