GraphGP: Scalable Gaussian Processes with Vecchia's Approximation
problem Naive Gaussian Process computation limits practical use
method GPU algorithm for Vecchia's approximation
result Linear time and memory requirements for nearly a billion parameters
Bayesian optimization technique scaled using Vecchia approximations.
problem Scalability issue with Gaussian process surrogate models in Bayesian optimization.
method Adapted Vecchia approximation from spatial statistics to Gaussian processes, developed improvements and extensions.
result Methods compared favorably to state-of-the-art on various test functions and reinforcement learning problems.
New iterative methods improve Vecchia-Laplace approximations for large data sets.
problem Inaccurate and slow Vecchia-Laplace approximations for large data sets.
method Iterative methods to improve Vecchia-Laplace approximations, including preconditioners and novel methods for predictive variances.
result Order of magnitude speed-up and threefold increase in prediction accuracy compared to state-of-the-art methods.
Proposes efficient Gaussian process approximations for large datasets.
problem Scalability issues in Gaussian processes for large data sets.
method Combines Vecchia approximations and inducing points methods.
result Efficient and accurate approximations for various data types.
New GPU algorithm speeds up Gaussian Process analysis.
problem Reducing computational complexity for large spatial datasets.
method Implemented three GPU methods for Vecchia Approximation.
result New GPU method outperforms existing methods.
Novel neural GP kernels learn stable, flexible covariance structures.
problem Scalable and flexible covariance kernels for Gaussian processes.
method Directly learn kriging coefficients and conditional standard deviations using deep neural architectures exploiting permutation-equivariant structure.
result Improved training stability and data efficiency with expressive, non-stationary kernels.
A scalable algorithm for GP regression selects relevant covariates efficiently.
problem Scalable variable selection in large GP regression models.
method VGPR algorithm using Vecchia approximation for sparse precision matrix, mini-batch subsampling.
result Improved scalability and accuracy in selecting relevant covariates.
Vecchia approximations provide the best accuracy-runtime trade-off for Gaussian process approximations.
problem High computational cost of Gaussian processes for large data sets.
method Systematic comparison of different Gaussian process approximations.
result Vecchia approximations consistently provide the best accuracy-runtime trade-off.
DVE uses GPs on DNN outputs to provide UQ without retraining.
problem Feature collapse in DNNs affects UQ methods.
method Deep Vecchia ensemble (DVE) of GPs on DNN hidden layers.
result Deterministic UQ possible in feature-collapsed DNNs.
New method speeds up analysis of computer experiments.
problem Computational infeasibility of direct GP inference for large datasets.
method Adapted Vecchia's ordered conditional approximation to scaled input space.
result Significant performance improvement over existing methods.
We derive a single pass algorithm for computing the gradient and Fisher information of Vecchia's Gaussian process loglikelihood approximation, which provides a computationally efficient means for applying the Fisher scoring algorithm for maximizing the loglikelihood. The advantages of the optimization techniques are de…
New iterative methods improve scalability of Gaussian process approximations for large data.
problem Scalability issues in Gaussian process approximations for large spatial data.
method Iterative methods combined with preconditioners to reduce computational costs.
result Preconditioners accelerate convergence and improve predictive variances.
A scalable Gaussian process clustering method for large datasets.
problem Infeasibility of Gaussian process clustering on large grids.
method Embedding Vecchia approximation in EM algorithm for scalability.
result Efficient Gaussian process clustering for large environmental applications.
Combines boosting with Gaussian process and mixed effects models.
problem Model misspecifications and independence assumptions in boosting.
method Relaxes zero or linearity assumption in Gaussian process and mixed effects models, and independence assumption in boosting.
result Increased prediction accuracy compared to existing approaches.
Improved spatial distribution learning with Bayesian transport maps and parametric shrinkage.
problem Learning non-Gaussian spatial distributions with limited training data.
method Proposed ShrinkTM approach using Bayesian transport maps with parametric shrinkage.
result ShrinkTM outperforms existing BTM, especially with few training samples.
This paper develops a fast algorithm for solving nonlinear PDEs using sparse Cholesky factorization.
problem Efficiently solving nonlinear PDEs with Gaussian processes and kernel methods.
method Sparse Cholesky factorization for near-linear complexity.
result Near-linear complexity algorithm for working with kernel matrices of nonlinear PDEs.
TERA method speeds up derivative Gaussian processes in high dimensions.
problem High-dimensional function evaluations and gradient computations are computationally expensive.
method TERA uses exact gradient reduction to decouple n and d from the computational cost. result TERA achieves state-of-the-art predictive accuracy with orders of magnitude faster computation.
New method preserves GCM spatial dependencies for better climate projections.
problem Systemic biases in GCM output and loss of spatial/temporal dependencies.
method SPECD approach using Vecchia approximation and semi-parametric quantile regression.
result SPECD preserves key marginal and joint distribution properties of precipitation and temperature.
In the current literature, the analytical tractability of discrete time option pricing models is guaranteed only for rather specific types of models and pricing kernels. We propose a very general and fully analytical option pricing framework, encompassing a wide class of discrete time models featuring multiple-componen…
Spatial statisticians and quantitative investors use the same mathematical object: a Schur complement, damped by one parameter.
problem The Schur complement is used in both spatial modeling and portfolio allocation, but the parameters are different.
method The Schur complement is interpreted as reliability shrinkage of a conditional Gaussian.
result The Schur complement is the same in both applications.
The study provides conditions for approximating Riemannian manifolds with polyhedral metrics.
problem Approximating Riemannian manifolds with polyhedral metrics.
method Conditions on curvature tensors for Lipschitz and local polyhedral approximations.
result Conditions are sufficient for local polyhedral approximations, conjectured to be sufficient for global approximations.
Geometric Gaussian approximations capture any distribution.
problem Approximating complex probability distributions.
method Geometric Gaussian approximations through diffeomorphisms or exponential maps.
result Geometric Gaussian approximations are universal, capturing any distribution.
Method approximates Riemannian barycenter on manifolds.
problem Computing the exact Riemannian barycenter is computationally expensive.
method Uses under- and over-approximations of Riemannian distance to compute an approximate barycenter.
result Approximation method is more efficient than exact methods and steepest descent.
We consider in this paper the optimal approximations of convex univariate functions with feed-forward Relu neural networks. We are interested in the following question: what is the minimal approximation error given the number of approximating linear pieces? We establish the necessary and sufficient conditions and uniqu…
Efficiently reduces tensor ranks using mean-field approximation.
problem Low-rank approximation of non-negative tensors.
method Mean-field approximation of tensor rank reduction.
result Our algorithm achieves faster and competitive tensor rank reduction.
Study approximates unknown function levels with queries.
problem Approximating unknown function levels through sequential queries.
method Introduce Bisect and Approximate algorithms to reduce to local function approximation.
result Rate-optimal sample complexity guarantees for H{ö}lder functions.
We study sparse approximate solutions to convex optimization problems. It is known that in many engineering applications researchers are interested in an approximate solution of an optimization problem as a linear combination of elements from a given system of elements. There is an increasing interest in building such …
Softmax attention approximates complex functions and subsumes many known universal approximators.
problem Universal approximation of continuous sequence-to-sequence functions.
method Interpolation-based analysis of attention's internal mechanism, showing its ability to approximate ReLU functions.
result Softmax attention is a universal approximator for continuous sequence-to-sequence functions.
Improved matrix approximation using randomized algorithms.
problem Finding better approximations of given matrices.
method Randomized algorithms to compute (HT) as an improved approximation. result Computed (HT) provides a better approximation than given F∗. Deviation inequalities for stochastic approximation methods.
problem Establishing bounds on the deviation of stochastic approximation methods.
method Martingale approximation method for separately Lipschitz functions.
result Established various deviation inequalities for stochastic approximation by averaging and minimization.
Neural approximate computing gains enormous energy-efficiency at the cost of tolerable quality-loss. A neural approximator can map the input data to output while a classifier determines whether the input data are safe to approximate with quality guarantee. However, existing works cannot maximize the invocation of the a…
Approximate symmetries of geodesic equations on 2-spheres are studied. These are the symmetries of the perturbed geodesic equations which represent approximate path of a particle rather than exact path. After giving the exact symmetries of the geodesic equations, two different approaches to study the approximate symmet…
We are concerned with an approximation problem for a symmetric positive semidefinite matrix due to motivation from a class of nonlinear machine learning methods. We discuss an approximation approach that we call {matrix ridge approximation}. In particular, we define the matrix ridge approximation as an incomplete matri…
Transformers use ReLUs to approximate softmax efficiently.
problem Analyzing resource usage in softmax transformer models.
method Translating ReLU approximation results to softmax attention mechanisms.
result Economic resource bounds for softmax attention mechanisms.
Approximating complex curves with simple parametric curves is widely used in CAGD, CG, and CNC. This paper presents an algorithm to compute a certified approximation to a given parametric space curve with cubic B-spline curves. By certified, we mean that the approximation can approximate the given curve to any given pr…
Recently, variational approximations such as the mean field approximation have received much interest. We extend the standard mean field method by using an approximating distribution that factorises into cluster potentials. This includes undirected graphs, directed acyclic graphs and junction trees. We derive generaliz…
Adaptive approximations improve variational inference for complex models.
problem Efficiently approximate marginal distributions and partition functions in complex probabilistic models.
method Two classes of adaptive approximations that include Bethe, tree-reweighted, and convex free energies.
result Proposed approximations automatically adapt to a given model and outperform existing methods.
Non-negative L1-approximating polynomials for Gaussian distributions are proven for certain classes of sets.
problem Existence of non-negative L1-approximating polynomials for Gaussian distributions. method Proving the existence of degree-k non-negative polynomials that approximate indicator functions of sets with Gaussian surface area in L1-norm. result Proves the existence of non-negative L1-approximating polynomials for certain classes of sets with Gaussian surface area. Paper introduces new approximations for lognormal sums, matching comonotonicity and moments.
problem Approximating sums of lognormal random variables accurately.
method Introduces new approximations based on weighted distribution theory, emphasizing comonotonicity and moment matching.
result Approximations perform better than classical methods, especially in the right tail of the distribution.
Paper analyzes normal approximation for two-timescale stochastic algorithms, revealing interaction between fast and slow timescales.
problem Non-asymptotic bounds for accuracy of normal approximation in linear two-timescale stochastic approximation algorithms.
method Established bounds for normal approximation in terms of convex distance, focusing on last iterate and Polyak-Ruppert averaging.
result Normal approximation rate for the last iterate improves with increased timescale separation, while it decreases in the averaged setting.
One-pass algorithm finds small subset for ℓp subspace approximation with additive error.
problem Finding a small subset of data points for ℓp subspace approximation. method One-pass subset selection with additive approximation guarantee for p∈[1,∞). result First one-pass algorithm with additive error for ℓp subspace approximation. We are interested in approximation of a multivariate function f(x1,…,xd) by linear combinations of products u1(x1)⋯ud(xd) of univariate functions ui(xi), i=1,…,d. In the case d=2 it is a classical problem of bilinear approximation. In the case of approximation in the L2 space the bili…
A new method for efficient Gaussian process inference using sparse approximations.
problem Scalable and accurate inference for latent Gaussian processes.
method Variational approximation with sparse inverse Cholesky factors and double Kullback-Leibler minimization.
result The proposed method can achieve highly accurate approximations with polylogarithmic time complexity.
In this paper, we propose a low-rank approximation method based on discrete least-squares for the approximation of a multivariate function from random, noisy-free observations. Sparsity inducing regularization techniques are used within classical algorithms for low-rank approximation in order to exploit the possible sp…
Neural network based approximate computing is a universal architecture promising to gain tremendous energy-efficiency for many error resilient applications. To guarantee the approximation quality, existing works deploy two neural networks (NNs), e.g., an approximator and a predictor. The approximator provides the appro…
The paper approximates supply curves using a one-step basis method.
problem Computing supply curves accurately and efficiently.
method Derives L2 approximation expression and proposes node selection procedure.
result Illustrates the approach with European electricity market bid curves.
The paper defines a new concept of approximability for Lagrangian submanifolds.
problem Understanding the approximability of Lagrangian submanifolds.
method Introducing a new notion of categorical approximability for metric spaces, showing it applies to specific types of Lagrangian submanifolds.
result Examples of Lagrangian submanifolds are found that are approximable but not precompact.
Boosting Nyström improves accuracy of matrix approximations.
problem Generating low-rank approximations of large matrices efficiently.
method Iteratively generate multiple weak Nyström approximations, combine them to form a strong approximation.
result Boosting Nyström yields more efficient and accurate low-rank approximations.