Develops a new theory for approximating functions on massive data.
problem Challenges in machine learning with massive data.
method eignets theory for local, stratified approximation.
result Solves inverse problems like finding data probability law and function smoothness.
We discuss Bayesian methods for learning Bayesian networks when data sets are incomplete. In particular, we examine asymptotic approximations for the marginal likelihood of incomplete data given a Bayesian network. We consider the Laplace approximation and the less accurate but more efficient BIC/MDL approximation. We …
New method improves Gaussian kernel approximations for high-frequency data.
problem Limited scalability of kernel-based models to large data sets.
method Local random feature approximations using Maclaurin expansions and polynomial sketches.
result Significant improvement in kernel approximations and downstream performance for high-frequency data.
Improves Laplace approximation for Bayesian inference on Riemannian manifolds.
problem Inaccurate Gaussian approximations for complex targets and finite-data posteriors.
method Develops alternative variants of the Laplace approximation using a Riemannian metric.
result Exact approximations at the limit of infinite data, improving practical performance.
We develop methods to efficiently approximate data in metric spaces without additional assumptions.
problem Efficiently approximating data in metric spaces without imposing structural assumptions.
method Identify discrete modulus of continuity, investigate consistency, propose algorithm, and develop approximation theory.
result Consistent approximation of data in metric spaces without structural assumptions.
Vecchia approximations provide the best accuracy-runtime trade-off for Gaussian process approximations.
problem High computational cost of Gaussian processes for large data sets.
method Systematic comparison of different Gaussian process approximations.
result Vecchia approximations consistently provide the best accuracy-runtime trade-off.
Paper applies ANOVA decomposition for interpretable data approximation.
problem High-dimensional data interpretation and dimensionality reduction.
method ANOVA decomposition and Grouped Transformations for interpretability.
result Ability to rank variable interactions and unimportant variables.
Paper improves tensor approximation for streaming data.
problem Challenges in finding accurate low-tubal-rank tensor approximations in streaming settings.
method Extends Frequent Directions for efficient low-tubal-rank tensor approximation.
result The new algorithm achieves arbitrarily small approximation error with linear sketch size growth.
Fast approximate inference for non-Gaussian data.
problem Efficient inference for non-Gaussian data.
method Laplace Matching for fast approximate inference in latent Gaussian models.
result Achieves high approximation quality with low computational cost.
New iterative methods improve Vecchia-Laplace approximations for large data sets.
problem Inaccurate and slow Vecchia-Laplace approximations for large data sets.
method Iterative methods to improve Vecchia-Laplace approximations, including preconditioners and novel methods for predictive variances.
result Order of magnitude speed-up and threefold increase in prediction accuracy compared to state-of-the-art methods.
Optimal smooth subspaces approximate large data sets efficiently.
problem Approximating large data sets with invariant subspaces.
method Smooth functions under lattice translations or crystallographic groups, with optimal selection of Paley-Wiener space.
result Optimal lattice selection enhances approximation efficiency.
This paper analyzes and improves GANs' approximation ability.
problem Theoretical and algorithmic analysis of GANs' approximation property.
method Theoretical analysis and SDG approach to enhance GANs' approximation ability.
result The generator of GANs can universally approximate the potential data distribution.
In order to avoid the curse of dimensionality, frequently encountered in Big Data analysis, there was a vast development in the field of linear and nonlinear dimension reduction techniques in recent years. These techniques (sometimes referred to as manifold learning) assume that the scattered input data is lying on a l…
We present effective numerical algorithms for locally recovering unknown governing differential equations from measurement data. We employ a set of standard basis functions, e.g., polynomials, to approximate the governing equation with high accuracy. Upon recasting the problem into a function approximation problem, we …
New method approximates CVaR with less data for heavy-tailed risks.
problem Lack of data for accurate CVaR approximation in heavy-tailed distributions.
method Importance sampling based extrapolation for heavy-tailed distributions.
result Statistically consistent approximations with reduced data requirements.
The paper projects unknown manifolds onto hyperspheres for efficient function approximation.
problem Function approximation from data on unknown manifolds with added errors.
method Projects unknown manifold onto hypersphere and uses localized spherical polynomial kernels.
result Optimal rates of approximation for rough functions are given.
The paper studies randomized approximations of Tukey's depth for log-concave isotropic data.
problem The challenge of approximating Tukey's depth in high dimensions.
method The study examines randomized algorithms for approximating Tukey's depth for log-concave isotropic data.
result Randomized algorithms correctly approximate maximal depth and close to zero depths but not intermediate depths.
Better Hessian approximations improve influence function attributions in deep learning.
problem Influence functions are difficult to compute due to ill-conditioned Hessians, leading to poor data attribution performance.
method Investigated the impact of Hessian approximation quality on influence-function attributions in a controlled setting.
result Better Hessian approximations consistently yield better influence score quality.
There are many methods developed to approximate a cloud of vectors embedded in high-dimensional space by simpler objects: starting from principal points and linear manifolds to self-organizing maps, neural gas, elastic maps, various types of principal curves and principal trees, and so on. For each type of approximator…
We provide bounds for kernel matrices and new approximations for high-dimensional data.
problem Approximating high-dimensional empirical kernel matrices.
method Decoupling results for U-statistics and non-commutative Khintchine inequality.
result New tighter approximations for inner-product kernel matrices.
Paper proposes a fast method for approximate data deletion in generative models.
problem Efficient data deletion in unsupervised learning models is an open problem.
method Density-ratio-based framework for generative models, fast method for approximate data deletion, statistical test.
result Theoretical guarantees and empirical demonstrations of the proposed methods across various generative models.
Paper improves kernel approximations for better statistical learning.
problem Improving kernel approximations for better statistical learning.
method Taylor series approximations of radial kernel functions.
result Establishes upper bounds for eigenfunctions, leading to better approximations.
Discrete approximation solves Björling's minimal surface problem.
problem Constructing minimal surfaces from real-analytic curves with specified normal fields.
method Approximate solution by discrete minimal surfaces and discrete isothermic surfaces.
result Approximation error is proportional to the square of the mesh size.
In much of the literature on function approximation by deep networks, the function is assumed to be defined on some known domain, such as a cube or a sphere. In practice, the data might not be dense on these domains, and therefore, the approximation theory results are observed to be too conservative. In manifold learni…
The vast quantity of information brought by big data as well as the evolving computer hardware encourages success stories in the machine learning community. In the meanwhile, it poses challenges for the Gaussian process (GP) regression, a well-known non-parametric and interpretable Bayesian model, which suffers from cu…
ASTRA improves TDA by more accurately approximating iHVP.
problem Improving insights into training data attribution.
method ASTRA uses EKFAC-preconditioner on Neumann series iterations to accurately approximate iHVP.
result Improving iHVP approximation significantly improves TDA performance.
Paper proves GDL models can approximate any continuous function on non-Euclidean data.
problem Processing non-Euclidean data with universal feedforward models.
method Introduces geometric deep learning framework for differentiable manifold geometries.
result GDL models can uniformly approximate any continuous function on compact sets.
The paper examines how kernel approximations affect Gaussian process regression in large data applications.
problem Effect of kernel approximations on Gaussian process regression in large data applications.
method Unified framework to analyze Gaussian process regression under computational and epistemic misspecification.
result Theoretical analysis of Gaussian process regression under various misspecifications.
Investigates the impact of finite VC dimension on neural network approximation and learning.
problem The influence of VC dimension on neural network approximation and learning from samples.
method Analysis of high-dimensional geometry and statistical learning theory, focusing on VC dimension.
result Finite VC dimension is beneficial for uniform convergence of empirical errors but not for approximation of functions from a probability distribution.
We consider the problem of efficiently approximating and encoding high-dimensional data sampled from a probability distribution ρ in RD, that is nearly supported on a d-dimensional set M - for example supported on a d-dimensional Riemannian manifold. Geometric Multi-Resolution Analysis (GM…
Neural approximate computing gains enormous energy-efficiency at the cost of tolerable quality-loss. A neural approximator can map the input data to output while a classifier determines whether the input data are safe to approximate with quality guarantee. However, existing works cannot maximize the invocation of the a…
Paper develops federated GLMM algorithms for analyzing hierarchical data.
problem Analyzing hierarchical data with non-independent observations in a federated setting.
method Developed two federated GLMM algorithms using Laplace and Gaussian Hermite approximations.
result Federated GLMM can handle hierarchical data and achieve comparable or superior performance.
This paper speeds up Gaussian process regression for autocorrelated data.
problem Temporal overfitting in Gaussian process models for autocorrelated data.
method Modifying existing Gaussian process approximations to handle blocked, de-correlated data.
result Proposed methods accelerate Gaussian process regression on autocorrelated data without sacrificing performance.
New ACV method speeds up CV in high dimensions with approximate low-rank data.
problem Accurate model assessment in high-dimensional, large data settings with expensive algorithms.
method Developed a new ACV algorithm that uses low-rank approximations of the Hessian matrix.
result The new method is fast and accurate in the presence of approximate low-rank data.
New findings on how convolutional architectures approximate time series data.
problem Understanding the approximation properties of convolutional architectures in time series modeling.
method Mathematical analysis of convolutional architectures applied to time series modeling.
result A new definition of spectrum-based regularity for measuring temporal relationships under convolutional approximation.
This paper introduces a spline-based method for nonparametric ADVI that handles complex posterior distributions.
problem Learning complex posterior distributions with skewness, multimodality, and bounded support.
method Develops a spline-based nonparametric approximation approach for ADVI.
result Establishes the asymptotic consistency of the derived lower bound for importance weighted autoencoder.
A new data-oblivious sketch for logistic regression reduces data size while maintaining approximation accuracy.
problem Efficiently solving logistic regression in one pass over a data stream.
method Data-oblivious sketching approach that reduces data size to poly(μdlog n) weighted points.
result Sketching reduces data size significantly and provides approximation guarantees.
A fast, approximate method for variable selection in GLMs tackles correlated data.
problem Variable selection in generalized linear models with correlated data.
method Replica method of statistical mechanics and vector approximate message passing.
result The proposed algorithm provides fast convergence and high approximation accuracy.
Combines pseudo-point and state space approximations for scalable GPs.
problem Handling large numbers of off-the-grid spatial data-points and long time-series.
method Combines pseudo-point approximations for spatial data with state space GP approximations for temporal data.
result Combined approach is more scalable and applicable to a greater range of spatio-temporal problems.
Collaborative filtering (CF) is a popular technique in today's recommender systems, and matrix approximation-based CF methods have achieved great success in both rating prediction and top-N recommendation tasks. However, real-world user-item rating matrices are typically sparse, incomplete and noisy, which introduce ch…
Paper offers a simple CDS approximation formula with high accuracy.
problem Lack of CDS levels for market appreciation of companies' default risk.
method Developed a global and transparent Equity-to-Credit (E2C) formula using random forest regression.
result Random forest regression with E2C formula achieves 87.3% out-of-sample accuracy in CDS approximations.
Study on approximability and generalization in machine learning.
problem Understanding how approximation affects learning and generalization in machine learning.
method Introducing a notion of sensitivity to analyze the impact of approximation operators on predictors and proving upper bounds on generalization.
result Proven that approximable target concepts are learnable with fewer labelled samples and sufficient unlabelled data.
This paper introduces Bayes Hilbert spaces for efficient posterior approximation.
problem Efficient posterior approximation in Bayesian models for large datasets.
method Develops Bayes Hilbert spaces for posterior approximation and connects them to Bayesian coresets and kernel-based distances.
result Bayes Hilbert spaces provide a novel framework for posterior approximation that is computationally efficient.
For optimization on large-scale data, exactly calculating its solution may be computationally difficulty because of the large size of the data. In this paper we consider subsampled optimization for fast approximating the exact solution. In this approach, one gets a surrogate dataset by sampling from the full data, and …
Paper develops a new kernel approximation framework.
problem High time and space complexity of kernel methods for large datasets.
method Perturbation-based kernel approximation framework using classical perturbation theory.
result Framework generalizes and improves upon existing methods.
Classifiers label data as belonging to one of a set of groups based on input features. It is challenging to obtain accurate classification performance when the feature distributions in the different classes are complex, with nonlinear, overlapping and intersecting supports. This is particularly true when training data …
The study explores various localized bases and their duals for scattered data approximation.
problem Scattered data approximation using radial basis functions.
method Examines different localized bases including Lagrange, Newton, and multiresolution versions, and their duals.
result Localized orthogonal bases, such as the Newton basis, offer symmetric preconditioners and are feasible for scattered data approximation.
Method approximates high-dimensional feature vectors for supervised learning.
problem Reducing storage and computational complexity in high-dimensional data.
method Approximates feature vectors using a kernel-based approach to reduce data size.
result Significant improvements in classification and regression tasks.