A new method of moments estimator goes beyond data reweighting.
problem Estimation of moment restrictions and conditional moment restrictions.
method Kernel Method of Moments (KMM) based on maximum mean discrepancy.
result KMM achieves competitive performance on conditional moment restriction tasks.
A new method for estimating causal parameters from observables reduces the need for finite moment conditions.
problem Estimating causal parameters from observational data with unknown or infinite moment conditions.
method Variational Method of Moments (VMM) for a general class of estimators, including kernel and neural net-based methods.
result VMM estimators are consistent, asymptotically normal, and semiparametrically efficient.
Paper provides unbiased spectral moment estimates from finite data.
problem Challenges in estimating spectral moments from limited data.
method Dynamic programming approach to estimate spectral moments of kernel integral operator.
result Demonstrates consistency with theoretical spectra and practical utility in neural networks.
We propose a new family of specification tests called kernel conditional moment (KCM) tests. Our tests are built on a novel representation of conditional moment restrictions in a reproducing kernel Hilbert space (RKHS) called conditional moment embedding (CMME). After transforming the conditional moment restrictions in…
A method learns representations for conditional moment models with controlled ill-posedness.
problem Efficient estimation of nonparametric conditional moment models with flexible models is challenging.
method Proposes a procedure that learns spectral representations with controlled measures of ill-posedness.
result The proposed method can efficiently estimate representations from data and is L2 consistent.
New KSDs control moments in approximations, improving diagnostics and tests.
problem Inability of standard KSDs to control moment convergence.
method Developed alternative diffusion KSDs under sufficient conditions.
result First KSDs to exactly characterize q-Wasserstein convergence.
Graph spectral techniques for measuring graph similarity, or for learning the cluster number, require kernel smoothing. The choice of kernel function and bandwidth are typically chosen in an ad-hoc manner and heavily affect the resulting output. We prove that kernel smoothing biases the moments of the spectral density.…
Kernel methods estimate causal effects with a single proxy for deterministic confounders.
problem Estimating causal effects with a single proxy for an unobserved confounder.
method Two kernel-based methods: two-stage regression and maximum moment restriction.
result Both kernel methods can consistently estimate the causal effect.
Graph spectra have been successfully used to classify network types, compute the similarity between graphs, and determine the number of communities in a network. For large graphs, where an eigen-decomposition is infeasible, iterative moment matched approximations to the spectra and kernel smoothing are typically used. …
New method improves estimation of complex models from conditional moment restrictions.
problem Estimation of complex models from conditional moment restrictions.
method Functional Generalized Empirical Likelihood (GEL) with a practical method.
result The method achieves state-of-the-art performance on two problems.
Kernel DRO uses RKHS to optimize under distributional uncertainty.
problem Optimizing under distributional uncertainty with limited knowledge.
method Kernel DRO using RKHS ambiguity sets and duality theory.
result Unified approach to robust and stochastic optimization.
A new method extracts features and reconstructs moments in dynamical systems using information geometry.
problem Reconstructing moments in dynamical systems efficiently and accurately.
method Information-geometric approach on spaces of probability measures.
result Moments can be expanded in eigenfunctions of a kernel integral operator, enabling nonparametric forecasting.
We present a novel framework for kernel learning with sequential data of any kind, such as time series, sequences of graphs, or strings. Our approach is based on signature features which can be seen as an ordered variant of sample (cross-)moments; it allows to obtain a "sequentialized" version of any static kernel. The…
The sequence of moments of a vector-valued random variable can characterize its law. We study the analogous problem for path-valued random variables, that is stochastic processes, by using so-called robust signature moments. This allows us to derive a metric of maximum mean discrepancy type for laws of stochastic proce…
New method approximates MMD using pseudo-differential operators and singular values.
problem Approximating MMD with pseudo-differential operators and singular values.
method Corresponding pseudo-differential operators to Mercer kernels, approximating p(x,y) with its first r singular values. result The new MMD distance measures the difference of two distributions with respect to r∗ local moments, where r∗ depends on singular values decay rate. We provide an approach for learning deep neural net representations of models described via conditional moment restrictions. Conditional moment restrictions are widely used, as they are the language by which social scientists describe the assumptions they make to enable causal inference. We formulate the problem of est…
New tests for distributional causal effects using improved kernel estimators.
problem Testing for higher-order moments and multidimensional outcomes affected by treatment.
method Improved kernel estimators based on doubly robust mean embeddings.
result New permutation-based tests for distributional causal effects with improved convergence rates.
Kernel tests assess equivalence between distributions without assuming specific moments.
problem Traditional goodness-of-fit tests fail to detect meaningful distributional differences.
method Proposes kernel-based tests using kernel Stein discrepancy and Maximum Mean Discrepancy.
result Tests assess the absence of meaningful distributional differences under controlled error rates.
Study on kernel tests for high-dimensional data, focusing on MMD and CLT.
problem Asymptotic behavior of kernel two-sample tests in high dimensions and large samples.
method Maximum mean discrepancy (MMD) with isotropic kernels, deriving asymptotic expansions and CLT.
result Interplay between moment discrepancy and dimension-and-sample orders in kernel tests.
Sharp policy value estimation for contextual bandits with unobserved confounders.
problem Estimating policy value under unobserved confounders with sensitivity analysis.
method Kernel method to approximate conditional moment constraints, leveraging f-divergence.
result Sharp lower bound of policy value, avoiding coarse relaxation of uncertainty set.
Develops a new method for estimating models with conditional moment restrictions.
problem Estimating models with conditional moment restrictions, especially non-parametric instrumental variable regression.
method Introduces a min-max criterion function to solve a zero-sum game between modeler and adversary, analyzing estimation rates for various hypothesis spaces.
result Shows that with regularization and rich test function spaces, estimation rates scale with the critical radius of hypothesis and test function spaces.
New kernels from neural networks show better performance than traditional methods.
problem Improving neural network performance on small datasets.
method Developed algebraic operations to create compositional kernels from neural network architectures.
result Compositional kernels achieve higher accuracy than neural tangent kernels and neural networks on small datasets.
We provide bounds for kernel matrices and new approximations for high-dimensional data.
problem Approximating high-dimensional empirical kernel matrices.
method Decoupling results for U-statistics and non-commutative Khintchine inequality.
result New tighter approximations for inner-product kernel matrices.
New method uses machine learning to improve statistical inference.
problem Performing inference on conditional functionals with scarce labeled data.
method Combines localization with prediction-based variance reduction.
result Valid and sharp confidence intervals for conditional functionals.
A new method for analyzing adaptive experiments using kernel treatment effects.
problem Efficiently analyzing adaptive experiments that adjust treatment assignments based on outcomes.
method Kernel Treatment Effects (KTE) framework combining RKHS scores and witness functions.
result Effective for both mean shifts and higher-moment differences, outperforming adaptive baselines.
Study heavy-tailed weights' impact on neural network's spectral distribution.
problem Analyzing spectral distribution of conjugate kernel matrices with heavy-tailed weights.
method Computed limiting eigenvalue distribution through moments, considering heavy-tailed distributions and nonlinear activation functions.
result Heavy-tailed weights induce strong correlations, leading to fundamentally different spectral behavior.
New GP-based method improves uncertainty quantification for causal functions.
problem Challenges in quantifying uncertainty for causal effects, especially for entire functions.
method GP-based approach using inner-product of observational functions in RKHS, with tractable posterior moments and calibration.
result Improves uncertainty quantification while maintaining causal effect estimation performance.
Deep Gaussian Processes (DGPs) combine the expressiveness of Deep Neural Networks (DNNs) with quantified uncertainty of Gaussian Processes (GPs). Expressive power and intractable inference both result from the non-Gaussian distribution over composition functions. We propose interpretable DGP based on approximating DGP …
Develops a deterministic method to approximate NSDEs for better uncertainty quantification.
problem Computational infeasibility of obtaining well-calibrated uncertainty from NSDEs.
method Bidimensional moment matching algorithm for approximating NSDE transition kernel.
result Deterministic approximation improves uncertainty calibration and prediction accuracy.
New findings on kernel regression in the quadratic regime, improving understanding of machine learning models.
problem Understanding kernel ridge regression in the quadratic asymptotic regime.
method Extended study of kernel regression to the quadratic regime, establishing approximation bounds and spectral distributions.
result Broad class of inner-product kernels exhibit behavior similar to a quadratic kernel, with precise asymptotic training and test errors characterized.
Paper develops efficient recursive learning for multi-channel systems with heterogeneous dynamics.
problem Accurately learning system dynamics in complex, multi-channel systems with nonlinear and noisy data.
method Formulates system as Gaussian process state-space models (GPSSMs), introduces heterogeneous multi-output kernel, and develops recursive inference framework.
result Matches SOTA offline GPSSMs in accuracy with 1/100 runtime, and outperforms SOTA online GPSSMs by 70% in accuracy under noise with 1/20 runtime.
New method detects changes in high-dimensional data from small samples.
problem Detecting changes in high-dimensional data with limited samples.
method Angular kernel scan framework for detecting marginal distributional shifts.
result Exact population mean factorization and asymptotically distribution-free test.
This paper introduces a new framework for quantifying predictive uncertainty for both data and models that relies on projecting the data into a Gaussian reproducing kernel Hilbert space (RKHS) and transforming the data probability density function (PDF) in a way that quantifies the flow of its gradient as a topological…
Estimates KRR risk from training data for various kernels and hyperparameters.
problem Predicting the generalization error of Kernel Ridge Regression.
method Introduces SCT and KARE to approximate KRR risk from training data.
result KARE provides an excellent approximation of KRR risk and helps select good kernels.
Kernel discriminant analysis uses nonlinear embeddings to improve classification.
problem Limited effectiveness of linear discriminant analysis in capturing nonlinear features.
method Study of nonlinear embeddings in kernel discriminant analysis using polynomial and Gaussian kernels, solving generalized eigenvalue problems.
result Polynomial and Gaussian discriminants capture class differences through population moments and randomized projections.
Study on U-statistics with heavy-tailed samples, providing tail bounds and LDP.
problem Deviation of U-statistics with heavy-tailed samples.
method Exponential tail bounds and Large Deviation Principle (LDP) for U-statistics.
result Obtained an exponential upper bound for U-statistics tail decay, showing two regions of decay.
Sharp heat kernel estimates on manifolds lead to solutions of the Parabolic Anderson model.
problem Well-posedness and intermittency of solutions to the Parabolic Anderson model on Riemannian manifolds.
method Sharp global heat kernel bounds and geodesic comparison geometry.
result Upper and lower moment bounds for solutions of the Parabolic Anderson model on general compact Riemannian manifolds.
Researchers develop a generalised geometric Brownian motion for better asset pricing.
problem Irregularities in simple geometric Brownian motion for asset dynamics.
method Introduce a memory kernel to generalise GBM, derive moments and probability density functions.
result The performance of kernels in pricing options depends on option maturity and moneyness.
The learning of domain-invariant representations in the context of domain adaptation with neural networks is considered. We propose a new regularization method that minimizes the discrepancy between domain-specific latent feature representations directly in the hidden activation space. Although some standard distributi…
Geometric regularisation improves statistical models by avoiding degeneracy loci.
problem Non-identifiability, singular information, and moment indeterminacy in statistical models.
method Develops the geometric regularisation of distribution-kernel pairs (T,φ) using Whitney, Thom, and Mather theorems. result Finite-dimensional weak transversality theorem for generic kernels, avoiding degeneracy strata of high codimension.
We design a new nonparametric method that allows one to estimate the matrix of integrated kernels of a multivariate Hawkes process. This matrix not only encodes the mutual influences of each nodes of the process, but also disentangles the causality relationships between them. Our approach is the first that leads to an …
Kernelized cumulants improve statistical analysis in high-dimensional spaces.
problem Statistical analysis in high-dimensional spaces with low variance estimators.
method Extending cumulants to RKHS using tensor algebra and kernel trick.
result Kernelized cumulants provide new all-purpose statistics with computational tractability.
Criterion extends identifiability for continuous mixtures of kernels.
problem Identify continuous mixtures of kernels.
method Generating-function accessibility criterion based on moment-generating functions or Laplace transforms.
result Criterion applies to mixtures of discrete and continuous variables.
The paper studies Lipschitz bounds for integral kernels under differentiability assumptions.
problem Understanding the Lipschitz continuity of feature maps associated with integral kernels.
method Analyzes differentiability assumptions to derive explicit formulas for Lipschitz constants and conditions for non-Lipschitz continuity.
result Explicit formulas and conditions for Lipschitz continuity of feature maps associated with various kernels.
Paper shows robustness of kernel-based pairwise learning without strict assumptions.
problem Statistical robustness of kernel-based pairwise learning under minimal conditions.
method No assumptions on input and output spaces; derives influence function and robustness.
result Qualitative robustness of kernel-based estimator established.
Researchers use Gaussian processes to approximate Lagrange multipliers for Maximum-Entropy distributions.
problem Finding Lagrange multipliers for Maximum-Entropy distributions is computationally challenging.
method Employed Gaussian processes to approximate the Lagrange multipliers as a map of moments. Optimized hyperparameters by maximizing log-likelihood.
result Data-driven Maximum-Entropy closure performs well in approximating non-equilibrium distributions.
Many pattern recognition methods rely on statistical information from centered data, with the eigenanalysis of an empirical central moment, such as the covariance matrix in principal component analysis (PCA), as well as partial least squares regression, canonical-correlation analysis and Fisher discriminant analysis. R…
The study analyzes prediction errors in systems with memory kernels, providing bounds and stability results.
problem Prediction errors in stochastic dynamical systems with memory kernels.
method Analysis of generalized Langevin equations (GLEs) with Volterra equations, integrating synchronized noise coupling and weighted norms.
result Prediction discrepancies decay at a rate determined by the memory kernel's decay, quantitatively bounded by kernel estimation errors.