Discovering the causal structure among a set of variables is a fundamental problem in many areas of science. In this paper, we propose Kernel Conditional Deviance for Causal Inference (KCDC) a fully nonparametric causal discovery method based on purely observational data. From a novel interpretation of the notion of as…
New method detects latent common causes from observational data.
problem Detecting latent common causes in observational data.
method Modified causal discovery algorithms to detect latent common causes.
result Successfully detects latent common causes in various noise regimes and real data.
New method uses kernel deviance measures to discover causal relationships in heterogeneous data.
problem Discovering causal relationships in complex, heterogeneous datasets.
method KIIM-HT, a novel score measure based on heterogeneous transformations of RKHS embeddings.
result KIIM-HT outperforms previous methods in causal discovery tasks.
Deviance-style normalization for sparse, jointly overdispersed count matrices
problem Jointly overdispersed count matrices
method Dirichlet-multinomial deviance residualization
result Preserves exact sparsity, evaluates in constant time, recovers multinomial residual
Reasoning based on causality, instead of association has been considered as a key ingredient towards real machine intelligence. However, it is a challenging task to infer causal relationship/structure among variables. In recent years, an Independent Mechanism (IM) principle was proposed, stating that the mechanism gene…
Deviance Voronoi residuals improve earthquake insurance risk assessment.
problem Assessing earthquake insurance risk using spatio-temporal point process models.
method Extended Voronoi residuals and created simulation-based approach.
result Proposed formula for country-wide minimum capital test.
Extends matrix factorization for deviance-based losses with GLM theory.
problem Improving data loss models beyond squared error.
method Adapts GLM theory to matrix factorization for deviance losses.
result Strong consistency and robustness of the proposed decomposition.
This paper aims to review the methodology behind the generalized linear models which are used in analyzing the actuarial situations instead of the ordinary multiple linear regression. We introduce how to assess the adequacy of the model which includes comparing nested models using the deviance and the scaled deviance. …
The paper addresses insurance pricing by improving machine learning models and metrics.
problem Lack of balance and confusion in insurance model performance metrics.
method Introduces autocalibration and Tweedie deviance minimization for insurance pricing models.
result Autocalibration corrects bias and ensures balance on local scales.
Paper introduces NICc for fast cluster-based validation of prediction models.
problem Validation of prediction models on clustered data.
method Derived NICc to approximate leave-one-cluster-out deviance for standard regression models.
result NICc provides more accurate model size and variable selection, especially with strong clustering.
This paper generalizes beta divergence beyond its classical form associated with power variance functions of Tweedie models. Generalized form is represented by a compact definite integral as a function of variance function of the exponential dispersion model. This compact integral form simplifies derivations of many pr…
Bayesian model clusters brain activity time series.
problem Heterogeneous multivariate time series in brain imaging.
method Group-based Bayesian mixture of smoothing splines with covariate effects.
result Distinct brain activity patterns identified.
EGO-MDA identifies optimal spectral-bands for process discrimination.
problem Optimal spectral-bands for process discrimination.
method EGO-MDA, an unsupervised method using EGO and Mixture Discriminant Analysis.
result EGO-MDA achieves at least 70% improvement in median deviance.
Develops a GMM method to estimate roughness in stochastic volatility models.
problem Estimating roughness in stochastic volatility models with fractional Brownian motion.
method GMM approach for log-normal models with integrated variance and noisy realized variance.
result Consistent and asymptotically normal parameter estimator with bias correction.
This work explores the connection between distances and kernels for conditional independence.
problem Measuring conditional independence in various fields like causal discovery and feature selection.
method Investigates the relationship between conditional independence measures induced by distances and reproducing kernels.
result Some kernel-based conditional independence measures are not equivalent to distance-based measures.
Conditional diffusion models can approximate target distributions well with Gaussian-mixture reverse kernels.
problem Approximating target distributions in conditional diffusion models.
method Using finite Gaussian mixtures with ReLU-network logits as reverse kernels, reducing the problem to static conditional density approximation.
result The resulting neural reverse-kernel class is dense in conditional KL divergence under exact terminal matching.
Heat kernel estimates on manifolds with mixed boundary conditions.
problem Estimating heat kernels on manifolds with ends and mixed boundary conditions.
method Global harmonic function construction and h-transform technique. result Two-sided heat kernel estimates for Riemannian manifolds with mixed boundary conditions.
Generative models use kernel smoothing for conditioning on small example sets.
problem Improving generative models' performance with limited conditioning examples.
method Showed that cross-attention conditioning is equivalent to kernel smoothing, specifically a Nadaraya--Watson kernel smoother.
result The approach predicts and confirms three failure regimes for kernel-based conditioning.
New recursive algorithm estimates conditional kernel mean embeddings in Hilbert space.
problem Estimating conditional distributions in RKHS for supervised learning.
method Recursive algorithm in L2 space for conditional kernel mean map. result Strong L2 consistency of recursive estimator proved. A new kernel-based CI test improves on existing methods.
problem Testing conditional independence (CI) in a broad range of dependencies.
method Regression-model-agnostic kernel-based CI test using reproducing kernel Hilbert spaces.
result GKCM outperforms state-of-the-art CI tests in simulations.
The paper compares heat kernels on manifolds with Robin boundary conditions.
problem Comparing heat kernels on manifolds with different boundary conditions.
method Proving comparison theorems for heat kernels on geodesic balls and minimal submanifolds.
result Eigenvalue comparison theorem for the first Robin eigenvalues on minimal submanifolds.
Conditional kernel mean embeddings are nonparametric models that encode conditional expectations in a reproducing kernel Hilbert space. While they provide a flexible and powerful framework for probabilistic inference, their performance is highly dependent on the choice of kernel and regularization hyperparameters. Neve…
New theoretical tools simplify kernel-based tests analysis.
problem Asymptotic behavior of kernel-based tests in various scenarios.
method Avoids complex expansions and limit theorems, works directly with Hilbert spaces random functionals.
result Framework leads to simpler analysis with minimal regularity conditions.
Federated learning calibrates insurance indices from renewable energy producers' data.
problem Calibrating parametric insurance indices under heterogeneous renewable energy production losses.
method Federated learning framework using Tweedie GLMs and distributed optimization.
result Federated learning recovers comparable index coefficients under moderate heterogeneity.
New test for conditional independence using kernel embeddings.
problem Testing conditional independence in high-dimensional settings.
method Analytic kernel embeddings, asymptotic distribution.
result New test outperforms existing methods in high-dimensional settings.
A new method compresses conditional distributions of labelled data.
problem No existing method directly compresses the conditional distribution of labelled data.
method Introduce Average Maximum Conditional Mean Discrepancy (AMCMD), derive a closed form estimator, and extend Kernel Herding (KH) to Average Conditional Kernel Herding (ACKH).
result Directly compressing conditional distributions outperforms joint distribution compression and greedy selection.
Maps between non-compact surfaces can have geometric kernels under certain conditions.
problem Understanding when maps between non-compact surfaces have geometric kernels.
method Using Brown's proper fundamental group to establish sufficient conditions for geometric kernels.
result Characterization of conjugacy classes in the proper fundamental group and sufficient conditions for geometric kernels.
Study Szegő kernel on non-compact CR manifolds with specific conditions.
problem Analyzing Szegő kernel on non-compact CR manifolds.
method Establish Szegő kernel asymptotic expansions on non-compact strictly pseudoconvex CR manifolds with transversal CR R-action under natural geometric conditions. result Szegő kernel asymptotic expansions established on non-compact CR manifolds.
The ratio of two probability densities can be used for solving various machine learning tasks such as covariate shift adaptation (importance sampling), outlier detection (likelihood-ratio test), and feature selection (mutual information). Recently, several methods of directly estimating the density ratio have been deve…
We offer a new, rigorous approach to conditional mean embeddings without operator constraints.
problem Lack of rigorous, operator-free approach to conditional mean embeddings.
method Measure-theoretic approach to conditional mean embeddings.
result Natural regression interpretation and universal consistency of empirical estimates.
Study optimizes CANN for actuarial tasks using RSM.
problem Optimizing hyperparameters for neural networks in actuarial science.
method Factorial design and response surface methodology (RSM).
result Reduced hyperparameter optimization from 288 to 188, achieving near-optimal performance.
Neural-Kernel CME tackles scalability and expressiveness challenges in conditional distribution representation.
problem Scalability and expressiveness challenges in kernel conditional mean embeddings.
method Combines deep learning with CMEs using a neural network optimization framework.
result Achieves competitive and often superior performance in conditional density estimation and RL.
New KCM tests improve specification testing via RKHS.
problem Improving specification tests for econometric models.
method Kernel conditional moment (KCM) tests based on RKHS.
result KCM tests have better finite-sample performance than existing tests.
Gradient descent benefits from tangent kernel advantages under specific conditions.
problem Comparing gradient descent with tangent kernel methods in learning.
method Analysis of gradient descent and tangent kernel methods under different conditions.
result Gradient descent can achieve small error only if tangent kernel methods have a non-trivial advantage, but this advantage can be very small.
The paper constructs optimal confidence bands for kernel gradient flow estimators.
problem Estimating generalization error and constructing confidence bands for kernel gradient flows.
method Established convergence rates and constructed optimal confidence bands under capacity-source condition.
result Optimal confidence bands for kernel gradient flows have shrinkage rates close to minimax optimal rates.
BENK estimates treatment effects with neural kernels for censored data.
problem Estimating heterogeneous treatment effects with censored time-to-event data.
method Proposes a method using the Beran estimator with neural kernels for survival functions.
result Shows improved accuracy compared to existing methods in various scenarios.
Coercivity condition ensures learning of interacting particle systems.
problem Ensuring identifiability of interaction functions in learning systems of interacting particles.
method Equivalence of coercivity condition to strictly positive definiteness of an integral kernel.
result For ergodic systems, the integral kernel is strictly positive definite, satisfying the coercivity condition.
Conditional kernel mean embeddings form an attractive nonparametric framework for representing conditional means of functions, describing the observation processes for many complex models. However, the recovery of the original underlying function of interest whose conditional mean was observed is a challenging inferenc…
Recent developments in system identification have brought attention to regularized kernel-based methods, where, adopting the recently introduced stable spline kernel, prior information on the unknown process is enforced. This reduces the variance of the estimates and thus makes kernel-based methods particularly attract…
Researchers approximate conditional expectation operators using kernel methods.
problem Statistical approximation of conditional expectation operators under minimal assumptions.
method Modifying the domain of the operator, approximating it by Hilbert-Schmidt operators in a reproducing kernel Hilbert space.
result The nonparametric estimate of the operator converges to a specific limiting object.
New conditions ensure MMDs separate and converge to target distributions.
problem Ensuring MMDs separate and converge to target distributions.
method Deriving new sufficient and necessary conditions for MMDs on separable metric spaces.
result First KSDs that exactly metrize weak convergence to P.
Study proves optimal controls for stochastic Volterra equations with singular kernels.
problem Existence of optimal controls for stochastic Volterra equations with singular kernels.
method Sufficient conditions based on integrability and growth hypotheses.
result Existence of optimal relaxed and strict controls under classical convexity assumptions.
Characterizes kernel interpolation in large dimensions, revealing optimal and sub-optimal regions.
problem Understanding the phase diagram of kernel interpolation in large dimensions.
method Characterization of variance and bias under various source conditions.
result Determined the (s,γ)-phase diagram of large-dimensional kernel interpolation. This work closes the theory-practice gap for distributed optimization methods by introducing a new regularity condition.
problem Existing convergence conditions for distributed optimization methods are violated by nearly all kernels used in practice.
method Introduces Hessian relative uniform continuity (HRUC) to guarantee convergence under mild conditions.
result Derives convergence guarantees for mirror descent-based gradient tracking without restrictive assumptions.
Paper tackles conditional expectation estimation using compactification operators.
problem Estimating conditional expectations from product of two random variables.
method Operator theoretic approach using kernel integral operators in reproducing kernel Hilbert space.
result Solutions allow numerical approximation and convergence of data-driven implementations.
Much of machine learning relies on comparing distributions with discrepancy measures. Stein's method creates discrepancy measures between two distributions that require only the unnormalized density of one and samples from the other. Stein discrepancies can be combined with kernels to define kernelized Stein discrepanc…
Sequential Kernel-based Conditional Independence Testing via Adaptive Betting
problem Testing conditional independence
method Testing-by-betting on an adaptively optimized Kernel Conditional Independence statistic
result Significantly reduces Type I error inflation while preserving high power
Paper develops a unified framework for measuring differences between conditional distributions.
problem Comparing conditional distributions in a unified and theoretically sound manner.
method Kernel embeddings and conditional maximum mean discrepancy (CMMD) framework.
result Established a coherent framework for measuring divergence between conditional distributions.