Proposes HSIC-Lasso for selective inference in non-linear data.
problem Detecting influential features in non-linear and high-dimensional data.
method Model-free HSIC-Lasso based on truncated Gaussians and polyhedral lemma.
result Tight control of type-I error even for small sample sizes.
HSIC-based method explains GNN structures.
problem Interpreting GNNs' complex dynamics.
method Flexible model-agnostic explanation using HSIC.
result Identifies crucial structures in GNNs.
New method selects key features from millions of biological data points.
problem Scalability issue in feature selection for ultra-high dimensional biological data.
method Scaled up HSIC Lasso to handle millions of features.
result Achieves high accuracy with only 20 out of one million features.
GraphLIME explains GNN models by selecting key features locally.
problem Explaining the effectiveness of GNN models is challenging due to complex nonlinear transformations.
method GraphLIME uses HSIC Lasso for nonlinear feature selection in GNN models.
result GraphLIME provides more descriptive explanations than existing methods.
HSIC bottleneck trains deep networks without backpropagation.
problem Training deep neural networks with exploding and vanishing gradients.
method HSIC bottleneck, alternative to cross-entropy loss and backpropagation.
result HSIC bottleneck achieves comparable performance to backpropagation.
Counterexamples show HSIC feature selection misses critical features.
problem Feature selection using HSIC misses important features.
method Feature selection via HSIC maximization.
result HSIC feature selection can miss critical features.
New research optimizes HSIC estimation rate for translation-invariant kernels.
problem Optimizing the rate of HSIC estimation for translation-invariant kernels.
method Proved minimax optimal rate of O(n−1/2) for HSIC estimation. result Optimality of various HSIC estimators proven.
Efficiently explains model outputs using HSIC, a dependence measure.
problem Efficiently explain model outputs for various architectures.
method HSIC, RKHS, Reproducing Kernel Hilbert Spaces, black-box attribution.
result Up to 8 times faster than previous methods while maintaining fidelity.
New method tests causal association using noise contrastive backdoor adjustment.
problem Testing causal association in complex settings with many confounders.
method Backdoor-HSIC (bd-HSIC) using HSIC for independence testing.
result Calibrated and powerful for binary and continuous treatments with many confounders.
Maximizes image representation dependence for self-supervised learning.
problem Learning meaningful image representations from unlabeled data.
method Maximizes Hilbert-Schmidt Independence Criterion (HSIC) between image transformations and identity.
result Matches state-of-the-art performance on ImageNet and other vision tasks.
The paper explores tensor product kernels and their characteristic properties.
problem Understanding when HSIC characterizes independence and MMD with tensor product kernel discriminates probability distributions.
method Study of various notions of characteristic property of tensor product kernels.
result The paper answers questions about the characteristic and universal properties of tensor product kernels.
Sensitivity Maps improve HSIC for better interpretability and scalability.
problem High computational cost and lack of interpretability in HSIC.
method Introduce Sensitivity Maps (SMs) for HSIC, approximate kernels using random features, and provide convergence bounds.
result RHSIC and SMs efficiently approximate HSIC and provide scalable solutions.
Empirical study on dependency in auto-encoding networks using HSIC.
problem Inaccurate measurement of mutual information in DNNs.
method Proposed using Hilbert-Schmidt Independence Criterion (HSIC) to measure dependency between layers in auto-encoding architectures.
result HSIC can measure dependence without density estimation, improving generalization evaluation.
New method speeds up HSIC for multiple variables.
problem Quadratic computational complexity of HSIC for multiple variables.
method Nyström approximation to HSIC for M≥2. result Consistent Nyström HSIC estimator for M≥2. A statistical test of independence may be constructed using the Hilbert-Schmidt Independence Criterion (HSIC) as a test statistic. The HSIC is defined as the distance between the embedding of the joint distribution, and the embedding of the product of the marginals, in a Reproducing Kernel Hilbert Space (RKHS). It has …
This paper improves HSIC-based dimensionality reduction for non-linear kernels.
problem Non-convexity of HSIC objective function for non-linear kernels limits optimization efficiency.
method Spectral optimization algorithm with local guarantees and principled initialization.
result Empirical improvements by a factor of 105 in runtime with lower errors. Unified framework for optimal kernel tests across MMD, HSIC, and KSD.
problem Optimal testing in kernel-based hypothesis testing frameworks.
method Unified derivation of minimax rates, adaptive kernel selection methods.
result Unified power results across MMD, HSIC, and KSD.
New methods integrate nonlinear, sparse, and multi-view aspects for high-dimensional data analysis.
problem Integrating nonlinear dependence, sparsity, and multi-view data in high-dimensional datasets.
method Proposes HSIC-SGCCA, SA-KGCCA, and TS-KGCCA methods for multi-view high-dimensional data analysis.
result HSIC-SGCCA outperforms competing methods in multi-view variable selection.
The article introduces practical estimators for kernel discrepancies.
problem Estimating kernel discrepancies accurately and efficiently.
method Presented various estimators for MMD, HSIC, and KSD, including V-statistics, U-statistics, and incomplete U-statistics. Stressed the importance of kernel bandwidth and introduced adaptive estimators.
result Adaptive estimators combining multiple estimators with various kernels address the problem of kernel selection.
We describe a novel non-parametric statistical hypothesis test of relative dependence between a source variable and two candidate target variables. Such a test enables us to determine whether one source variable is significantly more dependent on a first target variable or a second. Dependence is measured via the Hilbe…
This work improves independence tests for high-dimensional data.
problem Detecting subtle dependencies between high-dimensional random variables with complex distributions.
method Develops two approaches to learn powerful independence tests using variational mutual information and HSIC.
result Optimized HSIC tests generally outperform other approaches on detecting structured dependence.
A new non parametric approach to the problem of testing the independence of two random process is developed. The test statistic is the Hilbert Schmidt Independence Criterion (HSIC), which was used previously in testing independence for i.i.d pairs of variables. The asymptotic behaviour of HSIC is established when compu…
A fast kernel-based measure for sparse linguistic expressions.
problem Efficiently measuring co-occurrence in sparse linguistic data.
method Derives PHSIC from HSIC, estimates it linearly, and uses various kernels.
result Empirically, PHSIC outperforms PMI in accuracy and learning speed.
Method analyzes hyperparameters using HSIC for better neural network performance.
problem Complex hyperparameter spaces in deep learning.
method Goal-oriented sensitivity analysis using HSIC.
result Robust analysis index quantifies hyperparameters' impact.
New statistics improve kernel independence testing efficiency.
problem Improving efficiency in kernel independence testing.
method Adapting martingale MMD construction to joint independence problem.
result Two new statistics achieve finite-sample consistency with linear per-test cost.
Generative model for morphological continuum of normal and pathological states.
problem Identifying trends and features that separate normality and pathology in biomedical images.
method Wasserstein Auto-encoder with HSIC regularization for latent features.
result Model generates a continuum of morphological changes corresponding to side information.
Paper proposes an algorithm to learn DAGs with indirect dependencies.
problem Learning DAGs misses indirect dependencies in local variables.
method Two-phase algorithm using high-order HSIC for local optimization.
result OT algorithm outperforms existing methods in structure estimation.
New method disentangles hidden data structures using HSIC and supervision.
problem Tackles the challenge of interpreting high-dimensional data.
method Supervised Independent Subspace Principal Component Analysis (sisPCA) using HSIC.
result Identifies and separates hidden data structures effectively.
Proposes a new feature selection method integrating feature relationships.
problem Feature selection in machine learning models.
method Integrates feature-feature and feature-target relationships via penalized mRMR.
result Correctly identifies inactive features, reducing false discoveries.
A new kernel test avoids permutations for independence testing.
problem Intractable null distributions of kernel statistics.
method Developed xHSIC and xdCov, avoiding permutations.
result New tests have limiting Gaussian distributions under null.
This work proves the optimal estimation rates for popular kernel discrepancies.
problem Estimating the disagreement of distributions using kernel discrepancies.
method Proving minimax lower bounds for MMD, HSIC, and KSD.
result The minimax lower bound for estimation of MMD, HSIC, and KSD is \( n^{-1/2} \) on general topological spaces.
Discusses MultiFIT for multivariate dependence, comparing it to HSIC tests.
problem Comparing Multiscale Fisher's Independence Test (MultiFIT) to HSIC tests for multivariate dependence.
method Compares MultiFIT to HSIC tests, highlighting exact level control and performance limitations.
result Observes performance limitations of MultiFIT in terms of test power.
DI-SVM improves brain condition decoding performance via domain independence.
problem Transfer learning in brain imaging data with large p and small n.
method DI-SVM minimizes domain dependence via HSIC to learn common features.
result DI-SVM outperforms eight competing methods on brain decoding tasks.
HSIC loss improves robust regression and classification models.
problem Learning robust models in the absence of target domain data.
method Adapted HSIC loss for unsupervised covariate shift.
result Models outperform standard methods on covariate shift tasks.
Proposes a deep network for multi-class classification using spectral training and Gaussian kernel.
problem Multi-class classification with deep networks.
method Spectral training with linear weights and Gaussian kernel activation, constrained on Stiefel Manifold.
result Theoretical guarantee of global optimum and insight into network generalization.
CDSSL improves representation quality by integrating linear and nonlinear dependencies.
problem Scarcity of labeled data and neglect of nonlinear dependencies in SSL.
method CDSSL combines linear correlations and nonlinear dependencies using HSIC in RKHS.
result CDSSL enhances representation quality on diverse benchmarks.
Novel kernel-based PSI algorithm handles non-linearity and structured data.
problem Non-linearity and structured data in independence measures.
method Develops a PSI algorithm using HSIC, capable of handling non-linearity and structured data.
result Successfully identifies important features in real-world data.
This work learns kernels for structured prediction using polynomial transformations.
problem Learning effective kernel functions for structured prediction.
method Polynomial kernel transformations (Schoenberg transforms and Gegenbaur transforms) learned using HSIC and matrix decomposition.
result State-of-the-art results on real-world datasets.
Paper proposes a hybrid loss function for graph self-supervised learning.
problem Improving self-supervised representation learning for graphs.
method Hybrid loss function combining VICReg and HSIC.
result Hybrid loss function outperformed other methods in 4 out of 7 cases.
The paper uses a graph autoencoder to learn unbiased plant-pollinator interaction embeddings.
problem Sampling bias in citizen science data affects ecological network analysis.
method Bipartite graph variational autoencoder with HSIC for fairness.
result The method mitigates sampling bias and provides unbiased embeddings.
Framework for generating multiple clusterings from multi-view data.
problem Challenges in finding optimal clustering criteria and handling incomplete multi-view data.
method DiMVMC framework that optimizes multiple decoder deep networks to complete data views and generate shared representations.
result DiMVMC outperforms state-of-the-art competitors in generating multiple clusterings with high diversity and quality.
Cheap permutation tests speed up distribution testing without sacrificing accuracy.
problem Efficiently testing distribution differences and independence.
method Group datapoints into bins and permute only these bins, using stored sufficient statistics.
result Cheap permutation tests maintain the accuracy and optimality of standard tests but are significantly faster.
Proposes MVMC for multi-view clustering, enhancing diversity and quality.
problem Leveraging multi-view data for diverse clustering.
method Adapts multi-view self-representation learning, HSIC for redundancy reduction, matrix factorization.
result Generates multiple high-quality and diverse clusterings from multi-view data.
New bounds for Lasso and Group Lasso in high dimensions derived.
problem Estimation error bounds for Lasso and Group Lasso in high-dimensional settings.
method Recent advances in high-dimensional statistics to derive new L2 estimation upper bounds.
result Bounds match optimal minimax rate for Lasso and improve over existing results for Group Lasso.
PLS-Lasso integrates dimension reduction into regression for financial index tracking.
problem Dimension reduction and regression are traditionally treated separately in multivariate data analysis.
method PLS-Lasso integrates dimension reduction directly into the regression process, presenting two formulations: PLS-Lasso-v1 and PLS-Lasso-v2.
result PLS-Lasso-v1 and PLS-Lasso-v2 outperform Lasso in financial index tracking.
The paper examines Adaptive Lasso and Transfer Lasso, highlighting their differences and proposing a new method.
problem Comparing and contrasting Adaptive Lasso and Transfer Lasso.
method Theoretical analysis of asymptotic properties and introduction of a new method.
result The Transfer Lasso method reduces non-asymptotic estimation errors compared to Adaptive Lasso.
New proof shows weighted fused lasso has O(n^2) segments.
problem Proving the complexity of the weighted fused lasso.
method New proof showing the number of segments is O(n^2).
result The solution path of the weighted fused lasso has O(n^2) segments.
Bayesian Lasso Sparse model provides sparse estimates in linear and nonlinear regression.
problem Sparse learning in regression models.
method Develops a new sparse learning model using type-II maximum likelihood procedure.
result The BLS model provides sparse estimates and is more precise, especially with noisy data.