Paper proposes a new method to automatically select Gaussian kernel bandwidth for SVDD.
problem Selecting optimal Gaussian kernel bandwidth for SVDD is crucial but challenging.
method Automatic unsupervised method for selecting Gaussian kernel bandwidth.
result The selected bandwidth is competitive with existing methods and can be computed quickly.
This paper proposes a new method for automatically selecting the optimal kernel bandwidth in density estimation.
problem The challenge of selecting the optimal kernel bandwidth in unsupervised density estimation.
method The approach uses a topology-based loss function for automated bandwidth selection.
result Demonstrates the potential of the topology-based approach across different dimensions.
Reduces data transfer for neural network inference on limited bandwidth.
problem Limited bandwidth during neural network inference.
method Automatic selection of relevant input data parts.
result Significant reduction in data transfer without compromising model quality.
A new method for faster bandwidth selection in Gaussian kernel ridge regression.
problem Efficiently selecting the bandwidth in Gaussian kernel ridge regression.
method Formulated an approximate Jacobian expression for bandwidth selection, proposing a closed-form heuristic.
result Our method is as accurate as cross-validation and marginal likelihood maximization but up to six orders of magnitude faster.
New method selects optimal bandwidth for price return density estimation, impacting efficient market hypothesis evaluation.
problem Estimating the complexity of price return distributions using kernel density estimation.
method Proposes a new complexity measure to select optimal bandwidth, avoiding overfitting and underfitting.
result Optimal bandwidth selection leads to clearer evaluation of the efficient market hypothesis.
Algorithm selects variables and bandwidths for geographically weighted regression.
problem Estimating variable subsets and bandwidths for geographically weighted regression.
method Mathematical programming-based approach integrating variable selection and bandwidth estimation.
result Proposed algorithm provides stable spatially varying patterns with competitive explanatory power.
New approach to adaptively select bandwidths in nonparametric regression.
problem Adaptive bandwidth selection in nonparametric regression.
method Inspired by ℓ2-norms of interval projections, introduces a new bandwidth selection procedure. result Obtains non-asymptotic risk bounds for local polynomial regression methods that adapt to local Hölder exponent.
New method selects kernel bandwidth for SVDD and OCSVM.
problem Selecting optimal Gaussian kernel bandwidth for SVDD and OCSVM.
method Exploits low-rank representation of kernel matrix to suggest bandwidth.
result Method performs well for both low-dimensional and high-dimensional data.
New method screens important covariates in ultrahigh-dimensional data.
problem Handling ultrahigh-dimensional data for regression analysis.
method Favored smoothing bandwidth screening followed by iterative recovery.
result Screening method proves model selection consistency.
Changing kernel bandwidth during training improves kernel regression performance.
problem Improving kernel regression performance with varying model complexity.
method Investigated changing the bandwidth of a translational-invariant kernel during training for kernel regression using gradient descent.
result Kernel regression exhibits double descent behavior with decreasing model complexity (bandwidth).
A fast method for selecting Gaussian kernel bandwidth in kernel-based classifiers.
problem High computational complexity in estimating Gaussian kernel bandwidth.
method Developed based on reproducing kernel Hilbert space operators.
result Proposed method outperforms state-of-the-art methods in computational time and performance.
Study provides bounds for estimating intrinsic dimension using Gaussian kernels.
problem Estimating intrinsic dimension from data.
method Finite-sample concentration and anti-concentration bounds for Gaussian kernel sums.
result Explicit dependence on sample size, bandwidth, and geometric parameters.
We explore the performance of several automatic bandwidth selectors, originally designed for density gradient estimation, as data-based procedures for nonparametric, modal clustering. The key tool to obtain a clustering from density gradient estimators is the mean shift algorithm, which allows to obtain a partition not…
Proposes GRAB-MDM for robust multiview data fusion.
problem Limited theoretical guarantees for multiview fusion methods in noisy high-dimensional data.
method Generalized Robust Adaptive-Bandwidth Multiview Diffusion Maps (GRAB-MDM) with adaptive bandwidth selection.
result Adaptive bandwidths lead to robust recovery of shared intrinsic structure in noisy multiview data.
This work improves ASR noise robustness using parallel data and T/S learning.
problem Noise robustness in automatic speech recognition.
method Teacher-student learning with parallel clean and noisy data, logits selection.
result Best student model yields significant WER reductions in noisy conditions.
The study optimizes bandwidth for nonparametric modal clustering.
problem Optimizing bandwidth for nonparametric modal clustering.
method Asymptotic analysis of density-based partitions and bandwidth selection.
result Asymptotic approximation of a metric for partition distance.
Important information concerning a multivariate data set, such as clusters and modal regions, is contained in the derivatives of the probability density function. Despite this importance, nonparametric estimation of higher order derivatives of the density functions have received only relatively scant attention. Kernel …
New framework accelerates particle-based variational inference methods.
problem Improving the accuracy and speed of particle-based variational inference.
method Unified understanding of ParVIs through Wasserstein gradient flows, and acceleration framework based on the geometry of the Wasserstein space.
result Improved convergence and enhanced sample accuracy through the proposed acceleration framework and bandwidth-selection method.
Kernel Estimation is one of the most widely used estimation methods in non-parametric Statistics, having a wide-range of applications, including spot volatility estimation of stochastic processes. The selection of bandwidth and kernel function is of great importance, especially for the finite sample settings commonly e…
Support Vector Data Description (SVDD) provides a useful approach to construct a description of multivariate data for single-class classification and outlier detection with various practical applications. Gaussian kernel used in SVDD formulation allows flexible data description defined by observations designated as sup…
New algorithm reduces communication traffic in decentralized learning.
problem Communication bottleneck in decentralized learning for low-bandwidth workers.
method Sparsification and adaptive peer selection to reduce communication traffic.
result Significant reduction in communication traffic compared to existing methods.
The article derives a novel Gram-Charlier A (GCA) Series based Extended Rule-of-Thumb (ExROT) for bandwidth selection in Kernel Density Estimation (KDE). There are existing various bandwidth selection rules achieving minimization of the Asymptotic Mean Integrated Square Error (AMISE) between the estimated probability d…
PCRs compress data for deep learning, reducing training time.
problem Efficiently training deep learning models over large datasets.
method Combining progressive compression with an efficient storage layout.
result PCRs can tolerate up to 50% compression without significantly affecting training accuracy.
Efficiently clusters large datasets using low-density hyperplanes.
problem Clustering large datasets efficiently.
method Incremental estimation of low-density hyperplanes using stochastic gradient descent.
result The method automatically selects an appropriate number of clusters.
This paper improves bandwidth selectors for SPBNs to enhance their performance.
problem Suboptimal density estimation and reduced predictive performance in SPBNs due to normal rule bandwidth selection.
method Theoretical framework for state-of-the-art bandwidth selectors (cross-validation and plug-in methods) are established and evaluated.
result Cross-validation selectors outperform the normal rule, especially in high sample size scenarios.
The problem of adaptive noisy clustering is investigated. Given a set of noisy observations Zi=Xi+εi, i=1,...,n, the goal is to design clusters associated with the law of Xi's, with unknown density f with respect to the Lebesgue measure. Since we observe a corrupted sample, a direct approach as the popular …
Two adaptive kernel selection methods improve the accuracy of Kernelized Diffusion Maps.
problem Selecting an appropriate kernel for Kernelized Diffusion Maps.
method Two complementary approaches: variational outer loop and unsupervised cross-validation.
result Both methods improve the quality and stability of the recovered eigenfunctions.
Active Federated Learning selects clients to maximize efficiency.
problem Minimizing bandwidth usage and maximizing model accuracy in federated learning.
method Clients are selected with a probability conditioned on the current model and client data to maximize efficiency.
result Reduces the number of required training iterations by 20-70% while maintaining the same model accuracy.
Conditional density estimation generalizes regression by modeling a full density f(yjx) rather than only the expected value E(yjx). This is important for many tasks, including handling multi-modality and generating prediction intervals. Though fundamental and widely applicable, nonparametric conditional density estimat…
Optimal kernel improves estimation accuracy in modal statistical methods.
problem Estimation accuracy of kernel-based modal statistical methods depends on the kernel used.
method The study theoretically shows an optimal kernel that minimizes asymptotic error criterion.
result An optimal kernel minimizes the error criterion when using an optimal bandwidth.
Kernel method outperforms deep neural networks in speech enhancement.
problem Improving single-channel speech enhancement performance.
method Kernel regression with an exponential power kernel and EigenPro iterative method.
result Kernel method consistently outperforms deep neural networks in speech enhancement.
The article introduces practical estimators for kernel discrepancies.
problem Estimating kernel discrepancies accurately and efficiently.
method Presented various estimators for MMD, HSIC, and KSD, including V-statistics, U-statistics, and incomplete U-statistics. Stressed the importance of kernel bandwidth and introduced adaptive estimators.
result Adaptive estimators combining multiple estimators with various kernels address the problem of kernel selection.
A new method optimizes MMD test power by dynamically selecting kernels, overcoming traditional trade-offs.
problem Fixed kernels fail to distinguish certain distributions, leading to overfitting and variance collapse.
method Complexity-Penalized MMD (CP-MMD) criterion, derived from concentration inequality, optimizes kernel selection.
result CP-MMD maximizes true test power while ensuring unconditional Type-I validity, matching or exceeding state-of-the-art performance.
Ad-SVGD optimizes kernel parameters for SVGD, improving inference performance.
problem Efficiently approximating posterior distributions in Bayesian inference.
method Adaptive kernel selection for SVGD dynamics.
result Ad-SVGD outperforms standard heuristics in various tasks.
We propose a procedure for supervised classification that is based on potential functions. The potential of a class is defined as a kernel density estimate multiplied by the class's prior probability. The method transforms the data to a potential-potential (pot-pot) plot, where each data point is mapped to a vector of …
Paper presents an efficient method for selecting machine learning algorithms and hyper-parameters.
problem Efficient selection of machine learning algorithms and hyper-parameters is challenging for large datasets.
method Progressive sampling-based Bayesian optimization
result Significantly reduces search time, classification error rate, and error rate variability.
Optimizes audio codec selection with statistical guarantees.
problem Selecting the best audio encoding scheme for various data types.
method Supervised learning with uniform convergence theory.
result Rigorous statistical guarantees for codec selection.
CKA with Gaussian RBF kernels converges linearly as bandwidth increases.
problem Understanding the behavior of CKA with large bandwidth Gaussian kernels.
method Analyzing the convergence of CKA based on Gaussian RBF kernels in the large-bandwidth limit.
result CKA based on Gaussian RBF kernels converges linearly as bandwidth increases.
Scalable web crawling using noisy change-indicating signals.
problem Optimizing web page freshness with limited bandwidth and noisy side information.
method Proposes a scalable crawling algorithm that uses noisy side information optimally.
result Achieves constant total rate of crawling without spikes in bandwidth usage.
The paper explores the trade-off between recommendation system performance and bandwidth usage.
problem Balancing recommendation system performance with wireless bandwidth constraints.
method Analyzes two scenarios: multi-armed bandit with context and latent structure exploitation.
result Demonstrates a tradeoff between regret and bandwidth usage, with tight bounds for some instances.
New algorithm selects variables from large datasets.
problem Automatic selection of variables from large datasets.
method Uses Graphical Models and combines with OLS method.
result Outperforms LASSO method in forecasting models.
GP-ALPS automatically selects latent processes for multi-output GPs.
problem Manual selection of latent processes in multi-output GPs is time-consuming and prone to biases.
method Developed a variational inference scheme to automatically choose latent processes.
result Demonstrated suitability of GP-ALPS in preliminary experiments.
New method uses reinforcement learning to accurately estimate available network bandwidth.
problem Accurate and fast estimation of available bandwidth in networks with varying cross-traffic.
method Employed reinforcement learning, specifically the ε-greedy algorithm in a multi-armed bandit approach. result Proposed method identifies available bandwidth with high precision and converges under various challenging conditions.
We provide a way to infer about existence of topological circularity in high-dimensional data sets in Rd from its projection in R2 obtained through a fast manifold learning map as a function of the high-dimensional dataset X and a particular choice of a positive real σ known as band…
We investigate a Gaussian mixture model (GMM) with component means constrained in a pre-selected subspace. Applications to classification and clustering are explored. An EM-type estimation algorithm is derived. We prove that the subspace containing the component means of a GMM with a common covariance matrix also conta…
The paper bounds bandwidth and focal radius for manifolds with positive isotropic curvature.
problem Bounding bandwidth and focal radius for manifolds with positive isotropic curvature.
method Using spectral properties of a twisted de Rham-Hodge operator.
result Upper bounds on bandwidth and focal radius are derived for hypersurfaces in PIC manifolds.
New scheme for sparse feature selection in networked data.
problem Sparse feature selection in distributed, communication-restricted networks.
method Distributed sparse linear regression and feature selection method.
result True causal features can be reliably recovered with minimal bandwidth usage.
Consistency of the kernel density estimator requires that the kernel bandwidth tends to zero as the sample size grows. In this paper we investigate the question of whether consistency is possible when the bandwidth is fixed, if we consider a more general class of weighted KDEs. To answer this question in the affirmativ…