Algorithm selects variables and bandwidths for geographically weighted regression.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper proposes a new method for automatically selecting the optimal kernel bandwidth in density estimation.
Consistency of the kernel density estimator requires that the kernel bandwidth tends to zero as the sample size grows. In this paper we investigate the question of whether consistency is possible when the bandwidth is fixed, if we consider a more general class of weighted KDEs. To answer this question in the affirmativ…
An accurate and fast estimation of the available bandwidth in a network with varying cross-traffic is a challenging task. The accepted probing tools, based on the fluid-flow model of a bottleneck link with first-in, first-out multiplexing, estimate the available bandwidth by measuring packet dispersions. The estimation…
New method selects optimal bandwidth for price return density estimation, impacting efficient market hypothesis evaluation.
We provide a way to infer about existence of topological circularity in high-dimensional data sets in from its projection in obtained through a fast manifold learning map as a function of the high-dimensional dataset and a particular choice of a positive real known as band…
Study provides bounds for estimating intrinsic dimension using Gaussian kernels.
A streaming algorithm estimates quadratic covariation from financial data efficiently.
This paper improves bandwidth selectors for SPBNs to enhance their performance.
Estimators of information theoretic measures such as entropy and mutual information are a basic workhorse for many downstream applications in modern data science. State of the art approaches have been either geometric (nearest neighbor (NN) based) or kernel based (with a globally chosen bandwidth). In this paper, we co…
Estimates bandwidth for CMC initial data sets.
We explore the performance of several automatic bandwidth selectors, originally designed for density gradient estimation, as data-based procedures for nonparametric, modal clustering. The key tool to obtain a clustering from density gradient estimators is the mean shift algorithm, which allows to obtain a partition not…
The article derives a novel Gram-Charlier A (GCA) Series based Extended Rule-of-Thumb (ExROT) for bandwidth selection in Kernel Density Estimation (KDE). There are existing various bandwidth selection rules achieving minimization of the Asymptotic Mean Integrated Square Error (AMISE) between the estimated probability d…
Optimal kernel improves estimation accuracy in modal statistical methods.
New method tightens federated probe-logit distillation rates under varying bandwidths.
Most density-based clustering methods largely rely on how well the underlying density is estimated. However, density estimation itself is also a challenging problem, especially the determination of the kernel bandwidth. A large bandwidth could lead to the over-smoothed density estimation in which the number of density …
The paper analyzes Kernel Density Estimation in high dimensions with varying data and dimensionality.
Kernel Estimation is one of the most widely used estimation methods in non-parametric Statistics, having a wide-range of applications, including spot volatility estimation of stochastic processes. The selection of bandwidth and kernel function is of great importance, especially for the finite sample settings commonly e…
The article introduces practical estimators for kernel discrepancies.
Local Gaussian correlation struggles in tails but a new method improves it.
Kernel Density Estimation is a very popular technique of approximating a density function from samples. The accuracy is generally well-understood and depends, roughly speaking, on the kernel decay and local smoothness of the true density. However concrete statements in the literature are often invoked in very specific …
New approach to adaptively select bandwidths in nonparametric regression.
Proposes a new method for kernel density estimation using stagewise minimization and a simple dictionary.
The problem of adaptive noisy clustering is investigated. Given a set of noisy observations , , the goal is to design clusters associated with the law of 's, with unknown density with respect to the Lebesgue measure. Since we observe a corrupted sample, a direct approach as the popular …
Paper proves MS convergence for radially symmetric kernels with large bandwidths.
We propose a procedure for supervised classification that is based on potential functions. The potential of a class is defined as a kernel density estimate multiplied by the class's prior probability. The method transforms the data to a potential-potential (pot-pot) plot, where each data point is mapped to a vector of …
Proposes SD-KDE for density estimation using debiased kernel density with score-based adjustments.
Paper provides an upper bound for bias of Nadaraya-Watson kernel regression.
CKA with Gaussian RBF kernels converges linearly as bandwidth increases.
Changing kernel bandwidth during training improves kernel regression performance.
A new method for faster bandwidth selection in Gaussian kernel ridge regression.
Conditional density estimation generalizes regression by modeling a full density f(yjx) rather than only the expected value E(yjx). This is important for many tasks, including handling multi-modality and generating prediction intervals. Though fundamental and widely applicable, nonparametric conditional density estimat…
New method for spot volatility estimation with reduced microstructure noise.
The paper bounds bandwidth and focal radius for manifolds with positive isotropic curvature.
Support vector data description (SVDD) is a popular technique for detecting anomalies. The SVDD classifier partitions the whole space into an inlier region, which consists of the region near the training data, and an outlier region, which consists of points away from the training data. The computation of the SVDD class…
Proposes GRAB-MDM for robust multiview data fusion.
Study bandwidth-limited training and inference of language models.
The paper proves convergence of graph Laplacian with kNN self-tuned kernels.
Support vector data description (SVDD) is a popular anomaly detection technique. The SVDD classifier partitions the whole data space into an inlier region, which consists of the region near the training data, and an outlier region, which consists of points away from the training data. The computation of the SVDD classi…
We consider recommendation systems that need to operate under wireless bandwidth constraints, measured as number of broadcast transmissions, and demonstrate a (tight for some instances) tradeoff between regret and bandwidth for two scenarios: the case of multi-armed bandit with context, and the case where there is a la…
Study shows that ridgeless Gaussian kernel regression overfits even with varying bandwidth or dimensionality.
We study the density estimation problem with observations generated by certain dynamical systems that admit a unique underlying invariant Lebesgue density. Observations drawn from dynamical systems are not independent and moreover, usual mixing concepts may not be appropriate for measuring the dependence among these ob…
Kernel density estimation (KDE) is a popular statistical technique for estimating the underlying density distribution with minimal assumptions. Although they can be shown to achieve asymptotic estimation optimality for any input distribution, cross-validating for an optimal parameter requires significant computation do…
A two-step nonparametric method estimates financial systemic risk.
Kernel based methods have shown effective performance in many remote sensing classification tasks. However their performance significantly depend on its hyper-parameters. The conventional technique to estimate the parameter comes with high computational complexity. Thus, the objective of this letter is to propose an fa…
We propose a flexible nonparametric regression method for ultrahigh-dimensional data. As a first step, we propose a fast screening method based on the favored smoothing bandwidth of the marginal local constant regression. Then, an iterative procedure is developed to recover both the important covariates and the regress…
Support Vector Data Description (SVDD) provides a useful approach to construct a description of multivariate data for single-class classification and outlier detection with various practical applications. Gaussian kernel used in SVDD formulation allows flexible data description defined by observations designated as sup…
Estimates time-series drifts from i.i.d. data using a direct Nadaraya-Watson plug-in method.