Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,694 papers · 148 categories

Trend · papers per month

2525047551,007 · Jun 202019922001200920172026
48 results for generalized U-statistics

This paper simplifies computing higher-order UU-statistics efficiently.

problem The inefficiency of computing higher-order UU-statistics in practice.
method Decomposition, connection to Einstein summation, and treewidth-based complexity estimate.
result A new, more efficient algorithm to compute UU-statistics.

Paper develops efficient incomplete U-statistics for degenerate cases.

problem High computational cost and non-standard asymptotic behavior in degenerate U-statistics.
method Characterizes dependence structure using hypergraph theory and combinatorial designs, bypassing traditional Hoeffding decomposition.
result Derives a Berry-Esseen bound for incomplete U-statistics of deterministic designs, enabling Gaussian limiting distributions in degenerate cases.

Study on U-statistics with heavy-tailed samples, providing tail bounds and LDP.

problem Deviation of U-statistics with heavy-tailed samples.
method Exponential tail bounds and Large Deviation Principle (LDP) for U-statistics.
result Obtained an exponential upper bound for U-statistics tail decay, showing two regions of decay.

The paper provides bounds for high-dimensional U-statistics with novel order-explicit inequalities.

problem Bounding the deviation of high-dimensional U-statistics from their Hájek projections.
method Develops novel order-explicit moment inequalities for higher-order Hoeffding components.
result The maximum deviation of a high-dimensional U-statistic from its Hájek projection is of order Op(φbn1log2(dn))O_p(φb n^{-1}\log^2(dn)).

U-statistics improve gradient estimation in importance-weighted variational inference.

problem High variance in gradient estimation for importance-weighted variational inference.
method Use U-statistics to average base gradient estimators on overlapping batches of size m, achieving lower variance.
result U-statistic variance reduction leads to modest to significant improvements in inference performance.

Jackknife variance estimation validated for generalized U-statistics.

problem Uncertainty quantification for subsampling-based estimators.
method Jackknife variance estimation for generalized U-statistics with row-wise LrL^r weak law.
result Jackknife and delete-dd variance estimators are ratio-consistent for generalized U-statistics.

High-dimensional U-statistics show surprising phase transitions, impacting kernel-based tests.

problem Understanding phase transitions in high-dimensional U-statistics.
method Proved a convergence theorem for U-statistics of degree two in high dimensions.
result High-dimensional U-statistics can have non-Gaussian limits with larger variance and asymmetry.

New concentration inequality for U-statistics of Markov chains.

problem Proving a concentration inequality for U-statistics of order two in uniformly ergodic Markov chains.
method Inductive analysis using martingale techniques, uniform ergodicity, Nummelin splitting, and Bernstein's inequality.
result Recovery of convergence rate for U-statistics of independent random variables and canonical kernels, with improved results for dependent kernels.

The paper advances U-statistics in dependent settings, improving spectral estimation and goodness-of-fit tests.

problem Non-asymptotic analysis of U-statistics in dependent Markov chain settings.
method Proved new concentration and exponential inequalities for U-statistics, applied to spectral estimation, online algorithms, and goodness-of-fit tests.
result Established new results for spectral estimation, online algorithms, and goodness-of-fit tests in Markov chain settings.

Improved estimation of higher order integrals using shrinkage techniques.

problem Estimating higher order Bochner integrals in non-parametric settings.
method Shrinkage of U-statistic towards a target element, considering kernel degeneracy.
result Consistent shrinkage estimators with fast rates of convergence, even for non-degenerate kernels.

Efficient tests for various statistical problems using incomplete U-statistics.

problem Nonparametric tests for two-sample, independence, and goodness-of-fit problems.
method Proposes MMDAggInc, HSICAggInc, and KSDAggInc tests aggregating over multiple kernel bandwidths.
result Aggregated tests provide a solution to the kernel selection problem and achieve optimal rates.

This paper develops dimension-agnostic inference methods for high-dimensional data.

problem Understanding how classical inference methods behave in high-dimensional settings.
method Using variational representations, sample splitting, and self-normalization to create a refined test statistic.
result The resulting statistic has a Gaussian limiting distribution regardless of how dimensionality scales with sample size.

In this paper, we study the problem of computing UU-statistics of degree 22, i.e., quantities that come in the form of averages over pairs of data points, in the local model of differential privacy (LDP). The class of UU-statistics covers many statistical estimates of interest, including Gini mean difference, Kendal…

2019-10-09abs ↗pdf ↗

New method for MMD with unequal sample sizes improves test power.

problem Existing MMD methods assume equal sample sizes, discarding valuable data.
method Extended generalized U-statistics to handle unequal sample sizes.
result New asymptotic distributions and power optimization for MMD with unequal sample sizes.

Extends conformal prediction for controlling expected risk of monotone loss functions.

problem Controlling expected risk of monotone loss functions.
method Generalizes split conformal prediction with coverage guarantee, extending to distribution shift, quantile risk, multiple, adversarial, and expectations of U-statistics.
result Tight up to an O(1/n)\mathcal{O}(1/n) factor, with worked examples in computer vision and natural language processing.

Proposes a new method to analyze the distributional effects of treatments.

problem Analyzing the full distributional impact of treatments beyond just the mean.
method Uses kernel conditional mean embeddings and U-statistic regression to investigate the CoDiTE.
result Demonstrates the effectiveness of the proposed method through experiments.

Efficient and robust algorithms for decentralized estimation in networks are essential to many distributed systems. Whereas distributed estimation of sample mean statistics has been the subject of a good deal of attention, computation of UU-statistics, relying on more expensive averaging over pairs of observations, is…

2015-11-17abs ↗pdf ↗

This paper develops a general framework for analyzing asymptotics of VV-statistics. Previous literature on limiting distribution mainly focuses on the cases when nn \to \infty with fixed kernel size kk. Under some regularity conditions, we demonstrate asymptotic normality when kk grows with nn by utilizing existin…

2019-12-02abs ↗pdf ↗

New method for efficient matrix completion with nonignorable missing data.

problem Nonignorable missing data in matrix completion.
method Nuclear norm regularized U-statistic loss function and accelerated proximal gradient algorithm.
result Near minimax optimal statistical convergence rate for nonignorable missing data.

The article introduces practical estimators for kernel discrepancies.

problem Estimating kernel discrepancies accurately and efficiently.
method Presented various estimators for MMD, HSIC, and KSD, including V-statistics, U-statistics, and incomplete U-statistics. Stressed the importance of kernel bandwidth and introduced adaptive estimators.
result Adaptive estimators combining multiple estimators with various kernels address the problem of kernel selection.

USP test improves on Pearson's chi-squared and GG-test for independence.

problem Deficiencies in Pearson's chi-squared and GG-test for independence.
method USP test based on UU-statistic estimator of population dependence measure.
result USP test controls size, handles small cell counts, and detects minimal violations of independence.

Study efficient estimation of hidden subspaces in Gaussian Multi-index models.

problem Estimating hidden subspaces in Gaussian Multi-index models with low-dimensional projections.
method Introduced the generative leap exponent and developed an agnostic sequential estimation procedure using spectral U-statistics.
result Achieved optimal sample complexity of $n=Θ(d^{1 \vee \k/2})$ for efficient estimation.

We quantify uncertainty in Oja's algorithm's leading eigenvector estimation.

problem Estimating the error of Oja's algorithm's leading eigenvector from streaming data.
method Combining U-statistics, high-dimensional central limit theorems, and multiplier bootstrap.
result Established a weighted χ² approximation for the error between the eigenvector and algorithm output.

Paper proposes a method to estimate confidence bands for survival random forests.

problem No statistically valid and computationally feasible approach for estimating confidence bands for survival random forests.
method Extending recent developments in infinite-order incomplete U-statistics, the paper proposes an unbiased confidence band estimation.
result The proposed method accurately estimates the confidence band and achieves desired coverage rate.

New unbiased variance estimator for random forests using Hoeffding decomposition.

problem Uncertainty quantification in random forests with large kernel sizes and small sample sizes.
method Proposes a new Hoeffding decomposition view for variance estimation, establishing unbiased estimators and ratio consistency.
result Establishes the ratio consistency of the proposed variance estimator, justifying confidence interval coverage rates.

Paper resolves bias in ALFT training using generalized alignment games.

problem Systematic bias in estimating logarithmic rewards from small batches.
method Generalized Distributional Alignment Games, U-statistics, minimax polynomial estimators, Variance-Optimal Augmented Polynomial Optimization Program (AQP) Estimator.
result Proves optimal bias and accelerated convergence in ALFT training.

Paper analyzes CRL generalization under non-i.i.d. settings, providing bounds for practical data reuse.

problem Limited theoretical understanding of CRL generalization under non-i.i.d. data conditions.
method Inspired by U-statistics, derives generalization bounds for CRL under non-i.i.d. settings.
result Required number of samples scales logarithmically with class covering number.

Unified method for MMD variance estimation improves accuracy and computational efficiency.

problem Variance estimation for MMD in nonparametric testing.
method Unified finite-sample characterization of MMD variance through U-statistic and Hoeffding decomposition; exact acceleration method for univariate case.
result Unified estimators improve accuracy and computational efficiency for MMD variance.

Study improves generalization bounds for machine learning models in the presence of outliers.

problem Improving model robustness against outliers in machine learning.
method Median-of-Means (MoM) estimator and concentration properties analysis under contamination.
result Derives generalization guarantees for pairwise learning in contaminated data.

Study optimizes KSD estimation from samples, revealing Hilbert-Schmidt vs trace scales.

problem Optimizing estimation of Kernel Stein Discrepancy from samples.
method Identifying and comparing minimax scales for U-statistic and V-statistic.
result Hilbert-Schmidt norm of Stein covariance operator gives optimal scale.