Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

103207310413 · Jun 202019922001200920172026
48 results for closeness testing

We consider the closeness testing problem for discrete distributions. The goal is to distinguish whether two samples are drawn from the same unspecified distribution, or whether their respective distributions are separated in L1L_1-norm. In this paper, we focus on adapting the rate to the shape of the underlying distri…

2019-02-01abs ↗pdf ↗

A new measure scales MMD to assess distribution closeness.

problem Testing statistical significance of distribution closeness.
method Norm-adaptive MMD (NAMMD) for distributional discrepancy.
result NAMMD-based DCT has higher test power than MMD-based DCT.

Study improves sample complexity for distinguishing continuous distributions and causal relationships.

problem Distinguishing continuous distributions and causal relationships in the presence of unobserved confounding.
method Proposed an estimator of KL divergence based on von Mises expansion for closeness testing.
result Established sample complexity guarantees for causal discovery in non-linear models with continuous variables and unobserved confounding.

Bayes-optimal learning of deep random networks with Gaussian weights is studied.

problem Learning a target function corresponding to a deep, extensive-width, non-linear neural network with random Gaussian weights.
method Closed-form expressions for Bayes-optimal test error, ridge regression, kernel and random features regression are computed.
result Optimally regularized ridge regression and kernel regression achieve Bayes-optimal performances, while logistic loss yields a near-optimal test error for classification.

This work improves independence tests for high-dimensional data.

problem Detecting subtle dependencies between high-dimensional random variables with complex distributions.
method Develops two approaches to learn powerful independence tests using variational mutual information and HSIC.
result Optimized HSIC tests generally outperform other approaches on detecting structured dependence.

The paper classifies test configurations and derives a criterion for uniform K-stability of certain algebraic varieties.

problem Uniform K-stability of GG-varieties of complexity 1.
method Classification of GG-equivariant normal test configurations via combinatorial data and derivation of a criterion for uniform K-stability.
result Derivation of a criterion for uniform K-stability in terms of combinatorial data.

The paper develops a new method to test if two multidimensional distributions are equivalent or significantly different.

problem Testing equivalence of multidimensional distributions with sub-linear sample complexity.
method Uses generalized A_k distance and Ramsey theory to develop a computationally efficient closeness tester.
result First sub-linear sample complexity closeness tester for multidimensional distributions.

The paper presents an approximate formula for European mortgage options pricing.

problem Pricing European mortgage options with accuracy and efficiency.
method Approximation of the underlying price distribution using lognormal distributions and matching moments.
result The proposed formula provides a good approximation with high accuracy compared to Monte Carlo simulations.

In this paper we consider a Lagrange Multiplier-type test (LM) to detect change in the mean of time series with heteroskedasticity of unknown form. We derive the limiting distribution under the null, and prove the consistency of the test against the alternative of either an abrupt or smooth changes in the mean. We perf…

2011-02-26abs ↗pdf ↗

Optimal testing of discrete distributions with high probability, achieving sample complexity bounds.

problem Testing discrete distributions with high probability accuracy.
method Characterizing sample complexity as a function of parameters like δ, providing sample-optimal testers.
result Optimal algorithms for closeness and independence testing, achieving within constant factors of information-theoretic lower bounds.

Agent-based model compares different COVID-19 testing policies and their effectiveness.

problem Understanding how different testing policies reveal the true number of infected cases.
method Developed an agent-based simulation framework in Python to model various testing policies and interventions.
result Contact Tracing consistently captures more positive cases than Random Symptomatic Testing, and LBT performs similarly.

We propose a new setting for testing properties of distributions while receiving samples from several distributions, but few samples per distribution. Given samples from ss distributions, p1,p2,,psp_1, p_2, \ldots, p_s, we design testers for the following problems: (1) Uniformity Testing: Testing whether all the pip_i's are …

2019-11-17abs ↗pdf ↗

We build new test sets for the CIFAR-10 and ImageNet datasets. Both benchmarks have been the focus of intense research for almost a decade, raising the danger of overfitting to excessively re-used test sets. By closely following the original dataset creation processes, we test to what extent current classification mode…

2019-02-13abs ↗pdf ↗

We study distribution testing with communication and memory constraints in the following computational models: (1) The {\em one-pass streaming model} where the goal is to minimize the sample complexity of the protocol subject to a memory constraint, and (2) A {\em distributed model} where the data samples reside at mul…

2019-06-11abs ↗pdf ↗

The study examines property testing and estimation under non-identically distributed samples, finding necessary and sufficient sample complexities.

problem Property testing and estimation under non-identically distributed samples.
method Analysis of distributional property testing and estimation in settings with heterogeneous entities.
result Necessary and sufficient sample complexities for property testing and estimation under non-identically distributed samples.

Study robust hypothesis testing under Hellinger distance, proving lower bounds and providing tests.

problem Testing close variants of specified distributions robustly to Hellinger distance.
method Lower bound on slack factor, testing with Hellinger balls, symmetric chi-squared distance analysis.
result Lower bound on slack factor quantifies robustness under misspecification.

We present a new, practical algorithm to test whether a knot complement contains a closed essential surface. This property has important theoretical and algorithmic consequences; however, systematically testing it has until now been infeasibly slow, and current techniques only apply to specific families of knots. As a …

2012-12-07abs ↗pdf ↗

Sharp fractional Sobolev inequalities on closed manifolds identified.

problem Critical fractional Sobolev embedding on closed Riemannian manifolds.
method Intrinsic heat-kernel based framework, determining optimal coefficients, proving sharp inequalities.
result Sharp pp-power inequality and almost sharp inequality established.

The paper develops tests for comparing means in high dimensions with unknown covariance.

problem Testing if the mean of a high-dimensional distribution is close to zero or different from another.
method Develops nonasymptotic tests using concentration inequalities and operator norms.
result Obtains bounds on the minimal separation distance for controlling Type I and Type II errors.

Study characterizes training and test risks for MAP regression with Gaussian priors.

problem Understanding high-dimensional behavior of regularized linear regression with informative priors.
method Maximum a posteriori (MAP) regression with Gaussian priors, using random matrix theory.
result Closed-form risk formulas reveal the bias-variance-prior tradeoff and explain double descent.

Estimates and tests treatment effects on entire outcome distributions.

problem Treatment effects on entire outcome distributions, not just averages.
method Proposes a novel estimand and doubly robust estimator, develops a test.
result First test with provably valid type 1 error guarantees in this setting.

Study on continuous sequence classification with distribution uncertainty.

problem Classifying continuous sequences with varying distribution uncertainty.
method Proposes distribution-free tests for three test designs: fixed-length, sequential, and two-phase tests.
result Error probabilities decay exponentially fast for all test designs.

Test assesses if a linear classifier is random or significant.

problem Determining if a linear classifier captures meaningful differences between classes.
method Proposes a homogeneity test related to linear separability, establishes upper bounds for p-values.
result Upper bounds for p-values are highly accurate for normally distributed samples.

Alternative closed-form formula for spread call option prices under log-normal models.

problem Valuation of spread call options under log-normal models.
method Developed an alternative closed-form formula for spread call option prices.
result Our formula performs better for certain range of model parameters than existing closed-form formula.

Kernel methods are one of the mainstays of machine learning, but the problem of kernel learning remains challenging, with only a few heuristics and very little theory. This is of particular importance in methods based on estimation of kernel mean embeddings of probability measures. For characteristic kernels, which inc…

2016-03-07abs ↗pdf ↗

We show that 3-braid links with given (non-zero) Alexander or Jones polynomial are finitely many, and can be effectively determined. We classify among closed 3-braids strongly quasipositive and fibered ones, and show that 3-braid links have a unique incompressible Seifert surface. We also classify the positive braid wo…

2006-06-19abs ↗pdf ↗

Optimal AFs minimize RFR test error and sensitivity.

problem Finding optimal AFs for RFR to minimize test error and sensitivity.
method Closed-form solution for AFs minimizing test error and sensitivity under different functional parsimony.
result Optimal AFs can be linear, saturated linear, or Hermite polynomial expressions.

Testing independence is of significant interest in many important areas of large-scale inference. Using extreme-value form statistics to test against sparse alternatives and using quadratic form statistics to test against dense alternatives are two important testing procedures for high-dimensional independence. However…

2015-12-30abs ↗pdf ↗

We consider the problem of closeness testing for two discrete distributions in the practically relevant setting of \emph{unequal} sized samples drawn from each of them. Specifically, given a target error parameter ε>0\varepsilon > 0, m1m_1 independent draws from an unknown distribution p,p, and m2m_2 draws from an unkno…

2015-04-17abs ↗pdf ↗

There has been significant study on the sample complexity of testing properties of distributions over large domains. For many properties, it is known that the sample complexity can be substantially smaller than the domain size. For example, over a domain of size nn, distinguishing the uniform distribution from distrib…

2019-07-06abs ↗pdf ↗

Adversarial training achieves optimal test error for shallow networks.

problem Achieving optimal adversarial test error for general data distributions.
method Applying new Rademacher complexity bounds and properties of optimal adversarial predictors.
result Adversarial training can achieve optimal adversarial test error for general data distributions.

Proposes counterfactual explanations for deep two-sample tests on high-dimensional data.

problem Limited interpretability of deep two-sample tests on high-dimensional data.
method Combines diffusion autoencoder and pretrained deep two-sample test model to generate counterfactuals.
result Counterfactual transformations increase p-values, indicating closer distribution similarity.