Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

151302453604 · Jun 202019922001200920172026
48 results for Gaussian mean testing

A method for rank verification in multivariate Gaussian data, improving on existing approaches.

problem Determining the top KK means in multivariate Gaussian data with any covariance structure.
method Selective inference tools to generalize the two-sided difference-of-means test for any KK and covariance structure.
result The method provides a generalization for rank verification in multivariate Gaussian data with any covariance structure.

Near-optimal private tests for simple and MLR hypotheses developed under Gaussian differential privacy.

problem Developing private tests for simple and MLR hypotheses under Gaussian differential privacy.
method A private mean estimator with data-driven clamping bounds, constructing private test statistics.
result Private tests achieve the same asymptotic relative efficiency as non-private most powerful tests.

Diffusion models generate data with Gaussian Universality, matching linear model test errors.

problem Analyzing the performance of models trained on synthetic data generated by diffusion models.
method Investigates Gaussian Universality for data distributions generated via diffusion models, matching test errors of linear models trained on synthetic data to Gaussian Mixture models.
result The test error of a linear model trained on diffusion-generated data matches the test error of a linear model trained on Gaussian Mixture data with matching means and covariances per class.

Develops new e-processes and confidence sequences for Gaussian means with unknown variance.

problem Constructing valid t-tests and confidence sequences for Gaussian means with unknown variance.
method Explores generalized nonintegrable martingales and extended Ville's inequality, developing two new e-processes and confidence sequences.
result Analyzes the width of resulting confidence sequences with a polynomial dependence on error probability, proving it to be unavoidable and even better than classical fixed-sample t-tests.

Study on estimating Gaussian mean with missing data in high dimensions.

problem Estimating Gaussian mean in high dimensions with missing data due to realizable contamination.
method Statistical Query model, Low-Degree Polynomials, PTF tests, and algorithms.
result Established information-computation gap and developed efficient algorithms.

A novel kernel-based test detects equality versus singularity of two probability measures.

problem Detecting equality versus singularity of two probability distributions.
method Combines kernel mean and kernel covariance embeddings to construct a likelihood ratio test statistic.
result The test statistic satisfies a '0/\infty' law, vanishing under the null and diverging under the alternative.

We consider clustering based on significance tests for Gaussian Mixture Models (GMMs). Our starting point is the SigClust method developed by Liu et al. (2008), which introduces a test based on the k-means objective (with k = 2) to decide whether the data should be split into two clusters. When applied recursively, thi…

2019-10-07abs ↗pdf ↗

This paper introduces a novel clustering algorithm for heteroscedastic Gaussian data without needing to know the number of clusters.

problem Clustering heteroscedastic Gaussian data without prior knowledge of the number of clusters.
method Introduces a novel cost function and fixed-point analysis to estimate centroids, introduces Wald kernel for measurement plausibility, and derives CENTRE-X algorithm.
result CENTRE-X algorithm can estimate centroids without prior knowledge of the number of clusters and performs comparably to standard algorithms K-means and Mean-Shift.

GAAVI offers anytime-valid tests for CMF global null and contrasts.

problem Inference on the conditional mean function for high confidence decisions.
method Asymptotic anytime-valid tests for CMF global null and contrasts.
result Achieves asymptotic type-I error guarantees, power one, and optimal sample complexity.

Study evaluates two-sample tests for validating generative models in high dimensions.

problem Validating the performance and efficiency of non-parametric two-sample tests for high-dimensional generative models.
method Proposes and evaluates the sliced Wasserstein distance, mean of Kolmogorov-Smirnov statistics, and novel sliced Kolmogorov-Smirnov statistic.
result One-dimensional-based tests provide comparable sensitivity to other multivariate metrics but with lower computational cost.

The study examines conditions for achieving a simple lower bound in estimating mean from samples.

problem Achieving a simple lower bound for estimating the mean of a distribution.
method Analyzes conditions for nearly attaining Le Cam's two-point testing lower bound for mean estimation.
result An algorithm nearly attains the two-point testing rate for mixtures of symmetric, log-concave distributions with a common mean.

Gaussian Processes enhance financial forecasting by predicting mean-reverting time series with probability distributions.

problem Accurate long-term financial predictions with probability distributions.
method Functional and augmented data structures for Gaussian Processes.
result Gaussian Processes offer improved long-term predictions with probability distributions.

Study approximates weak error for specific stochastic models with rough and Gaussian mean-reverting volatility.

problem Approximating weak error for specific stochastic models with rough and Gaussian mean-reverting volatility.
method Used Euler type scheme with integrated kernels to study weak convergence rate.
result Obtained weak convergence rate of min(3α1,1)\min(3α-1,1) for discretised rough Ornstein-Uhlenbeck process and stochastic rough volatility model.

Kernel embeddings separate distinct probability distributions, simplifying testing.

problem Testing equality of non-atomic probability distributions.
method Kernel covariance embeddings and Gaussian measures in reproducing kernel Hilbert spaces.
result Testing for singularity between Gaussian measures is equivalent to testing for equality of non-atomic probability distributions.

The study examines the universality of Gaussian data in high-dimensional generalized linear estimation.

problem Understanding when Gaussian data suffices for high-dimensional generalized linear estimation.
method Sharp asymptotic expressions for test and training errors in high-dimensional Gaussian mixture data with labels from a single-index model.
result The universality of Gaussian data in error estimation depends on the alignment between target weights and mixture cluster means and covariances.

Proposes TAGI for efficient Gaussian inference in Bayesian neural networks.

problem Efficient inference in Bayesian neural networks with complex architectures.
method Analytical method for tractable approximate Gaussian inference (TAGI).
result Matches performance of gradient-based methods with O(n)\mathcal{O}(n) computational complexity.

Optimizes cover parameter in Mapper algorithm for better visualization.

problem Tuning the cover parameter in Mapper algorithm to generate a ``nice'' graph.
method Optimizes cover by repeatedly splitting using statistical tests and Gaussian mixture model.
result Algorithm generates covers that retain dataset essence while being faster.

New method for estimating median and mean with high probability privacy.

problem Estimating median and mean with differential privacy.
method Propose, Test, Release (PTR) mechanism with concentration inequalities.
result First sub-Gaussian high probability bounds for differentially private median and mean estimation.

New calibration bands for various distributions improve testing for auto-calibration.

problem Testing for auto-calibration in finite samples is challenging.
method Construct calibration bands for the exponential dispersion family using finite sample properties.
result Calibration bands allow for various tests for calibration and auto-calibration.

Kernel methods are one of the mainstays of machine learning, but the problem of kernel learning remains challenging, with only a few heuristics and very little theory. This is of particular importance in methods based on estimation of kernel mean embeddings of probability measures. For characteristic kernels, which inc…

2016-03-07abs ↗pdf ↗

Study on kernel tests for high-dimensional data, focusing on MMD and CLT.

problem Asymptotic behavior of kernel two-sample tests in high dimensions and large samples.
method Maximum mean discrepancy (MMD) with isotropic kernels, deriving asymptotic expansions and CLT.
result Interplay between moment discrepancy and dimension-and-sample orders in kernel tests.

Proposes a method to select variables for kernel two-sample tests.

problem Determining whether two samples have the same distribution using informative variables.
method A framework based on kernel maximum mean discrepancy (MMD) for selecting a subset of variables.
result The sample size requirements for the three kernels depend on the number of selected variables, not the data dimension.

GaussDetect-LiNGAM eliminates Gaussianity tests for causal discovery.

problem Causal direction identification without Gaussianity assumptions.
method Leverages the equivalence between noise Gaussianity and residual independence in reverse regression.
result Gaussianity tests replaced with robust kernel-based independence tests.

Novel optimization method detects change points in Gaussian data.

problem Detecting change points in univariate Gaussian data sequences.
method Continuous optimization for best subset selection (COMBSS) applied to a reformulated statistical inverse problem.
result Adaptation and evaluation of COMBSS for offline normal mean multiple change-point detection.

Robust covariance testing requires significantly more samples in contaminated data.

problem Testing the covariance matrix of a high-dimensional Gaussian in the presence of contamination.
method We study the problem in the Huber's contamination model, distinguishing between the identity matrix and matrices far from it in Frobenius norm.
result The sample complexity of covariance testing increases dramatically to Ω(d2)Ω(d^2) in the contaminated setting.

The paper develops tests for comparing means in high dimensions with unknown covariance.

problem Testing if the mean of a high-dimensional distribution is close to zero or different from another.
method Develops nonasymptotic tests using concentration inequalities and operator norms.
result Obtains bounds on the minimal separation distance for controlling Type I and Type II errors.

Develops a method for reverse stress testing in multivariate scenarios.

problem Reconstructing a multivariate stress scenario from a single exogenous shock.
method Maximizing conditional density under three distributional assumptions.
result Simulated scenarios are economically coherent and reproduce risk-reward asymmetry.

The paper confirms two groups of gamma-ray bursts using a new nonparametric metric.

problem Determining the number of inherent groups in gamma-ray bursts.
method A new nonparametric interpoint distance-based measure, combined with clustering methods.
result Confirms two groups of short and long gamma-ray bursts.

New sparse Gaussian process method tackles unconstrained regression problems.

problem Dealing with physical systems that satisfy inequality constraints.
method Extends constrained Gaussian process by redefining hat basis functions.
result Reduces computational complexity from O(n3)O(n^{3}) to O(nm2)O(nm^{2}).

A simple method treats heteroscedastic variance variatively, improving model calibration and sample quality.

problem Brittle optimization impacts model likelihoods for mean and variance estimation.
method Proposes a variational approach to heteroscedastic variance, improving predictive mean and variance calibration.
result The proposed method significantly improves parameter calibration and sample quality for regression and VAEs.

The paper analyzes how data augmentation affects the test error in regression models.

problem Understanding the impact of data augmentation on the test error in regression models.
method Characterizes the test error in terms of population quantities and augmentation statistics.
result Provides a tight characterization of the test error in mean squared error.

The study uncovers the breakdown of Gaussian universality in high-dimensional empirical risk minimization.

problem Understanding the breakdown of Gaussian universality in high-dimensional empirical risk minimization.
method Extending the Convex Gaussian Min-Max Theorem to non-Gaussian settings, deriving asymptotic min-max characterizations, and proving asymptotic equivalence of regularizers.
result The projection of the ERM estimator onto a test covariate approximately follows a Gaussian convolution under certain conditions.

Improved inference for heterogeneous multi-output Gaussian processes using natural gradient optimization.

problem Challenges in adaptive gradient optimization for multi-output Gaussian processes.
method Introducing a fully natural gradient scheme to overcome optimization issues.
result Better local optima solutions and higher test performance rates compared to adaptive gradient methods.