New algorithm controls type I error in NP classification under label noise.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Despite the great success of deep neural networks, the adversarial attack can cheat some well-trained classifiers by small permutations. In this paper, we propose another type of adversarial attack that can cheat classifiers by significant changes. For example, we can significantly change a face but well-trained neural…
We formulate statistical watermarking as hypothesis testing and establish near-optimal bounds.
FactTest assesses LLM factuality with Type I error control.
DP synthetic data may inflate statistical test results, caution advised.
In this paper, we propose a generalized scale mixture family of distributions, namely the Power Exponential Scale Mixture (PESM) family, to model the sparsity inducing priors currently in use for sparse signal recovery (SSR). We show that the successful and popular methods such as LASSO, Reweighted and Reweigh…
We study almost-calibrated, -equivariant Lagrangian mean curvature flow in , and prove structural theorems about the Type I and Type II blowups of finite-time singularities. In particular, we prove that any Type I blowup of such a flow must be a special Lagrangian pair of transversely intersecting p…
Study detects signals in spiked Wigner models using log likelihood ratio.
Optimal classification rules control error rates in multiclass mixture models.
Sparse linear (or generalized linear) models combine a standard likelihood function with a sparse prior on the unknown coefficients. These priors can conveniently be expressed as a maximization over zero-mean Gaussians with different variance hyperparameters. Standard MAP estimation (Type I) involves maximizing over bo…
The Neyman-Pearson (NP) paradigm in binary classification seeks classifiers that achieve a minimal type II error while enforcing the prioritized type I error controlled under some user-specified level . This paradigm serves naturally in applications such as severe disease diagnosis and spam detection, where people h…
New method uses CDMs to improve CI testing without distributional assumptions.
In regression settings where explanatory variables have very low correlations and there are relatively few effects, each of large magnitude, we expect the Lasso to find the important variables with few errors, if any. This paper shows that in a regime of linear sparsity---meaning that the fraction of variables with a n…
WHOMP optimizes randomized controlled trials by minimizing subgroup bias.
This paper addresses the challenges in classifying textual data obtained from open online platforms, which are vulnerable to distortion. Most existing classification methods minimize the overall classification error and may yield an undesirably large type I error (relevant textual messages are classified as irrelevant)…
DP-SPRT improves privacy in sequential tests with near-optimal error rates.
The paper tackles data misappropriation in LLMs by embedding watermarks and testing for their presence.
We systematically analyse the necessary and sufficient conditions for the preservation of supersymmetry for bosonic geometries of the form R^{1,9-d} \times M_d, in the common NS-NS sector of type II string theory and also type I/heterotic string theory. The results are phrased in terms of the intrinsic torsion of G-str…
Entropy asymmetry affects regularization in ERM, leading to biased solutions.
Numerical simulations show stability of Type-II singularities in noncompact hypersurfaces.
Local singularity analysis for Ricci flows with applications to bounded scalar curvature.
Gaussian graphical model is a graphical representation of the dependence structure for a Gaussian random vector. It is recognized as a powerful tool in different applied fields such as bioinformatics, error-control codes, speech language, information retrieval and others. Gaussian graphical model selection is a statist…
Proves product metrics are Yamabe metrics under small flat torus conditions.
We provide a condition for spatial curves which rules out the development of a type I singularity. The condition is that after the last time for which an inflection point develops, if the torsion is ever everywhere non-negative, the curve cannot develop a type I singularity.
This review explores resampling techniques for imbalanced binary classification.
Most existing binary classification methods target on the optimization of the overall classification risk and may fail to serve some real-world applications such as cancer diagnosis, where users are more concerned with the risk of misclassifying one specific class than the other. Neyman-Pearson (NP) paradigm was introd…
Motivated by problems of anomaly detection, this paper implements the Neyman-Pearson paradigm to deal with asymmetric errors in binary classification with a convex loss. Given a finite collection of classifiers, we combine them and obtain a new classifier that satisfies simultaneously the two following properties with …
A list of possible holonomy groups contained the exceptional, non-compact Lie group was provided by Fino and Kath. The classification is due to the corresponding holonomy algebras and divided into Type I, II and III, depending on the dimension of the socle being 1,2 or 3, respectively. It was also sh…
In this paper we investigate the singularities of Lagrangian mean curvature flows in by means of smooth singularity models. Type I singularities can only occur at certain times determined by invariants in the cohomology of the initial data. In the type II case, these smooth singularity models are asympto…
SONAR improves outlier detection for streaming data with strong theoretical guarantees.
We characterize the asymptotic performance of nonparametric one- and two-sample testing. The exponential decay rate or error exponent of the type-II error probability is used as the asymptotic performance metric, and an optimal test achieves the maximum rate subject to a constant level constraint on the type-I error pr…
New methods improve anomaly detection with reduced false positives.
We show that a rescale limit at any degenerate singularity of Ricci flow in dimension 3 is a steady gradient soliton. In particular, we give a geometric description of type I and type II singularities.
The paper defines new types of positivity and proves properties of Schur forms for vector bundles.
Identifying statistical dependence between the features and the label is a fundamental problem in supervised learning. This paper presents a framework for estimating dependence between numerical features and a categorical label using generalized Gini distance, an energy distance in reproducing kernel Hilbert spaces (RK…
Curve Shortening Flow preserves circularity for convex projections.
New test for conditional independence using kernel embeddings.
Study on pseudo-Riemannian metrics on Lie groups, finding new non-Einstein examples.
We investigate the problems of identity and closeness testing over a discrete population from random samples. Our goal is to develop efficient testers while guaranteeing Differential Privacy to the individuals of the population. We describe an approach that yields sample-efficient differentially private testers for the…
Efficient tests achieve best error rates in high-dimensional hypothesis testing.
In high-dimensional linear models, the sparsity assumption is typically made, stating that most of the parameters are equal to zero. Under the sparsity assumption, estimation and, recently, inference have been well studied. However, in practice, sparsity assumption is not checkable and more importantly is often violate…
Constructs two types of Eguchi-Hanson metrics with negative scalar curvature.
We consider the weak detection problem in a rank-one spiked Wigner data matrix where the signal-to-noise ratio is small so that reliable detection is impossible. We propose a hypothesis test on the presence of the signal by utilizing the linear spectral statistics of the data matrix. The test is data-driven and does no…
The paper examines translating solitons and their relation to Lagrangian mean curvature flows with zero Maslov class.
New tests for binary classification regression functions without distribution assumptions.
Study ancient solutions on noncompact steady Ricci solitons, proving types of ancient solutions.
We characterize the asymptotic performance of nonparametric goodness of fit testing. The exponential decay rate of the type-II error probability is used as the asymptotic performance metric, and a test is optimal if it achieves the maximum rate subject to a constant level constraint on the type-I error probability. We …
In bankruptcy prediction, the proportion of events is very low, which is often oversampled to eliminate this bias. In this paper, we study the influence of the event rate on discrimination abilities of bankruptcy prediction models. First the statistical association and significance of public records and firmographics i…