Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

6351,2701,9042,539 · Jun 202019922001200920172026
48 results for Type-I and Type-II errors

Despite the great success of deep neural networks, the adversarial attack can cheat some well-trained classifiers by small permutations. In this paper, we propose another type of adversarial attack that can cheat classifiers by significant changes. For example, we can significantly change a face but well-trained neural…

2018-09-03abs ↗pdf ↗

We formulate statistical watermarking as hypothesis testing and establish near-optimal bounds.

problem Statistical watermarking in the context of hypothesis testing.
method Formulated as a hypothesis testing problem, using coupling of output tokens and rejection regions.
result Established nearly matching upper and lower bounds on the number of i.i.d. tokens required for small Type I and Type II errors.

FactTest assesses LLM factuality with Type I error control.

problem Lack of rigorous factuality verification for LLMs.
method Formulates factuality testing as hypothesis testing, ensuring Type I and II error control.
result Improves model accuracy by over 40% in abstaining from unknown questions.

DP synthetic data may inflate statistical test results, caution advised.

problem Inflated Type I errors in statistical tests on DP-synthetic data.
method Evaluation of Mann-Whitney U test, t-test, chi-squared test, and median test on DP-synthetic data generated from real-world and simulated datasets using various DP-synthetic data generation methods.
result A large portion of evaluation results showed inflated Type I errors, especially at low privacy levels.

We study almost-calibrated, O(n)O(n)-equivariant Lagrangian mean curvature flow in Cn\mathbb{C}^n, and prove structural theorems about the Type I and Type II blowups of finite-time singularities. In particular, we prove that any Type I blowup of such a flow must be a special Lagrangian pair of transversely intersecting p…

2019-10-14abs ↗pdf ↗

Study detects signals in spiked Wigner models using log likelihood ratio.

problem Detecting signals in rank-one spiked Wigner models with non-Gaussian noise.
method Proved asymptotic normality of log likelihood ratio and computed error thresholds.
result Optimal signal-to-noise ratio threshold for reliable detection.

Optimal classification rules control error rates in multiclass mixture models.

problem Classifying observations in multiclass mixture models while controlling error rates.
method Finding optimal classification rules by searching an optimal region in the observation space, using Maximum A Posteriori (MAP) rule and heuristic computation.
result The FDR-like optimal rule can be significantly less conservative than thresholded MAP rules.

Sparse linear (or generalized linear) models combine a standard likelihood function with a sparse prior on the unknown coefficients. These priors can conveniently be expressed as a maximization over zero-mean Gaussians with different variance hyperparameters. Standard MAP estimation (Type I) involves maximizing over bo…

2012-07-10abs ↗pdf ↗

The Neyman-Pearson (NP) paradigm in binary classification seeks classifiers that achieve a minimal type II error while enforcing the prioritized type I error controlled under some user-specified level αα. This paradigm serves naturally in applications such as severe disease diagnosis and spam detection, where people h…

2018-02-07abs ↗pdf ↗

New method uses CDMs to improve CI testing without distributional assumptions.

problem Testing conditional independence when the conditional distribution is unknown.
method Uses conditional diffusion models (CDMs) to approximate XZX|Z and a classifier-based CMI estimator.
result Proposed method performs better than GAN-based CI tests and controls type I and II errors.

In regression settings where explanatory variables have very low correlations and there are relatively few effects, each of large magnitude, we expect the Lasso to find the important variables with few errors, if any. This paper shows that in a regime of linear sparsity---meaning that the fraction of variables with a n…

2015-11-05abs ↗pdf ↗

DP-SPRT improves privacy in sequential tests with near-optimal error rates.

problem Privacy constraints in sequential probability ratio tests.
method A wrapper for SPRT that uses a private mechanism to determine when to stop based on predefined intervals.
result DP-SPRT achieves near-optimal error rates and privacy guarantees.

The paper tackles data misappropriation in LLMs by embedding watermarks and testing for their presence.

problem Detecting data misappropriation in LLMs trained on copyrighted data.
method Embedding watermarks, formulating as hypothesis testing, developing statistical framework, constructing test statistics, determining optimal thresholds, controlling errors, establishing asymptotic optimality.
result The proposed statistical testing framework effectively detects data misappropriation in LLMs.

We systematically analyse the necessary and sufficient conditions for the preservation of supersymmetry for bosonic geometries of the form R^{1,9-d} \times M_d, in the common NS-NS sector of type II string theory and also type I/heterotic string theory. The results are phrased in terms of the intrinsic torsion of G-str…

2003-02-19abs ↗pdf ↗

Entropy asymmetry affects regularization in ERM, leading to biased solutions.

problem Analyzing the impact of relative entropy asymmetry in ERM regularization.
method Examined Type-I and Type-II ERM-RER, comparing their solutions and properties.
result Type-II ERM-RER regularization introduces a strong bias against training data.

Numerical simulations show stability of Type-II singularities in noncompact hypersurfaces.

problem Stability of Type-II singularities in noncompact hypersurfaces with rotationally-symmetric perturbations.
method Adaptation of the overlap method to include angular dependence.
result MCF of noncompact hypersurfaces with angular dependence behaves similarly to rotationally-symmetric perturbations, developing Type-II or Type-I singularities.

Local singularity analysis for Ricci flows with applications to bounded scalar curvature.

problem Understanding the nature of singularities in Ricci flows.
method Local singularity analysis, introducing Type I and Type II singular points, and proving curvature blow-up rates.
result Ricci curvature must blow up at least at a Type I rate near singular points of a Ricci flow.

Gaussian graphical model is a graphical representation of the dependence structure for a Gaussian random vector. It is recognized as a powerful tool in different applied fields such as bioinformatics, error-control codes, speech language, information retrieval and others. Gaussian graphical model selection is a statist…

2017-01-09abs ↗pdf ↗

Proves product metrics are Yamabe metrics under small flat torus conditions.

problem Yamabe metrics on product spaces with small flat tori.
method Extends earlier results to Type~I and Type~II Yamabe constants, QQ-curvature problems, and isoperimetric-ratio type problems.
result Product metrics are Yamabe metrics for sufficiently small flat tori.

This review explores resampling techniques for imbalanced binary classification.

problem Imbalanced classes lead to poor prediction results in classification.
method Classical, cost-sensitive, and Neyman-Pearson paradigms with resampling techniques and classification methods.
result Complex dynamics among resampling techniques, base methods, metrics, and imbalance ratios.

Most existing binary classification methods target on the optimization of the overall classification risk and may fail to serve some real-world applications such as cancer diagnosis, where users are more concerned with the risk of misclassifying one specific class than the other. Neyman-Pearson (NP) paradigm was introd…

2015-08-13abs ↗pdf ↗

Motivated by problems of anomaly detection, this paper implements the Neyman-Pearson paradigm to deal with asymmetric errors in binary classification with a convex loss. Given a finite collection of classifiers, we combine them and obtain a new classifier that satisfies simultaneously the two following properties with …

2011-02-28abs ↗pdf ↗

A list of possible holonomy groups contained the exceptional, non-compact Lie group G2\mathrm{G}_2^{*} was provided by Fino and Kath. The classification is due to the corresponding holonomy algebras and divided into Type I, II and III, depending on the dimension of the socle being 1,2 or 3, respectively. It was also sh…

2019-04-05abs ↗pdf ↗

In this paper we investigate the singularities of Lagrangian mean curvature flows in Cm\mathbf{C}^m by means of smooth singularity models. Type I singularities can only occur at certain times determined by invariants in the cohomology of the initial data. In the type II case, these smooth singularity models are asympto…

2015-05-07abs ↗pdf ↗

SONAR improves outlier detection for streaming data with strong theoretical guarantees.

problem Outlier detection for non-stationary streaming data with high Type I/II errors.
method SONAR is an efficient SGD-based OCSVM solver with strong convex regularization and lifelong learning guarantees.
result SONAR outperforms traditional OCSVM in Type I/II error rates under non-stationary data.

We characterize the asymptotic performance of nonparametric one- and two-sample testing. The exponential decay rate or error exponent of the type-II error probability is used as the asymptotic performance metric, and an optimal test achieves the maximum rate subject to a constant level constraint on the type-I error pr…

2019-08-27abs ↗pdf ↗

The paper defines new types of positivity and proves properties of Schur forms for vector bundles.

problem Defining and characterizing new types of positivity for vector bundles.
method Introducing and characterizing two types of strongly decomposable positivity, proving properties of Schur forms.
result Schur forms of strongly decomposable positive vector bundles are positive or weakly positive, answering a question of Griffiths.

Identifying statistical dependence between the features and the label is a fundamental problem in supervised learning. This paper presents a framework for estimating dependence between numerical features and a categorical label using generalized Gini distance, an energy distance in reproducing kernel Hilbert spaces (RK…

2019-06-05abs ↗pdf ↗

Study on pseudo-Riemannian metrics on Lie groups, finding new non-Einstein examples.

problem Characterizing and finding non-Einstein pseudo-Riemannian metrics on Lie groups.
method Analyzing left invariant metrics, using double extension process, and constructing examples.
result Construction of infinitely many new explicit examples of non-Einstein pseudo-Riemannian metrics on Lie groups.

Efficient tests achieve best error rates in high-dimensional hypothesis testing.

problem Achieving optimal error rates in computationally efficient hypothesis testing.
method Linear spectral statistics and low-degree likelihood ratio analysis.
result An efficient test achieves the best possible error rates among all computationally efficient tests.

In high-dimensional linear models, the sparsity assumption is typically made, stating that most of the parameters are equal to zero. Under the sparsity assumption, estimation and, recently, inference have been well studied. However, in practice, sparsity assumption is not checkable and more importantly is often violate…

2016-10-07abs ↗pdf ↗

We consider the weak detection problem in a rank-one spiked Wigner data matrix where the signal-to-noise ratio is small so that reliable detection is impossible. We propose a hypothesis test on the presence of the signal by utilizing the linear spectral statistics of the data matrix. The test is data-driven and does no…

2018-09-28abs ↗pdf ↗

We study the statistical decision process of detecting the signal from a `signal+noise' type matrix model with an additive Wigner noise. We propose a hypothesis test based on the linear spectral statistics of the data matrix, which does not depend on the distribution of the signal or the noise. The test is optimal unde…

2020-01-16abs ↗pdf ↗

The paper examines translating solitons and their relation to Lagrangian mean curvature flows with zero Maslov class.

problem Understanding the behavior of Lagrangian translating solitons near Type II singularities.
method Analyzes necessary conditions for blow-up limits and applies to open questions.
result Provides a necessary condition for blow-up limits of Lagrangian mean curvature flows with zero Maslov class.

New tests for binary classification regression functions without distribution assumptions.

problem Testing regression functions in binary classification without distributional assumptions.
method Conditional kernel mean embeddings and resampling-based framework.
result Distribution-free hypothesis tests with exact type I error control.