New bounds for Neyman-Pearson region using -divergences.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Adapts Neyman-Pearson classification for both source and target distribution shifts.
Combines cost-sensitive and Neyman-Pearson paradigms for better binary classification.
Neyman-Pearson testing improves goodness of fit in detecting new physics.
Characterizes distribution-free rates in unbalanced classification problems.
The paper tackles Neyman-Pearson classification control issues.
New method corrects bias in density ratio estimation for missing data.
Develops NPMC method for noisy labels, improving multiclass classification accuracy.
Develops algorithms for multi-class Neyman-Pearson classification with cost sensitivity.
Most existing binary classification methods target on the optimization of the overall classification risk and may fail to serve some real-world applications such as cancer diagnosis, where users are more concerned with the risk of misclassifying one specific class than the other. Neyman-Pearson (NP) paradigm was introd…
A neural network for online NP classification with reduced complexity.
Motivated by problems of anomaly detection, this paper implements the Neyman-Pearson paradigm to deal with asymmetric errors in binary classification with a convex loss. Given a finite collection of classifiers, we combine them and obtain a new classifier that satisfies simultaneously the two following properties with …
Optimal selective classification using likelihood ratios improves model reliability.
Motivated by optimal investment problems in mathematical finance, we consider a variational problem of Neyman-Pearson type for law-invariant robust utility functionals and convex risk measures. Explicit solutions are found for quantile-based coherent risk measures and related utility functionals. Typically, these solut…
In the problem of domain adaptation for binary classification, the learner is presented with labeled examples from a source domain, and must correctly classify unlabeled examples from a target domain, which may differ from the source. Previous work on this problem has assumed that the performance measure of interest is…
New algorithm controls type I error in NP classification under label noise.
Unified framework for Bayes-optimal classifiers under group fairness.
Robust hypothesis testing designs a test for worst-case distributions using kernel methods.
Value-at-Risk (VaR) and Conditional Value-at-Risk (CVaR) are popular risk measures from academic, industrial and regulatory perspectives. The problem of minimizing CVaR is theoretically known to be of Neyman-Pearson type binary solution. We add a constraint on expected return to investigate the Mean-CVaR portfolio sele…
A new method validates generative models in high-dimensional data.
Paper introduces exact credible sets for classification problems.
New algorithm for precise changepoint localization without assumptions.
Optimal classification requires choosing the right group symmetries, contrary to intuition.
This paper addresses the challenges in classifying textual data obtained from open online platforms, which are vulnerable to distortion. Most existing classification methods minimize the overall classification error and may yield an undesirably large type I error (relevant textual messages are classified as irrelevant)…
Robust test for distributions under Hellinger distance, simpler than optimal tests.
This paper proposes a novel optimization principle and its implementation for unsupervised anomaly detection in sound (ADS) using an autoencoder (AE). The goal of unsupervised-ADS is to detect unknown anomalous sound without training data of anomalous sound. Use of an AE as a normal model is a state-of-the-art techniqu…
Minimizes indecisions in selective classification to control misclassification rates.
New method improves certified robustness for classifier confidence.
We study nonzero-sum hypothesis testing games that arise in the context of adversarial classification, in both the Bayesian as well as the Neyman-Pearson frameworks. We first show that these games admit mixed strategy Nash equilibria, and then we examine some interesting concentration phenomena of these equilibria. Our…
Study tests whether trade-off functions are above or below benchmarks using finite samples.
Deep learning models are considered to be state-of-the-art in many offline machine learning tasks. However, many of the techniques developed are not suitable for online learning tasks. The problem of using deep learning models with sequential data becomes even harder when several loss functions need to be considered si…
The Neyman-Pearson (NP) paradigm in binary classification seeks classifiers that achieve a minimal type II error while enforcing the prioritized type I error controlled under some user-specified level . This paradigm serves naturally in applications such as severe disease diagnosis and spam detection, where people h…
The issue of constructing a risk minimizing hedge under an additional almost-surely type constraint on the shortfall profile is examined. Several classical risk minimizing problems are adapted to the new setting and solved. In particular, the bankruptcy threat of optimal strategies appearing in the classical risk minim…
Develops methods for fair classification under linear disparity constraints.
We compute exact values respectively bounds of "distances" - in the sense of (transforms of) power divergences and relative entropy - between two discrete-time Galton-Watson branching processes with immigration GWI for which the offspring as well as the immigration is arbitrarily Poisson-distributed (leading to arbitra…
Environmental acoustic sensing involves the retrieval and processing of audio signals to better understand our surroundings. While large-scale acoustic data make manual analysis infeasible, they provide a suitable playground for machine learning approaches. Most existing machine learning techniques developed for enviro…
New algorithm detects outliers from rare abnormal data.
A novel unified Bayesian framework for network detection is developed, under which a detection algorithm is derived based on random walks on graphs. The algorithm detects threat networks using partial observations of their activity, and is proved to be optimum in the Neyman-Pearson sense. The algorithm is defined by a …
Model change detection is studied, in which there are two sets of samples that are independently and identically distributed (i.i.d.) according to a pre-change probabilistic model with parameter , and a post-change model with parameter , respectively. The goal is to detect whether the change in the model is sign…
Study benchmarks TSC algorithms in distinguishing diffusions using the likelihood ratio test.
Hypothesis testing plays a central role in statistical inference, and is used in many settings where privacy concerns are paramount. This work answers a basic question about privately testing simple hypotheses: given two distributions and , and a privacy level , how many i.i.d. samples are needed to…
Extends likelihood ratio exponential families to analyze various optimization methods.
Quantum states can be learned efficiently using gentle measurements.
A new test method improves goodness-of-fit tests for copulas.
FedSGM tackles constrained federated learning with unified framework.
Feature selection aims to select the smallest subset of features for a specified level of performance. The optimal achievable classification performance on a feature subset is summarized by its Receiver Operating Curve (ROC). When infinite data is available, the Neyman- Pearson (NP) design procedure provides the most e…
Network detection is an important capability in many areas of applied research in which data can be represented as a graph of entities and relationships. Oftentimes the object of interest is a relatively small subgraph in an enormous, potentially uninteresting background. This aspect characterizes network detection as …
RS-Del provides robustness for sequence classifiers against edit distance attacks.