New bounds for Neyman-Pearson region using -divergences.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper tackles Neyman-Pearson classification control issues.
Optimal selective classification using likelihood ratios improves model reliability.
New method corrects bias in density ratio estimation for missing data.
In the problem of domain adaptation for binary classification, the learner is presented with labeled examples from a source domain, and must correctly classify unlabeled examples from a target domain, which may differ from the source. Previous work on this problem has assumed that the performance measure of interest is…
Unified framework for Bayes-optimal classifiers under group fairness.
Adapts Neyman-Pearson classification for both source and target distribution shifts.
Most existing binary classification methods target on the optimization of the overall classification risk and may fail to serve some real-world applications such as cancer diagnosis, where users are more concerned with the risk of misclassifying one specific class than the other. Neyman-Pearson (NP) paradigm was introd…
Motivated by problems of anomaly detection, this paper implements the Neyman-Pearson paradigm to deal with asymmetric errors in binary classification with a convex loss. Given a finite collection of classifiers, we combine them and obtain a new classifier that satisfies simultaneously the two following properties with …
Value-at-Risk (VaR) and Conditional Value-at-Risk (CVaR) are popular risk measures from academic, industrial and regulatory perspectives. The problem of minimizing CVaR is theoretically known to be of Neyman-Pearson type binary solution. We add a constraint on expected return to investigate the Mean-CVaR portfolio sele…
Combines cost-sensitive and Neyman-Pearson paradigms for better binary classification.
Motivated by optimal investment problems in mathematical finance, we consider a variational problem of Neyman-Pearson type for law-invariant robust utility functionals and convex risk measures. Explicit solutions are found for quantile-based coherent risk measures and related utility functionals. Typically, these solut…
Neyman-Pearson testing improves goodness of fit in detecting new physics.
Characterizes distribution-free rates in unbalanced classification problems.
Optimal classification requires choosing the right group symmetries, contrary to intuition.
Robust hypothesis testing designs a test for worst-case distributions using kernel methods.
Develops NPMC method for noisy labels, improving multiclass classification accuracy.
Develops algorithms for multi-class Neyman-Pearson classification with cost sensitivity.
Robust test for distributions under Hellinger distance, simpler than optimal tests.
A neural network for online NP classification with reduced complexity.
Minimizes indecisions in selective classification to control misclassification rates.
A new method validates generative models in high-dimensional data.
The issue of constructing a risk minimizing hedge under an additional almost-surely type constraint on the shortfall profile is examined. Several classical risk minimizing problems are adapted to the new setting and solved. In particular, the bankruptcy threat of optimal strategies appearing in the classical risk minim…
Develops methods for fair classification under linear disparity constraints.
New algorithm controls type I error in NP classification under label noise.
This paper proposes a novel optimization principle and its implementation for unsupervised anomaly detection in sound (ADS) using an autoencoder (AE). The goal of unsupervised-ADS is to detect unknown anomalous sound without training data of anomalous sound. Use of an AE as a normal model is a state-of-the-art techniqu…
Paper introduces exact credible sets for classification problems.
New algorithm for precise changepoint localization without assumptions.
Study tests whether trade-off functions are above or below benchmarks using finite samples.
Study benchmarks TSC algorithms in distinguishing diffusions using the likelihood ratio test.
This paper addresses the challenges in classifying textual data obtained from open online platforms, which are vulnerable to distortion. Most existing classification methods minimize the overall classification error and may yield an undesirably large type I error (relevant textual messages are classified as irrelevant)…
New method improves certified robustness for classifier confidence.
In recent years, constrained optimization has become increasingly relevant to the machine learning community, with applications including Neyman-Pearson classification, robust optimization, and fair machine learning. A natural approach to constrained optimization is to optimize the Lagrangian, but this is not guarantee…
FedSGM tackles constrained federated learning with unified framework.
Extends likelihood ratio exponential families to analyze various optimization methods.
Hypothesis testing plays a central role in statistical inference, and is used in many settings where privacy concerns are paramount. This work answers a basic question about privately testing simple hypotheses: given two distributions and , and a privacy level , how many i.i.d. samples are needed to…
We study nonzero-sum hypothesis testing games that arise in the context of adversarial classification, in both the Bayesian as well as the Neyman-Pearson frameworks. We first show that these games admit mixed strategy Nash equilibria, and then we examine some interesting concentration phenomena of these equilibria. Our…
Network detection is an important capability in many areas of applied research in which data can be represented as a graph of entities and relationships. Oftentimes the object of interest is a relatively small subgraph in an enormous, potentially uninteresting background. This aspect characterizes network detection as …
Deep learning models are considered to be state-of-the-art in many offline machine learning tasks. However, many of the techniques developed are not suitable for online learning tasks. The problem of using deep learning models with sequential data becomes even harder when several loss functions need to be considered si…
Quantum states can be learned efficiently using gentle measurements.
The Neyman-Pearson (NP) paradigm in binary classification seeks classifiers that achieve a minimal type II error while enforcing the prioritized type I error controlled under some user-specified level . This paradigm serves naturally in applications such as severe disease diagnosis and spam detection, where people h…
We compute exact values respectively bounds of "distances" - in the sense of (transforms of) power divergences and relative entropy - between two discrete-time Galton-Watson branching processes with immigration GWI for which the offspring as well as the immigration is arbitrarily Poisson-distributed (leading to arbitra…
New method detects watermarks in LLM-generated text with human edits.
Environmental acoustic sensing involves the retrieval and processing of audio signals to better understand our surroundings. While large-scale acoustic data make manual analysis infeasible, they provide a suitable playground for machine learning approaches. Most existing machine learning techniques developed for enviro…
New algorithm detects outliers from rare abnormal data.
A novel unified Bayesian framework for network detection is developed, under which a detection algorithm is derived based on random walks on graphs. The algorithm detects threat networks using partial observations of their activity, and is proved to be optimum in the Neyman-Pearson sense. The algorithm is defined by a …
Feature selection aims to select the smallest subset of features for a specified level of performance. The optimal achievable classification performance on a feature subset is summarized by its Receiver Operating Curve (ROC). When infinite data is available, the Neyman- Pearson (NP) design procedure provides the most e…
CTI produces efficient prediction intervals with guaranteed coverage.