Study on error probability for classification of heavy-tailed renewal processes.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Inverse classification, the process of making meaningful perturbations to a test point such that it is more likely to have a desired classification, has previously been addressed using data from a single static point in time. Such an approach yields inflated probability estimates, stemming from an implicitly made assum…
One of the central themes in the classification task is the estimation of class posterior probability at a new point . The vast majority of classifiers output a score for , which is monotonically related to the posterior probability via an unknown relationship. There are many attempts in the literature …
Binary classification models get more efficient predictive probabilities.
Focal loss improves classification but not class-posterior probability estimation.
Probability distributions produced by the cross-entropy loss for ordinal classification problems can possess undesired properties. We propose a straightforward technique to constrain discrete ordinal probability distributions to be unimodal via the use of the Poisson and binomial probability distributions. We evaluate …
New method corrects skewed confidence for PbN classification.
Study on error probabilities of machine learning classification techniques using large deviations theory.
This work assesses DNNs for estimating conditional probabilities.
A method for classifying points with minimal queries using Hermite polynomials.
A method to monitor probability predictions for calibration loss in image classification models.
Develops a new minimax probability machine for imbalanced classification tasks.
We study the problem of supervised learning for both binary and multiclass classification from a unified geometric perspective. In particular, we propose a geometric regularization technique to find the submanifold corresponding to a robust estimator of the class probability . The regularization term meas…
PosCal training improves classification models by calibrating posterior probabilities.
New method uses graph generative models for graph classification.
New tests for binary classification regression functions without distribution assumptions.
The law of total probability may be deployed in binary classification exercises to estimate the unconditional class probabilities if the class proportions in the training set are not representative of the population class proportions. We argue that this is not a conceptually sound approach and suggest an alternative ba…
In this paper, we present a novel sequential paradigm for classification in crowdsourcing systems. Considering that workers are unreliable and they perform the tests with errors, we study the construction of decision trees so as to minimize the probability of mis-classification. By exploiting the connection between pro…
Bayesian model improves classification performance with flexible uncertainty modeling.
With the widespread success of deep neural networks in science and technology, it is becoming increasingly important to quantify the uncertainty of the predictions produced by deep learning. In this paper, we introduce a new method that attaches an explicit uncertainty statement to the probabilities of classification u…
Reduces false positives in classifying rare online platforms.
We provide a general theoretical analysis of expected out-of-sample utility, also referred to as decision-theoretic classification, for non-decomposable binary classification metrics such as F-measure and Jaccard coefficient. Our key result is that the expected out-of-sample utility for many performance metrics is prov…
Develops geometric framework for uncertainty-aware multi-class classification.
The softmax representation of probabilities for categorical variables plays a prominent role in modern machine learning with numerous applications in areas such as large scale classification, neural language modeling and recommendation systems. However, softmax estimation is very expensive for large scale inference bec…
We present a novel modulation level classification (MLC) method based on probability distribution distance functions. The proposed method uses modified Kuiper and Kolmogorov-Smirnov distances to achieve low computational complexity and outperforms the state of the art methods based on cumulants and goodness-of-fit test…
The majority of traditional classification ru les minimizing the expected probability of error (0-1 loss) are inappropriate if the class probability distributions are ill-defined or impossible to estimate. We argue that in such cases class domains should be used instead of class distributions or densities to construct …
New bounds for balanced classification improve understanding of imbalanced datasets.
Work in the classification literature has shown that in computing a classification function, one need not know the class membership of all observations in the training set; the unlabeled observations still provide information on the marginal distribution of the feature set, and can thus contribute to increased classifi…
Paper analyzes risk bounds for in-context learning in multiclass classification.
Problems of interpolation, classification, and clustering are considered. In the tenets of Radon--Nikodym approach , where the is a linear function on input attributes, all the answers are obtained from a generalized eigenproblem $|f|ψ^{[i]}\rangle =…
Proposes novel wSVMs for sparse learning and accurate probability estimation.
Proposes MCLLO for assessing and recalibrating multiclass probability predictions.
Develops a new framework for estimating joint probability distributions.
Guiding the design of neural networks is of great importance to save enormous resources consumed on empirical decisions of architectural parameters. This paper constructs shallow sigmoid-type neural networks that achieve 100% accuracy in classification for datasets following a linear separability condition. The separab…
Study classifies submanifolds in probability simplex.
New method calibrates classifier probabilities with guaranteed coverage.
Proposes an accuracy-preserving calibration method for DNNs.
Motivated by problems of anomaly detection, this paper implements the Neyman-Pearson paradigm to deal with asymmetric errors in binary classification with a convex loss. Given a finite collection of classifiers, we combine them and obtain a new classifier that satisfies simultaneously the two following properties with …
CIPNN model tackles continuous latent variables, solving intractable posterior problems.
FJS method improves multinomial classification accuracy.
We consider a high dimensional binary classification problem and construct a classification procedure by minimizing the empirical misclassification risk with a penalty on the number of selected features. We derive non-asymptotic probability bounds on the estimated sparsity as well as on the excess misclassification ris…
Subspace models play an important role in a wide range of signal processing tasks, and this paper explores how the pairwise geometry of subspaces influences the probability of misclassification. When the mismatch between the signal and the model is vanishingly small, the probability of misclassification is determined b…
Simplifies machine learning validation using kNN and conditional probability algorithms.
This project explores several Machine Learning methods to predict movie genres based on plot summaries. Naive Bayes, Word2Vec+XGBoost and Recurrent Neural Networks are used for text classification, while K-binary transformation, rank method and probabilistic classification with learned probability threshold are employe…
Multi-label classification aims to classify instances with discrete non-exclusive labels. Most approaches on multi-label classification focus on effective adaptation or transformation of existing binary and multi-class learning approaches but fail in modelling the joint probability of labels or do not preserve generali…
Study ruin probabilities in risk processes on stochastic networks.
New approach for classification using trigonometric polynomial kernels from signal processing.
Develops methods for integrating multivariate normals and computing classification measures.