We consider a problem of learning a binary classifier only from positive data and unlabeled data (PU learning) and estimating the class-prior in unlabeled data under the case-control scenario. Most of the recent methods of PU learning require an estimate of the class-prior probability in unlabeled data, and it is estim…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New method needed for class prior estimation when covariates are reduced.
We consider the problem of estimating the class prior in an unlabeled dataset. Under the assumption that an additional labeled dataset is available, the class prior can be estimated by fitting a mixture of class-wise data distributions to the unlabeled data distribution. However, in practice, such an additional labeled…
Estimates class prior for unlabeled data using kernel embedding.
Adversarial meta-learning computes Gamma-minimax estimators for vague prior knowledge.
Friedman's method performs well for estimating class distributions.
New priors can update posteriors without re-estimating likelihoods.
The problem of developing binary classifiers from positive and unlabeled data is often encountered in machine learning. A common requirement in this setting is to approximate posterior probabilities of positive and negative classes for a previously unseen data point. This problem can be decomposed into two steps: (i) t…
We introduce Fisher consistency in the sense of unbiasedness as a desirable property for estimators of class prior probabilities. Lack of Fisher consistency could be used as a criterion to dismiss estimators that are unlikely to deliver precise estimates in test datasets under prior probability and more general dataset…
Recent advances in weakly supervised classification allow us to train a classifier only from positive and unlabeled (PU) data. However, existing PU classification methods typically require an accurate estimate of the class-prior probability, which is a critical bottleneck particularly for high-dimensional data. This pr…
GS-BSE improves label shift estimation by smoothing priors on a graph.
Work in the classification literature has shown that in computing a classification function, one need not know the class membership of all observations in the training set; the unlabeled observations still provide information on the marginal distribution of the feature set, and can thus contribute to increased classifi…
Bayesian framework estimates label shift for improved classifier performance.
Regression Prior Networks improve ensemble performance on regression tasks.
We investigate the problem of estimating the causal effect of a treatment on individual subjects from observational data, this is a central problem in various application domains, including healthcare, social sciences, and online advertising. Within the Neyman Rubin potential outcomes model, we use the Kullback Leibler…
This paper simplifies finding least favorable priors by reducing dimensionality.
Bottlenecks of binary classification from positive and unlabeled data (PU classification) are the requirements that given unlabeled patterns are drawn from the test marginal distribution, and the penalty of the false positive error is identical to the false negative error. However, such requirements are often not fulfi…
Novel prior for orthogonal functions improves functional component estimation.
The aim of this paper is to provide some theoretical understanding of quasi-Bayesian aggregation methods non-negative matrix factorization. We derive an oracle inequality for an aggregated estimator. This result holds for a very general class of prior distributions and shows how the prior affects the rate of convergenc…
A new method ReCPE removes the need for a distributional assumption in PU learning.
The quantification problem consists of determining the prevalence of a given label in a target population. However, one often has access to the labels in a sample from the training population but not in the target population. A common assumption in this situation is that of prior probability shift, that is, once the la…
In this paper, we consider a class of prescribed Weingarten curvature equations. Under some sufficient condition, we obtain an existence result by the standard degree theory based on the a prior estimates for the solutions to the prescribed Weingarten curvature equations.
We develop a classification algorithm for estimating posterior distributions from positive-unlabeled data, that is robust to noise in the positive labels and effective for high-dimensional data. In recent years, several algorithms have been proposed to learn from positive-unlabeled data; however, many of these contribu…
This report works out the details of a closed-form, fully Bayesian, multiclass, openset, generative pattern classifier using multivariate Gaussian likelihoods, with conjugate priors. The generative model has a common within-class covariance, which is proportional to the between-class covariance in the conjugate prior. …
SJS model predicts label shifts in multinomial datasets.
A novel minimax classifier tackles imbalanced datasets with few minority samples.
Point estimation of class prevalences in the presence of data set shift has been a popular research topic for more than two decades. Less attention has been paid to the construction of confidence and prediction intervals for estimates of class prevalences. One little considered question is whether or not it is necessar…
PriorVAE uses VAEs to efficiently encode spatial priors for small-area estimation.
Bayesian neural networks' performance varies with prior choice, affecting their ability to identify unknowns.
GANs as priors improve uncertainty quantification in complex fields.
A simple regularization method improves model generalization.
One of the central themes in the classification task is the estimation of class posterior probability at a new point . The vast majority of classifiers output a score for , which is monotonically related to the posterior probability via an unknown relationship. There are many attempts in the literature …
In real-world classification problems, the class balance in the training dataset does not necessarily reflect that of the test dataset, which can cause significant estimation bias. If the class ratio of the test dataset is known, instance re-weighting or resampling allows systematical bias correction. However, learning…
This paper improves multi-class calibration methods using mutual information maximization-based binning.
Bayesian approach improves ODE solution accuracy.
This paper introduces hierarchical Gaussian process priors for neural networks to capture weight correlations and inductive biases.
Reduces quantifier variance with accuracy optimization of base classifier.
Bayesian metalearning improves performance in linear bandits with misspecified priors.
We consider robust covariance estimation with group symmetry constraints. Non-Gaussian covariance estimation, e.g., Tyler scatter estimator and Multivariate Generalized Gaussian distribution methods, usually involve non-convex minimization problems. Recently, it was shown that the underlying principle behind their succ…
Proposes a new method to control FDR using frequentist-assisted horseshoe for high-dimensional testing.
New algorithms improve rank one signal estimation from noisy data.
Ensemble approaches for uncertainty estimation have recently been applied to the tasks of misclassification detection, out-of-distribution input detection and adversarial attack detection. Prior Networks have been proposed as an approach to efficiently \emph{emulate} an ensemble of models for classification by paramete…
New algorithm proves convergence for MAP estimation with denoisers.
Improved flow-based models capture dependencies better with multi-scale autoregressive priors.
This paper studies the problem of learning with augmented classes (LAC), where augmented classes unobserved in the training data might emerge in the testing phase. Previous studies generally attempt to discover augmented classes by exploiting geometric properties, achieving inspiring empirical performance yet lacking t…
Two EM algorithms estimate prior distributions in mixture of linear regressions.
Paper proposes an efficient algorithm for nonnegative binary matrix factorization.
The Bayesian framework is a well-studied and successful framework for inductive reasoning, which includes hypothesis testing and confirmation, parameter estimation, sequence prediction, classification, and regression. But standard statistical guidelines for choosing the model class and prior are not always available or…