The paper extends logistic regression for unbounded majority classes and derives asymptotic properties.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new oversampling framework generates minority samples by perturbing majority classes.
M2m method improves deep learning performance on class-imbalanced datasets.
Insider threat detection is getting an increased concern from academia, industry, and governments due to the growing number of malicious insider incidents. The existing approaches proposed for detecting insider threats still have a common shortcoming, which is the high number of false alarms (false positives). The chal…
Proposes BMME for optimizing nonsmooth nonconvex problems with block structure.
We explain why numbers occurring in the classification of polygon spaces coincide with numbers of self-dual equivalence classes of threshold functions, or of regular Boolean functions, or of decisive weighted majority games.
Paper introduces a new performance metric for class imbalance datasets.
In many situations, classes of data points of primary interest also happen to be those that are least numerous. A well-known example is detection of fraudulent transactions among the collection of all financial transactions, the vast majority of which are legitimate. These types of problems fall under the label of `rar…
Class imbalance classification is a challenging research problem in data mining and machine learning, as most of the real-life datasets are often imbalanced in nature. Existing learning algorithms maximise the classification accuracy by correctly classifying the majority class, but misclassify the minority class. Howev…
New bounds on majority voting's accuracy for multi-class classification problems.
Estimates for polynomial operators using determinant majorization and subharmonics.
RaRecognize learns to recognize rare classes in a stream of data.
Personal income distribution in the USA has a well-defined two-class structure. The majority of population (97-99%) belongs to the lower class characterized by the exponential Boltzmann-Gibbs ("thermal") distribution, whereas the upper class (1-3% of population) has a Pareto power-law ("superthermal") distribution. By …
Framework learns to transform majority to minority samples for balanced classification.
Class-imbalance refers to classification problems in which many more instances are available for certain classes than for others. Such imbalanced datasets require special attention because traditional classifiers generally favor the majority class which has a large number of instances. Ensemble of classifiers have been…
Image classification datasets are often imbalanced, characteristic that negatively affects the accuracy of deep-learning classifiers. In this work we propose balancing GAN (BAGAN) as an augmentation tool to restore balance in imbalanced datasets. This is challenging because the few minority-class images may not be enou…
Class-imbalance refers to classification problems in which many more instances are available for certain classes than for others. Such imbalanced datasets require special attention because traditional classifiers generally favor the majority class which has a large number of instances. Ensemble of classifiers have been…
Class imbalance problems manifest in domains such as financial fraud detection or network intrusion analysis, where the prevalence of one class is much higher than another. Typically, practitioners are more interested in predicting the minority class than the majority class as the minority class may carry a higher misc…
Unified approach for federated learning using MM optimization.
Majority Vote is optimal for reliable data labeling under certain conditions.
This paper tackles imbalanced data in binary classification problems.
Logistic regression is a widely used method in several fields. When applying logistic regression to imbalanced data, for which majority classes dominate over minority classes, all class labels are estimated as `majority class.' In this article, we use an F-measure optimization method to improve the performance of logis…
ARCADe detects anomalies in a sequence of tasks with limited data.
Paper establishes sufficient condition for comparing linear combinations of infinite-mean risks.
A number of classification problems need to deal with data imbalance between classes. Often it is desired to have a high recall on the minority class while maintaining a high precision on the majority class. In this paper, we review a number of resampling techniques proposed in literature to handle unbalanced datasets …
The majority of traditional classification ru les minimizing the expected probability of error (0-1 loss) are inappropriate if the class probability distributions are ill-defined or impossible to estimate. We argue that in such cases class domains should be used instead of class distributions or densities to construct …
We present SemEval-2019 Task 8 on Fact Checking in Community Question Answering Forums, which features two subtasks. Subtask A is about deciding whether a question asks for factual information vs. an opinion/advice vs. just socializing. Subtask B asks to predict whether an answer to a factual question is true, false or…
In machine learning, Domain Adaptation (DA) arises when the distribution gen- erating the test (target) data differs from the one generating the learning (source) data. It is well known that DA is an hard task even under strong assumptions, among which the covariate-shift where the source and target distributions diver…
Analysis shows data imbalance slows learning curves for minority and majority classes.
Majorization-minimization algorithms consist of iteratively minimizing a majorizing surrogate of an objective function. Because of its simplicity and its wide applicability, this principle has been very popular in statistics and in signal processing. In this paper, we intend to make this principle scalable. We introduc…
Class imbalance problem has been a challenging research problem in the fields of machine learning and data mining as most real life datasets are imbalanced. Several existing machine learning algorithms try to maximize the accuracy classification by correctly identifying majority class samples while ignoring the minorit…
Learning from many real-world datasets is limited by a problem called the class imbalance problem. A dataset is imbalanced when one class (the majority class) has significantly more samples than the other class (the minority class). Such datasets cause typical machine learning algorithms to perform poorly on the classi…
This paper proposes an approach to detect emotion from human speech employing majority voting technique over several machine learning techniques. The contribution of this work is in two folds: firstly it selects those features of speech which is most promising for classification and secondly it uses the majority voting…
Crowdsourcing has become an effective and popular tool for human-powered computation to label large datasets. Since the workers can be unreliable, it is common in crowdsourcing to assign multiple workers to one task, and to aggregate the labels in order to obtain results of high quality. In this paper, we provide finit…
Decision trees can be biased towards minority class, contrary to belief.
WOTBoost improves minority class accuracy in imbalanced datasets.
Model for detecting rare labels in imbalanced crowdsourcing data.
AREBA algorithm improves learning from imbalanced, nonstationary data.
We investigate connectedness within and across two major groups or assets: i) five popular cryptocurrencies, and ii) six major asset classes plus two commonly employed risk factors. Granger-causality tests uncover six direct channels of causality from the elements of the mainstream assets/risk factors group to digital …
LoRAS improves model performance on imbalanced datasets by better oversampling the minority class.
Proposes a method to balance imbalanced image datasets using capsule-GAN.
Local semi-supervised method improves brain tissue classification in child MRI.
New bounds for optimal transport using Gaussian processes and rate-distortion functions.
A new graph-based sampling method improves classification of imbalanced COVID-19 datasets.
WDL models density curves using Wasserstein distance and flexible mixture models.
In an -framework, we present a few extension theorems for linear operators. We focus the attention on majorant preserving and sandwich preserving types of extensions. These results are then applied to the study of price systems derived by a reasonable restriction of the class of equivalent martingale measures…
Paper introduces a new identifiability criterion for DAGs using conditional variances.
In many healthcare settings, intuitive decision rules for risk stratification can help effective hospital resource allocation. This paper introduces a novel variant of decision tree algorithms that produces a chain of decisions, not a general tree. Our algorithm, -Carving Decision Chain (ACDC), sequentially carves o…