Theoretical and empirical study on SMOTE rebalancing strategy for imbalanced data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Deep SMOTE improves SMOTE's stability and accuracy in imbalanced classification.
The Synthetic Minority Oversampling TEchnique (SMOTE) is widely-used for the analysis of imbalanced datasets. It is known that SMOTE frequently over-generalizes the minority class, leading to misclassifications for the majority class, and effecting the overall balance of the model. In this article, we present an approa…
The paper analyzes SMOTE for imbalanced classification, providing theoretical bounds and guidelines.
This paper surveys various data balancing methods for imbalanced datasets.
SMOTE-DP enhances synthetic data privacy without sacrificing utility.
SMOTE is one of the oversampling techniques for balancing the datasets and it is considered as a pre-processing step in learning algorithms. In this paper, four new enhanced SMOTE are proposed that include an improved version of KNN in which the attribute weights are defined by mutual information firstly and then they …
The paper addresses classification imbalance by framing it as a transfer learning problem.
Proposes a new technique for handling imbalanced data.
Learning from class-imbalanced data continues to be a common and challenging problem in supervised learning as standard classification algorithms are designed to handle balanced class distributions. While different strategies exist to tackle this problem, methods which generate artificial data to achieve a balanced cla…
Study improves detection of cryptocurrency pump-and-dump schemes.
Proposes GMOTE for better handling imbalanced data.
This work develops a high precision fault diagnosis classifier using XAI insights.
A novel time series imputation technique using tSMOTE for handling missing data.
We aim at developing and improving the imbalanced business risk modeling via jointly using proper evaluation criteria, resampling, cross-validation, classifier regularization, and ensembling techniques. Area Under the Receiver Operating Characteristic Curve (AUC of ROC) is used for model comparison based on 10-fold cro…
A new algorithm reduces imbalanced data classification errors in multi-class settings.
Forums play an important role in providing a platform for community interaction. The introduction of irrelevant content or spam by individuals for commercial and social gains tends to degrade the professional experience presented to the forum users. Automated moderation of the relevancy of posted content is desired. Ma…
This paper tackles imbalanced classification with weakly supervised oversampling.
Machine learning has automated much of financial fraud detection, notifying firms of, or even blocking, questionable transactions instantly. However, data imbalance starves traditionally trained models of the content necessary to detect fraud. This study examines three separate factors of credit card fraud detection vi…
A new CA-GAN architecture improves minority class data generation in health datasets.
A new algorithm improves credit scoring accuracy for imbalanced data.
This paper tackles imbalanced data in binary classification problems.
INGB improves oversampling for noisy imbalanced datasets.
Proposes a new data augmentation method for imbalanced datasets in both classification and regression.
WOTBoost improves minority class accuracy in imbalanced datasets.
Bayesian network framework assesses urban risks across multiple domains.
Machine Learning has been steadily gaining traction for its use in Anomaly-based Network Intrusion Detection Systems (A-NIDS). Research into this domain is frequently performed using the KDD~CUP~99 dataset as a benchmark. Several studies question its usability while constructing a contemporary NIDS, due to the skewed r…
Study predicts heart failure patient survival using stacked ensemble ML.
Purpose: Malicious web domain identification is of significant importance to the security protection of Internet users. With online credibility and performance data, this paper aims to investigate the use of machine learning tech-niques for malicious web domain identification by considering the class imbalance issue (i…
Machine learning models predict bluebottles' presence on beaches, addressing class imbalance and unreliable absence data.
This paper builds a machine learning model to predict credit defaults for unsecured lending.
In this paper we propose the use of Generative Adversarial Networks (GAN) to generate artificial training data for machine learning tasks. The generation of artificial training data can be extremely useful in situations such as imbalanced data sets, performing a role similar to SMOTE or ADASYN. It is also useful when t…
The study reduces a personality measurement instrument to 10 features with minimal loss of accuracy.
The latest techniques from Neural Networks and Support Vector Machines (SVM) are used to investigate geometric properties of Complete Intersection Calabi-Yau (CICY) threefolds, a class of manifolds that facilitate string model building. An advanced neural network classifier and SVM are employed to (1) learn Hodge numbe…
Cardiovascular diseases are one of the most common causes of death in the world. Prevention, knowledge of previous cases in the family, and early detection is the best strategy to reduce this fact. Different machine learning approaches to automatic diagnostic are being proposed to this task. As in most health problems,…
Imbalanced Learning is an important learning algorithm for the classification models, which have enjoyed much popularity on many applications. Typically, imbalanced learning algorithms can be partitioned into two types, i.e., data level approaches and algorithm level approaches. In this paper, the focus is to develop a…
Data imbalance remains one of the most widespread problems affecting contemporary machine learning. The negative effect data imbalance can have on the traditional learning algorithms is most severe in combination with other dataset difficulty factors, such as small disjuncts, presence of outliers and insufficient numbe…
Class imbalance problems manifest in domains such as financial fraud detection or network intrusion analysis, where the prevalence of one class is much higher than another. Typically, practitioners are more interested in predicting the minority class than the majority class as the minority class may carry a higher misc…
Fake engagement is one of the significant problems in Online Social Networks (OSNs) which is used to increase the popularity of an account in an inorganic manner. The detection of fake engagement is crucial because it leads to loss of money for businesses, wrong audience targeting in advertising, wrong product predicti…
A novel two-stage resampling method improves CNN training on imbalanced colorectal cancer image data.
Study predicts startup outcomes like funding, patenting, IPOs using machine learning.
Generative Adversarial Networks (GANs) have been used in many different applications to generate realistic synthetic data. We introduce a novel GAN with Autoencoder (GAN-AE) architecture to generate synthetic samples for variable length, multi-feature sequence datasets. In this model, we develop a GAN architecture with…
A new sampling method balances imbalanced data using gamma distribution.
Study evaluates three class imbalance techniques across diverse datasets.
RSmote improves PINNs accuracy with less memory usage.
Framework learns to transform majority to minority samples for balanced classification.
Deep learning detects traffic accidents in real time using spatiotemporal data.
The real-time crash likelihood prediction has been an important research topic. Various classifiers, such as support vector machine (SVM) and tree-based boosting algorithms, have been proposed in traffic safety studies. However, few research focuses on the missing data imputation in real-time crash likelihood predictio…