Improved multiclass classification with class-weighted nearest neighbors.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We investigate how asymmetrizing an impurity function affects the choice of optimal node splits when growing a decision tree for binary classification. In particular, we relax the usual axioms of an impurity function and show how skewing an impurity function biases the optimal splits to isolate points of a particular c…
U-Det improves lung nodule segmentation in CT images.
Imputation of missing data is a common application in various classification problems where the feature training matrix has missingness. A widely used solution to this imputation problem is based on the lazy learning technique, -nearest neighbor (kNN) approach. However, most of the previous work on missing data does…
We introduce a minorization-maximization approach to optimizing common measures of discovery significance in high energy physics. The approach alternates between solving a weighted binary classification problem and updating class weights in a simple, closed-form manner. Moreover, an argument based on convex duality sho…
This paper introduces new loss functions for balanced multi-class classification.
Hedge funds have long been viewed as a veritable "black box" of investing since outsiders may never view the exact composition of portfolio holdings. Therefore, the ability to estimate an informative set of asset weights is highly desirable for analysis. We present a compositional state space model for estimation of an…
This paper classifies strongly nilpotent special multi-flags and their Goursat counterparts.
New bounds for balanced classification improve understanding of imbalanced datasets.
This paper evaluates six strategies for mitigating imbalanced data: oversampling, undersampling, ensemble methods, specialized algorithms, class weight adjustments, and a no-mitigation approach referred to as the baseline. These strategies were tested on 58 real-life binary imbalanced datasets with imbalance rates rang…
We tackle imbalanced classification by weighting losses and derive robust risks.
Study evaluates three class imbalance techniques across diverse datasets.
Training classification models on imbalanced data tends to result in bias towards the majority class. In this paper, we demonstrate how variable discretization and cost-sensitive logistic regression help mitigate this bias on an imbalanced credit scoring dataset, and further show the application of the variable discret…
IB-GAN improves multivariate time series classification under imbalance.
New approach improves domain adaptation with label shift assumptions.
One major challenge in the medication of Parkinson's disease is that the severity of the disease, reflected in the patients' motor state, cannot be measured using accessible biomarkers. Therefore, we develop and examine a variety of statistical models to detect the motor state of such patients based on sensor data from…
Enhances Random Forest for imbalanced functional data classification.
Modified CTGAN-Plus-Features method optimizes asset allocation with CVaR constraint.
Synthetic data augmentation can improve imbalanced classification metrics.
PLD distills knowledge using choice-theoretic Plackett-Luce model.
In analyses of rare-events, regardless of the domain of application, class-imbalance issue is intrinsic. Although the challenges are known to data experts, their explicit impact on the analytic and the decisions made based on the findings are often overlooked. This is in particular prevalent in interdisciplinary resear…
In this paper, we propose a Dual Focal Loss (DFL) function, as a replacement for the standard cross entropy (CE) function to achieve a better treatment of the unbalanced classes in a dataset. Our DFL method is an improvement on the recently reported Focal Loss (FL) cross-entropy function, which proposes a scaling metho…