Study compares BERT and XLNet for multi-class categorization of product descriptions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study categorizes mutual funds using natural language processing from unstructured data.
Automated feature selection is important for text categorization to reduce the feature size and to speed up the learning process of classifiers. In this paper, we present a novel and efficient feature selection framework based on the Information Theory, which aims to rank the features with their discriminative capacity…
This research sets limits on how complex multi-class learning problems can be.
Develops a machine learning framework for identifying authorship in texts.
Paper develops new algorithms for unsupervised multi-class domain adaptation.
Paper investigates methods to improve classification by inducing a hierarchy from flat labels.
Product categorization using text data for eCommerce is a very challenging extreme classification problem with several thousands of classes and several millions of products to classify. Even though multi-class text classification is a well studied problem both in academia and industry, most approaches either deal with …
Develops algorithms for multi-class Neyman-Pearson classification with cost sensitivity.
This paper reviews metrics for evaluating multi-class classification models.
This work presents a new strategy for multi-class classification that requires no class-specific labels, but instead leverages pairwise similarity between examples, which is a weaker form of annotation. The proposed method, meta classification learning, optimizes a binary classifier for pairwise similarity prediction a…
We consider a problem of risk estimation for large-margin multi-class classifiers. We propose a novel risk bound for the multi-class classification problem. The bound involves the marginal distribution of the classifier and the Rademacher complexity of the hypothesis class. We prove that our bound is tight in the numbe…
Paper proposes GEG to enhance fairness in binary and multi-class classification.
Combines neural networks and STL for multi-class time-series classification.
A new method identifies class-specific covariates in multi-class prediction tasks.
A new algorithm reduces imbalanced data classification errors in multi-class settings.
Symmetrizes loss functions to improve neural network robustness against noisy labels.
Paper proposes MMVFL for multi-class VFL with multiple participants.
New calibration measures for multi-class classification improve model accuracy.
Due to myriads of classes, designing accurate and efficient classifiers becomes very challenging for multi-class classification. Recent research has shown that class structure learning can greatly facilitate multi-class learning. In this paper, we propose a novel method to learn the class structure for multi-class clas…
Conditional generators learn the data distribution for each class in a multi-class scenario and generate samples for a specific class given the right input from the latent space. In this work, a method known as "Versatile Auxiliary Classifier with Generative Adversarial Network" for multi-class scenarios is presented. …
We address the problem of multi-class classification in the case where the number of classes is very large. We propose a double sampling strategy on top of a multi-class to binary reduction strategy, which transforms the original multi-class problem into a binary classification problem over pairs of examples. The aim o…
We study consistency of learning algorithms for a multi-class performance metric that is a non-decomposable function of the confusion matrix of a classifier and cannot be expressed as a sum of losses on individual data points; examples of such performance metrics include the macro F-measure popular in information retri…
Paper tackles cybersecurity attack detection with an ensemble approach.
New method calibrates multi-class predictions efficiently without sacrificing accuracy.
New method learns multi-class from single-class data with confidences.
Improves probability estimates for small datasets in multi-class problems.
The study analyzes multi-class teacher-student perceptron performance and generalization errors.
Develops risk-averse fair multi-class classification methods.
We consider the problem of multi-class classification and a stochastic opti- mization approach to it. We derive risk bounds for stochastic mirror descent algorithm and provide examples of set geometries that make the use of the algorithm efficient in terms of error in k.
Improves ROC/AUC for multi-class classification.
Extends PCVM for multi-class classification with improved accuracy.
Study extends learnability equivalence to multi-class and regression, overcoming binary classification limits.
Proposes CLIQUE for improved local variable importance in multi-class classification.
Domain generalization is the problem of assigning labels to an unlabeled data set, given several similar data sets for which labels have been provided. Despite considerable interest in this problem over the last decade, there has been no theoretical analysis in the setting of multi-class classification. In this work, w…
Multi-class classification with a very large number of classes, or extreme classification, is a challenging problem from both statistical and computational perspectives. Most of the classical approaches to multi-class classification, including one-vs-rest or multi-class support vector machines, require the exact estima…
Recent advances in neuroscience have revealed many principles about neural processing. In particular, many biological systems were found to reconfigure/recruit single neurons to generate multiple kinds of decisions. Such findings have the potential to advance our understanding of the design and optimization process of …
In the last few years, many different performance measures have been introduced to overcome the weakness of the most natural metric, the Accuracy. Among them, Matthews Correlation Coefficient has recently gained popularity among researchers not only in machine learning but also in several application fields such as bio…
Improved algorithms solve multi-period multi-class packing problems with bandit feedback.
OPLTs online train label trees for multi-label and multi-class classification.
The number of possible methods of generalizing binary classification to multi-class classification increases exponentially with the number of class labels. Often, the best method of doing so will be highly problem dependent. Here we present classification software in which the partitioning of multi-class classification…
New algorithms for multi-class classification with abstention.
ETGP improves multi-class classification efficiency.
Machine Learning has become very famous currently which assist in identifying the patterns from the raw data. Technological advancement has led to substantial improvement in Machine Learning which, thus helping to improve prediction. Current Machine Learning models are based on Classical Theory, which can be replaced b…
Recent studies in the literature have paid much attention to the sparsity in linear classification tasks. One motivation of imposing sparsity assumption on the linear discriminant direction is to rule out the noninformative features, making hardly contribution to the classification problem. Most of those work were focu…
StructureBoost improves gradient boosting for complex categorical variables efficiently.
Categorical bundles provide a natural framework for gauge theories involving multiple gauge groups. Unlike the case of traditional bundles there are distinct notions of triviality, and hence also of local triviality, for categorical bundles. We study categorical principal bundles that are product bundles in the categor…
Machine learning (ML) algorithms and machine learning based software systems implicitly or explicitly involve complex flow of information between various entities such as training data, feature space, validation set and results. Understanding the statistical distribution of such information and how they flow from one e…