Binary PheNorm extends phenotype labeling for EHRs using binary silver labels.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Proposes MLPSVM for multi-label learning, improving on binary relevance.
Proposes top-label calibration and M2B framework for multiclass to binary calibration.
Recently the deep learning techniques have achieved success in multi-label classification due to its automatic representation learning ability and the end-to-end learning framework. Existing deep neural networks in multi-label classification can be divided into two kinds: binary relevance neural network (BRNN) and thre…
Binary classification improves with a small fraction of corrupted labels.
This work analyzes two methods for combining multiple binary labels in bipartite ranking.
Paper proposes Pcomp classification for binary classification with pairwise confidence comparisons.
Efficiently poisons offline RLHF models by flipping preference labels.
Paper estimates FPR of Bayes classifier using soft labels.
Sharp bounds on uniform generalization errors in binary linear classification.
Ridge regression shows different behaviors in binary classification with noisy labels.
New research shows that binary classification can be done with noisy data, but only if there are clean samples available.
Data labeling is currently a time-consuming task that often requires expert knowledge. In research settings, the availability of correctly labeled data is crucial to ensure that model predictions are accurate and useful. We propose relatively simple machine learning-based models that achieve high performance metrics in…
Tests for classifier independence without ground truth labels.
Class imbalance is an intrinsic characteristic of multi-label data. Most of the labels in multi-label data sets are associated with a small number of training examples, much smaller compared to the size of the data set. Class imbalance poses a key challenge that plagues most multi-label learning methods. Ensemble of Cl…
Bird sound data collected with unattended microphones for automatic surveys, or mobile devices for citizen science, typically contain multiple simultaneously vocalizing birds of different species. However, few works have considered the multi-label structure in birdsong. We propose to use an ensemble of classifier chain…
We study the natural map eta between a group of binary planar trees whose leaves are labeled by elements of a free abelian group H and a certain group D(H) derived from the free Lie algebra over H. Both of these groups arise in several different topological contexts. The map eta is known to be an isomorphism over Q, bu…
New method improves LLM judge accuracy by accounting for dependencies in aggregated binary labels.
Multi-label classification studies the task where each example belongs to multiple labels simultaneously. As a representative method, Ranking Support Vector Machine (Rank-SVM) aims to minimize the Ranking Loss and can also mitigate the negative influence of the class-imbalance issue. However, due to its stacking-style …
Convolutional neural network (CNN)-based feature learning has become state of the art, since given sufficient training data, CNN can significantly outperform traditional methods for various classification tasks. However, feature learning becomes more difficult if some training labels are noisy. With traditional regular…
Proposes CEP to better represent financial products' carbon impact.
The paper studies how to use AI-generated labels in econometrics to avoid bias.
In many applications of classifier learning, training data suffers from label noise. Deep networks are learned using huge training data where the problem of noisy labels is particularly relevant. The current techniques proposed for learning deep networks under label noise focus on modifying the network architecture and…
This paper develops convex surrogates for optimizing the multi-label F-measure.
We consider a query-based data acquisition problem for binary classification of unknown labels, which has diverse applications in communications, crowdsourcing, recommender systems and active learning. To ensure reliable recovery of unknown labels with as few number of queries as possible, we consider an effective quer…
The paper presents Imbalance-XGBoost, a Python package that combines the powerful XGBoost software with weighted and focal losses to tackle binary label-imbalanced classification tasks. Though a small-scale program in terms of size, the package is, to the best of the authors' knowledge, the first of its kind which prov…
We propose a new problem formulation which is similar to, but more informative than, the binary multiple-instance learning problem. In this setting, we are given groups of instances (described by feature vectors) along with estimates of the fraction of positively-labeled instances per group. The task is to learn an ins…
New method calibrates multi-class predictions efficiently without sacrificing accuracy.
We consider sequential decision making problems for binary classification scenario in which the learner takes an active role in repeatedly selecting samples from the action pool and receives the binary label of the selected alternatives. Our problem is motivated by applications where observations are time consuming and…
Quantum circuits represent binary classification trees with binary features.
This paper investigates the problem of active learning for binary label prediction on a graph. We introduce a simple and label-efficient algorithm called S2 for this task. At each step, S2 selects the vertex to be labeled based on the structure of the graph and all previously gathered labels. Specifically, S2 queries f…
This paper investigates semi-supervised hashing methods using variational autoencoders.
New loss functions improve extreme classification with missing labels.
Simplifies multi-label classification with stochastic sketch strategy.
The number of possible methods of generalizing binary classification to multi-class classification increases exponentially with the number of class labels. Often, the best method of doing so will be highly problem dependent. Here we present classification software in which the partitioning of multi-class classification…
New method estimates hidden binary mixture model centers efficiently.
In this paper, we investigate the degree to which the encoding of a -VAE captures label information across multiple architectures on Binary Static MNIST and Omniglot. Even though they are trained in a completely unsupervised manner, we demonstrate that a -VAE can retain a large amount of label information, even w…
New methods lift weak supervision to structured prediction, providing robustness guarantees.
We propose an adversarial training procedure for learning a causal implicit generative model for a given causal graph. We show that adversarial training can be used to learn a generative model with true observational and interventional distributions if the generator architecture is consistent with the given causal grap…
New method corrects skewed confidence for PbN classification.
The paper develops bounds for predictive values in binary classification.
Estimates binary labels from dependent data using Markov Random Fields.
In this paper, we propose a compositional nonparametric method in which a model is expressed as a labeled binary tree of nodes, where each node is either a summation, a multiplication, or the application of one of the basis functions to one of the covariates. We show that in order to recover a labeled bi…
Develops NPMC method for noisy labels, improving multiclass classification accuracy.
The paper explores how machine learning models can be learnable despite label shifts.
Paper proposes methods to improve SVM classifiers in noisy data scenarios.
Unified framework designs LK structures using integer twists on non-manifold meshes.
We propose the Autoencoding Binary Classifiers (ABC), a novel supervised anomaly detector based on the Autoencoder (AE). There are two main approaches in anomaly detection: supervised and unsupervised. The supervised approach accurately detects the known anomalies included in training data, but it cannot detect the unk…