A deep learning method for XML with autoencoder and ranking loss.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Data collection is a major bottleneck in machine learning and an active research topic in multiple communities. There are largely two reasons data collection has recently become a critical issue. First, as machine learning is becoming more widely-used, we are seeing new applications that do not necessarily have enough …
Deep neural networks (DNNs) trained on large-scale datasets have exhibited significant performance in image classification. Many large-scale datasets are collected from websites, however they tend to contain inaccurate labels that are termed as noisy labels. Training on such noisy labeled datasets causes performance de…
Paper proposes training deep nets on noisy labels without manual annotation.
Extracts geometric information from point-clouds for multiclass classification.
Modern machine learning-based recognition approaches require large-scale datasets with large number of labelled training images. However, such datasets are inherently difficult and costly to collect and annotate. Hence there is a great and growing interest in automatic dataset collection methods that can leverage the w…
This paper considers the problem of inferring image labels from images when only a few annotated examples are available at training time. This setup is often referred to as low-shot learning, where a standard approach is to re-train the last few layers of a convolutional neural network learned on separate classes for w…
Collecting large training datasets, annotated with high-quality labels, is costly and time-consuming. This paper proposes a novel framework for training deep convolutional neural networks from noisy labeled datasets that can be obtained cheaply. The problem is formulated using an undirected graphical model that represe…
Deep neural networks have established as a powerful tool for large scale supervised classification tasks. The state-of-the-art performances of deep neural networks are conditioned to the availability of large number of accurately labeled samples. In practice, collecting large scale accurately labeled datasets is a chal…
Unlike images or videos data which can be easily labeled by human being, sensor data annotation is a time-consuming process. However, traditional methods of human activity recognition require a large amount of such strictly labeled data for training classifiers. In this paper, we present an attention-based convolutiona…
New methods detect targets from imprecisely labeled hyperspectral data.
BERT-based word embeddings improve active learning for text datasets.
Modern computing and communication technologies can make data collection procedures very efficient. However, our ability to analyze large data sets and/or to extract information out from them is hard-pressed to keep up with our capacities for data collection. Among these huge data sets, some of them are not collected f…
Plud system reduces labeling time and produces accurate models for uncategorized images.
Crowdsourcing utilizes the wisdom of crowds for collective classification via information (e.g., labels of an item) provided by labelers. Current crowdsourcing algorithms are mainly unsupervised methods that are unaware of the quality of crowdsourced data. In this paper, we propose a supervised collective classificatio…
Work shows hallucination detection by LLMs is impossible without expert feedback.
Convolutional neural networks (CNNs) have been successfully applied to many recognition and learning tasks using a universal recipe; training a deep model on a very large dataset of supervised examples. However, this approach is rather restrictive in practice since collecting a large set of labeled images is very expen…
New method pools labels from similar data items to improve learning from small samples.
Paper explains why small-loss criterion works for learning from noisy labels.
DALES offers a large annotated aerial LiDAR dataset for 3D deep learning.
Inherent risk scoring is an important function in anti-money laundering, used for determining the riskiness of an individual during onboarding fraudulent transactions occur. It is, however, often fraught with two challenges: (1) inconsistent notions of what constitutes as high or low risk by experts a…
Paper tackles robust training with noisy labels using a trusted set and pseudo labels.
Collecting labeled data is costly and thus a critical bottleneck in real-world classification tasks. To mitigate this problem, we propose a novel setting, namely learning from complementary labels for multi-class classification. A complementary label specifies a class that a pattern does not belong to. Collecting compl…
Collective classification has been intensively studied due to its impact in many important applications, such as web mining, bioinformatics and citation analysis. Collective classification approaches exploit the dependencies of a group of linked objects whose class labels are correlated and need to be predicted simulta…
Self-training avoids spurious features in domain adaptation.
Paper proposes semi-supervised learning for bearing anomaly detection.
X-Transformer improves deep transformer performance on extreme multi-label text classification.
XR-Transformer accelerates XMC by recursively fine-tuning on multi-resolution objectives.
Doubly robust method reduces label cost for noisy crowdsourced data.
A method for collecting human supervision that combines rules and instance labels.
Deep learning fails in classifying handwritten historical documents, traditional methods perform better.
This work tackles missing annotations in large sensor datasets.
Paper proposes a method to adapt classifiers using complementary labels instead of true labels.
Clarinet uses complementary labels to train classifiers with less source data.
Proposes methods to recover labels from shuffled networks using graph averages.
Boost GNNs for node classification by incorporating label dependencies.
Active inference uses machine learning to prioritize data labeling for more efficient statistical inference.
Due to their ubiquitous and pervasive nature, Wi-Fi networks have the potential to collect large-scale, low-cost, and disaggregate data on multimodal transportation. In this study, we develop a semi-supervised deep residual network (ResNet) framework to utilize Wi-Fi communications obtained from smartphones for the pur…
Regression problems are pervasive in real-world applications. Generally a substantial amount of labeled samples are needed to build a regression model with good generalization ability. However, many times it is relatively easy to collect a large number of unlabeled samples, but time-consuming or expensive to label them…
Paper introduces new techniques for real-time sensor data labelling.
New framework improves fairness in small data settings.
Collective learning leverages human collaboration for semi-supervised learning.
Paper uses user engagement signals to automatically label training data for AI assistants.
Proposes a method to create predictive sets from partially labeled data.
We investigate ways in which to improve the interpretability of LDA topic models by better analyzing and visualizing their outputs. We focus on examining what we refer to as topic similarity networks: graphs in which nodes represent latent topics in text collections and links represent similarity among topics. We descr…
Crowdsourcing has become a popular method for collecting labeled training data. However, in many practical scenarios traditional labeling can be difficult for crowdworkers (for example, if the data is high-dimensional or unintuitive, or the labels are continuous). In this work, we develop a novel model for crowdsourcin…
OpinionRank uses graph-based ranking to improve unreliable crowdsourced labels.
Response time improves alignment with diverse human preferences.