Improves cross-lingual NER by projecting entities from one language to another.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study uncovers financial trends from cross-lingual news data.
\emph{Sentiment Quantification} (i.e., the task of estimating the relative frequency of sentiment-related classes -- such as \textsf{Positive} and \textsf{Negative} -- in a set of unlabelled documents) is an important topic in sentiment analysis, as the study of sentiment-related quantities and trends across a populati…
RoBERTa outperforms other pre-trained models in NER tasks.
The authors advocate for more rigorous unsupervised cross-lingual learning methods.
Paper explores zero-shot cross-lingual reading comprehension using pre-trained multi-lingual model.
A new method sorts models to find the best one with minimal risk.
Structural correspondence learning (SCL) is an effective method for cross-lingual sentiment classification. This approach uses unlabeled documents along with a word translation oracle to automatically induce task specific, cross-lingual correspondences. It transfers knowledge through identifying important features, i.e…
Paper develops a new unsupervised scoring function for cross-lingual document alignment.
FGN improves Chinese NER by integrating glyph information and interactive context.
CCAligned creates a massive web document dataset for cross-lingual research.
Develops a cross-lingual hate speech detection model using pre-trained Transformers.
Deep learning model improves Vietnamese NER accuracy.
BERT model improves cross-lingual document retrieval.
Paper reproduces and enhances a method for cross-lingual word embeddings.
Paper refines cross-lingual word embeddings using Manhattan norm.
Study shows partially-typed NER datasets can match fully-typed ones in model performance.
LPL optimizes embeddings to align local neighborhoods, improving cross-lingual word alignment.
Paper fine-tunes LLaMA-3-8B for financial NER using instruction and LoRA.
Paper presents a method to train NER models without labelled data using weak supervision.
The paper introduces uncertainty quantification for NER models.
NMIXX fine-tunes embeddings for finance, outperforming general models in Korean.
Recently, due to the increasing popularity of social media, the necessity for extracting information from informal text types, such as microblog texts, has gained significant attention. In this study, we focused on the Named Entity Recognition (NER) problem on informal text types for Turkish. We utilized a semi-supervi…
We introduce BilBOWA (Bilingual Bag-of-Words without Alignments), a simple and computationally-efficient model for learning bilingual distributed representations of words which can scale to large monolingual datasets and does not require word-aligned parallel training data. Instead it trains directly on monolingual dat…
State-of-the-art named entity recognition (NER) systems have been improving continuously using neural architectures over the past several years. However, many tasks including NER require large sets of annotated data to achieve such performance. In particular, we focus on NER from clinical notes, which is one of the mos…
Recent approaches based on artificial neural networks (ANNs) have shown promising results for named-entity recognition (NER). In order to achieve high performances, ANNs need to be trained on a large labeled dataset. However, labels might be difficult to obtain for the dataset on which the user wants to perform NER: la…
We present a weakly-supervised data augmentation approach to improve Named Entity Recognition (NER) in a challenging domain: extracting biomedical entities (e.g., proteins) from the scientific literature. First, we train a neural NER (NNER) model over a small seed of fully-labeled examples. Second, we use a reference s…
We solve a key problem in cross-lingual learning using a novel approach.
Improved NER performance on imbalanced data.
Simple framework decouples word alignment and multilingual embedding mapping.
Proposes a method to improve deep active learning for NER tasks.
This paper describes an approach for automatic construction of dictionaries for Named Entity Recognition (NER) using large amounts of unlabeled data and a few seed examples. We use Canonical Correlation Analysis (CCA) to obtain lower dimensional embeddings (representations) for candidate phrases and classify these phra…
Many information retrieval algorithms rely on the notion of a good distance that allows to efficiently compare objects of different nature. Recently, a new promising metric called Word Mover's Distance was proposed to measure the divergence between text passages. In this paper, we demonstrate that this metric can be ex…
Named-entity recognition (NER) aims at identifying entities of interest in a text. Artificial neural networks (ANNs) have recently been shown to outperform existing NER systems. However, ANNs remain challenging to use for non-expert users. In this paper, we present NeuroNER, an easy-to-use named-entity recognition tool…
Introduces PELP for graph-enhanced word embeddings.
Motivated by the need to automate medical information extraction from free-text radiological reports, we present a bi-directional long short-term memory (BiLSTM) neural network architecture for modelling radiological language. The model has been used to address two NLP tasks: medical named-entity recognition (NER) and …
NERS improves RL by sampling diverse transitions considering local and global contexts.
State-of-the-art sequence labeling systems traditionally require large amounts of task-specific knowledge in the form of hand-crafted features and data pre-processing. In this paper, we introduce a novel neutral network architecture that benefits from both word- and character-level representations automatically, by usi…
This paper introduces PyDCI, a new implementation of Distributional Correspondence Indexing (DCI) written in Python. DCI is a transfer learning method for cross-domain and cross-lingual text classification for which we had provided an implementation (here called JaDCI) built on top of JaTeCS, a Java framework for text …
Study quantifies gender bias in language models across 7 languages.
MedCAT extracts valuable medical information from unstructured text.
Cross-lingual Text Classification (CLC) consists of automatically classifying, according to a common set C of classes, documents each written in one of a set of languages L, and doing so more accurately than when naively classifying each document via its corresponding language-specific classifier. In order to obtain an…
A framework for multi-label sentiment analysis in 100 languages with dynamic weighting.
GTI network learns linguistic features for multi-task sequence tagging.
The work investigates deep generative models, which allow us to use training data from one domain to build a model for another domain. We propose the Variational Bi-domain Triplet Autoencoder (VBTA) that learns a joint distribution of objects from different domains. We extend the VBTAs objective function by the relativ…
Re-speaking is a mechanism for obtaining high quality subtitles for use in live broadcast and other public events. Because it relies on humans performing the actual re-speaking, the task of estimating the quality of the results is non-trivial. Most organisations rely on humans to perform the actual quality assessment, …
For many natural language processing (NLP) tasks the amount of annotated data is limited. This urges a need to apply semi-supervised learning techniques, such as transfer learning or meta-learning. In this work we tackle Named Entity Recognition (NER) task using Prototypical Network - a metric learning technique. It le…
Hybrid AI and rule-based framework de-identifies medical imaging data.