Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

225449674898 · Jun 202019922001200920182026
48 results for neural word embeddings

Unsupervised model learns word and context embeddings from character sequences.

problem Learning meaningful word and context embeddings from unlabeled data.
method Character-aware neural architecture that jointly learns word and context embeddings.
result Compact encoders achieve high performance in downstream tasks.

This paper reviews different word embeddings for sentiment classification using deep learning.

problem Handling large textual data with simple ML algorithms.
method Word embedding strategies implemented on an Amazon Review Dataset.
result Different word embeddings improve accuracy in sentiment classification.

Paper distills word embeddings to reduce dimensionality without sacrificing accuracy.

problem Reducing neural model size for practical deployment.
method Teacher-student model-based embedding distillation with ensemble learning.
result Significant reduction in model size (80x faster and lighter) with minimal accuracy loss.

New insights into word embeddings reveal linear relationships behind analogy phenomena.

problem Understanding the linear behavior of word embeddings in analogy tasks.
method Derive a probabilistic definition of paraphrasing, interpret as word transformation, and prove linear relationships.
result Existence of linear relationships between W2V-type embeddings underpinning analogy phenomena.

Enhanced word embeddings boost multiclass text classification accuracy.

problem Improving multiclass text classification accuracy using pre-trained embeddings.
method Proposed word-class embeddings (WCEs) to enhance pre-trained word embeddings.
result WCEs significantly improve multiclass text classification accuracy.

Neural model improves text normalization for non-English languages.

problem Improving text normalization in non-English languages with limited data.
method Sequence-to-sequence model with character and word embeddings, using pre-trained word embeddings with subword information.
result Achieved state-of-the-art F1 score on Arabic language correction dataset.

CNNs and biomedical word embeddings detect ADRs in literature.

problem Detecting Adverse Drug Reactions (ADRs) in biomedical literature is time-consuming.
method Used token-level convolutions and biomedical word embeddings in CNN architectures.
result CNNs outperform traditional models and word embeddings improve performance.

With a simple architecture and the ability to learn meaningful word embeddings efficiently from texts containing billions of words, word2vec remains one of the most popular neural language models used today. However, as only a single embedding is learned for every word in the vocabulary, the model fails to optimally re…

2017-06-08abs ↗pdf ↗

A new neural network model extends word embedding vectors with MeSH concepts for biomedical semantic similarity.

problem Eliciting semantic similarity between biomedical concepts remains challenging.
method Proposes a MeSH-gram neural network model that extends skip-gram by using MeSH descriptors.
result MeSH-gram outperforms skip-gram and is comparable to best methods but requires more computation and external resources.

A new neural network for text classification reduces parameters with improved accuracy.

problem Reducing the number of parameters in text classification models.
method Compositional coding, capsule network, k-means routing algorithm.
result The proposed method achieves competitive accuracy with significantly fewer parameters.

Proposes a new feature-based evaluation method for explaining Deep Learning models in text classification.

problem Lack of consideration for linguistic dependencies in existing attribution-based explanations.
method Investigates perturbations based on embedded features removal from intermediate layers of Convolutional Neural Networks.
result Visualization tool assists analysts in understanding model predictions better.

Paper uses neural word embeddings to analyze UN speeches for policy preferences and voting behavior.

problem Analyzing policy preferences and paradigm shifts in international politics.
method Applied neural word embeddings (Word2vec) to UN General Debate speeches.
result Found statistical relation between speech semantic content and voting behavior, contrary to hypothesis.

Paper develops neural network for Mandarin polyphone disambiguation.

problem Homograph problem in Mandarin Chinese text-to-speech.
method Bidirectional RNN for context, prediction network for mapping embeddings to pronunciations.
result Achieves 94.69% accuracy on polyphonic character dataset.

Most existing word embedding approaches do not distinguish the same words in different contexts, therefore ignoring their contextual meanings. As a result, the learned embeddings of these words are usually a mixture of multiple meanings. In this paper, we acknowledge multiple identities of the same word in different co…

2016-11-29abs ↗pdf ↗

Neural Machine Translation (MT) has reached state-of-the-art results. However, one of the main challenges that neural MT still faces is dealing with very large vocabularies and morphologically rich languages. In this paper, we propose a neural MT system using character-based embeddings in combination with convolutional…

2016-03-02abs ↗pdf ↗

Word embeddings are a powerful approach for capturing semantic similarity among terms in a vocabulary. In this paper, we develop exponential family embeddings, a class of methods that extends the idea of word embeddings to other types of high-dimensional data. As examples, we studied neural data with real-valued observ…

2016-08-02abs ↗pdf ↗

BERT-based word embeddings improve active learning for text datasets.

problem Efficiently labelling large text datasets for machine learning.
method Evaluation of text representation mechanisms (BERT vs. bag of words) in active learning.
result BERT-based word embeddings significantly improve active learning performance.

Paper uses JIVE to decompose word embeddings, improving sentiment analysis performance.

problem Improving sentiment analysis performance on word embeddings.
method Joint and individual variance explained (JIVE) method for decomposition.
result Mapping word embeddings into joint components improves sentiment analysis performance.

New methods for unsupervised learning of word and entity representations.

problem Learning distributed representations of words and entities from text and knowledge bases.
method MVLSA for words and NVSE for entities, both unsupervised learning methods.
result MVLSA and NVSE outperform state-of-the-art models in word and entity representation learning.

This paper tackles rare word problem in low-resource language pairs using NMT.

problem Rare word problem in neural machine translation, especially for low-resource languages.
method Three solutions: enhanced source context, morphology learning, and wordnet synonyms.
result Significant improvements in BLEU scores (+1.0 points) on English-Vietnamese and Japanese-Vietnamese.

WME generates document embeddings from word embeddings, outperforming state-of-the-art techniques.

problem Lack of unsupervised document embeddings from pre-trained word embeddings.
method Word Mover's Embedding (WME) approach.
result WME consistently matches or outperforms state-of-the-art techniques on various text classification and similarity tasks.

Paper analyzes word embedding composition using tensor decomposition.

problem Given vector representations of two words, compute a vector for the entire phrase.
method Generative model with low rank Tucker decomposition of word embedding correlations.
result Word embeddings and a core tensor can be derived from the Tucker decomposition.

Proposes MorphMine for unsupervised morpheme segmentation to improve word embeddings.

problem Lack of semantic information in word-level analysis for infrequent and out-of-vocabulary words.
method MorphMine applies a parsimony criterion to hierarchically segment words into the fewest number of morphemes.
result MorphMine segments words into human-verified morphemes and improves word embedding quality.

Improved NER on Turkish tweets using semi-supervised learning and word embeddings.

problem Named Entity Recognition on informal Turkish text types.
method Semi-supervised learning with neural networks and word embeddings.
result Achieved better F-score performances than previous Turkish NER systems.

GroupReduce compresses neural language models by reducing embedding and softmax matrices.

problem Large vocabulary size in neural language models leads to high model size and memory usage.
method Vocabulary-partition based low-rank matrix approximation and frequency distribution of tokens.
result Significantly outperforms traditional compression methods with 6.6 times compression rate for embedding and softmax matrices.

Word embeddings in hyperbolic space outperform Euclidean ones.

problem Improving word embeddings for better performance.
method Learning word embeddings in hyperbolic space using skip-gram architecture and hyperbolic distance objective function.
result Hyperbolic word embeddings show potential, especially in low dimensions, but not clear superiority over Euclidean embeddings.

Unified framework for word embedding models using noise examples.

problem Improving word embedding models with negative sampling.
method Formulated a Word-Context Classification (WCC) framework that generalizes SkipGram word embedding models.
result The best noise distribution is the data distribution, improving both performance and training speed.