Paper learns identity-sensitive word embeddings from text corpora.
problem Lack of context-aware word embeddings.
method Constructs a heterogeneous network of words and identities, then embeds into a low-dimensional space.
result Identity-sensitive word embeddings capture different meanings of words.
This paper assesses biases in contextualized word representations.
problem Analyzing biases in contextualized word representations.
method Proposes assessing bias at the contextual word level, capturing contextual effects of bias.
result Demonstrates evidence of bias in contextual word models, including racial bias and exacerbated effects for intersectional minorities.
We bound the value of the Casson invariant of any integral homology 3-sphere M by a constant times the distance-squared to the identity, measured in any word metric on the Torelli group $\T$, of the element of $\T$ associated to any Heegaard splitting of M. We construct examples which show this bound is asymptotica…
Unsupervised method improves word vectors by suppressing high variance features.
problem Improving semantic information in word vectors.
method Using conceptors to suppress high variance features in word vectors.
result Post-processed word vectors outperform existing alternatives in lexical evaluation tasks.
Study polynomial trace identities in $SL(2,\IC)$ using quaternion algebras.
problem Understanding polynomial trace identities in $SL(2,\IC)$ and their applications.
method Use quaternion algebras over indefinites and their units to study discrete subgroups of $SL(2,\IC)$.
result Obtained structure theorems for quaternion algebras and new polynomial trace identities.
The well-known fact that any genus g symplectic Lefschetz fibration X4→S2 is given by a word that is equal to the identity element in the mapping class group and each of whose elements is given by a positive Dehn twist, provides an intimate relationship between words in the mapping class group and 4-manif…
New post-processing methods improve word embedding performance.
problem Boosting the performance of word embeddings for similarity and analogy tasks.
method Optimizing a semi-Riemannian manifold with Centralised Kernel Alignment (CKA) to shrink the covariance matrix towards a scaled identity matrix.
result Improved performance on downstream tasks after smoothing the spectrum of word vectors.
We show that for any given n, there exists a sequence of words a_k in the generators sigma_1, ... sigma_{n-1} of the braid group B_n, representing the identity element of B_n, such that the number of braid relations of the form sigma_i sigma_{i+1} sigma_i = sigma_{i+1} sigma_i sigma_{i+1} needed to pass from a_k to the…
We show that every co--orientable taut foliation F of an orientable, atoroidal 3-manifold admits a transverse essential lamination. If this transverse lamination is a foliation G, the pair F,G are the unstable and stable foliation respectively of an Anosov flow. Otherwise, F admits a pair of transverse very full genuin…
In this paper we prove that the space of flat metrics (nonpositively curved Euclidean cone metrics) on a closed, oriented surface is marked length spectrally rigid. In other words, two flat metrics assigning the same lengths to all closed curves differ by an isometry isotopic to the identity. The novel proof suggests a…
Locally-contextual CRFs improve sequence labeling performance.
problem Improving sequence labeling with contextual embeddings.
method Locally-contextual nonlinear CRFs using deep neural networks.
result Consistently outperforms linear chain CRF and previous state of the art.
We show that one can skip the skew-symmetry assumption in the definition of Nambu-Poisson brackets. In other words, a n-ary bracket on the algebra of smooth functions which satisfies the Leibniz rule and a n-ary version of the Jacobi identity must be skew-symmetric. A similar result holds for a non-antisymmetric versio…
Unsupervised MT struggles with morphologically rich languages.
problem Limitations of unsupervised machine translation on morphologically rich languages.
method Adversarial unsupervised alignment of word embedding spaces for bilingual dictionary induction.
result A simple trick exploiting weak supervision from identical words improves unsupervised bilingual dictionary induction performance.
We discuss locally simply transitive affine actions of Lie groups G on finite-dimensional vector spaces such that the commutator subgroup [G,G] is acting by translations. In other words, we consider left-symmetric algebras satisfying the identity [x,y].z=0. We derive some basic characterizations of such left-symmetric …
Model captures author language diffusion over time.
problem Lack of author identity and temporal context in language models.
method Temporal language model conditioning on author and temporal vectors.
result Beat temporal and non-temporal baselines, learns time-varying author representations.
The diameter of a disc filling a loop in the universal covering of a Riemannian manifold may be measured extrinsically using the distance function on the ambient space or intrinsically using the induced length metric on the disc. Correspondingly, the diameter of a van Kampen diagram filling a word that represents the i…
Improved neural keyphrase generation by beam search with reward functions.
problem Sequence length bias and beam diversity issues in neural keyphrase generation.
method Beam search decoding strategy with word-level and ngram-level reward functions.
result Significant improvement in generating diverse and accurate keyphrases.
Seq-CVAE learns a latent space for each word position to capture sentence intention.
problem Capturing diversity in image captioning models.
method Seq-CVAE learns a sequential latent space for each word position, mimicking future sentence summaries.
result Significantly improves diversity metrics on MSCOCO dataset compared to baselines.
Mitigates bias in text classification by weighting instances.
problem Unintended biases in text classification datasets based on demographic terms.
method Instance weighting to recover non-discrimination distribution.
result Effective mitigation of unintended biases without sacrificing generalization.
Given a principal G-bundle P→M and two C1 curves in M with coinciding endpoints, we say that the two curves are holonomically equivalent if the parallel transport along them is identical for any smooth connection on P. The main result in this paper is that if G is semi-simple, then the two curves are h…
A new method selects anchor words for better topic discovery in text corpora.
problem Selecting anchor words for improved topic modeling in text corpora.
method Proposes a new greedy method to find a minimum edge-weight anchor clique in a word similarity graph.
result The proposed method outperforms existing methods on topic quality and is faster.
There are certain families of words and word sequences (words in the generators of a two-generator group) that arise frequently in the Teichm{ü}ller theory of hyperbolic three-manifolds and Kleinian and Fuchsian groups and in the discreteness problem for two generator matrix groups. We survey some of the families of su…
Probabilistic FastText captures multiple word senses and sub-word structures.
problem Capturing multiple word senses and sub-word structures in word embeddings.
method Probabilistic FastText uses Gaussian mixture densities to represent words, sharing statistical strength across sub-word structures and capturing different word senses.
result Probabilistic FastText outperforms existing models on word-similarity benchmarks and discerning different meanings.
We prove that two countable locally finite-by-abelian groups G,H endowed with proper left-invariant metrics are coarsely equivalent if and only if their asymptotic dimensions coincide and the groups are either both finitely-generated or both are infinitely generated. On the other hand, we show that each countable group…
FRAGE learns word embeddings without frequency bias, improving performance across NLP tasks.
problem Word embeddings are biased towards word frequency, affecting performance for rare words.
method Adversarial training to learn Frequency-Agnostic word Embedding (FRAGE).
result FRAGE achieves higher performance than baselines in all four NLP tasks.
Geometrically transforms word embeddings into a common space for better comparison.
problem Comparing embeddings from different sources is challenging.
method Applies orthogonal rotations and Mahalanobis scaling to transform embeddings into a shared latent space.
result The method improves word similarity and analogy tasks.
We discuss a topological approach to words introduced by the author. Words on an arbitrary alphabet are approximated by Gauss words and then studied up to natural modifications inspired by the Reidemeister moves on knot diagrams. This leads us to a notion of homotopy for words. We introduce several homotopy invariants …
Random groups prove length constraints on product of conjugates.
problem Quantify products of conjugates in random groups.
method Sharp van Kampen diagram argument and boundary block-counting.
result Prove a sharp inequality for products of conjugates in random groups.
Discriminative model identifies readers and assesses comprehension from eye movements.
problem Inferring readers' identities and estimating their text comprehension from eye movements.
method Generative model of gaze patterns, Fisher-score representation, Fisher-SVM with Fisher kernel.
result SVM with Fisher kernel excels at identifying readers, but not comprehending text.
New word distributions capture multiple meanings and outperform existing methods.
problem Capturing semantic information for words with multiple meanings.
method Gaussian mixtures with an energy-based max-margin objective.
result Multimodal word distributions outperform word2vec and Gaussian embeddings.
Paper proposes a new word embedding method optimizing word similarity.
problem Optimizing word similarity in embedding space.
method Two-step random walks between words via topics to learn an optimal embedding simplex.
result Our method outperforms existing approaches in various queries.
Paper finds better words for topic models by reranking top words.
problem Top words in topic models are not always representative.
method Reranking words by considering marginal probability over every topic.
result Reranked top words are more representative of topics.
The abstract explains how word and relation representations capture semantic meaning.
problem Understanding how word and relation representations capture semantic meaning.
method Theoretical justification and extension of geometric relationships between word embeddings and knowledge graph representations.
result The geometric relationships between word embeddings correspond to semantic relations between words and entities in knowledge graphs.
Approaches KL divergence for learning multi-sense word distributions.
problem Capturing the polysemy and uncertainty of words in word embeddings.
method Modeling words as multi-sense Gaussian mixtures and using KL divergence for learning.
result The proposed approach effectively captures word entailment and distribution similarity.
Paper proposes MMD-Sense-Analysis for detecting word sense shifts.
problem Detecting and interpreting shifts in word meanings over time.
method Leverages Maximum Mean Discrepancy (MMD) to identify and explain word sense changes.
result Demonstrates effectiveness of MMD-Sense-Analysis through empirical results.
Bayesian algorithm improves word representations using semantic taxonomy.
problem Improving word representations in semantic taxonomy.
method Bayesian Hierarchical Words Representation (BHWR) learning algorithm combining Variational Bayes and semantic taxonomy modeling.
result BHWR produces better representations for rare words.
New method improves ASR word confidence for diverse applications.
problem Mitigating ASR errors and improving word error rate.
method Heterogeneous Word Confusion Network (HWCN) with score calibration.
result Word sequence with best overall confidence is more accurate than 1-best result.
Acoustic Neighbor Embeddings map speech and text to fixed dimensions for phonetic confusability.
problem Mapping speech and text to fixed dimensions for phonetic confusability.
method Adapting SNE to sequential inputs, training two encoder neural networks.
result More accurate results with low-dimensional embeddings in word recognition tasks.
Word2vec improved but lacks multi-meaning words; ConEc creates new embeddings.
problem Lack of meaningful embeddings for words with multiple meanings and OOV words.
method Context encoders (ConEc) extend word2vec by multiplying embeddings with context vectors.
result ConEc creates embeddings for OOV words and words with multiple meanings based on local contexts.
New method uses word subspaces and term-frequency to improve text classification.
problem Lack of semantic meaning in bag-of-words features.
method Proposes word subspaces and term-frequency weighted word subspaces for text classification.
result Improved text classification performance compared to state-of-the-art algorithms.
Recent work on learning multilingual word representations usually relies on the use of word-level alignements (e.g. infered with the help of GIZA++) between translated sentences, in order to align the word embeddings in different languages. In this workshop paper, we investigate an autoencoder model for learning multil…
Proposes MorphMine for unsupervised morpheme segmentation to improve word embeddings.
problem Lack of semantic information in word-level analysis for infrequent and out-of-vocabulary words.
method MorphMine applies a parsimony criterion to hierarchically segment words into the fewest number of morphemes.
result MorphMine segments words into human-verified morphemes and improves word embedding quality.
Estimator Vectors learns OOV word embeddings using subword and context clues.
problem Lack of OOV word representations in neural network models.
method Jointly learns word, subword, and context clue representations.
result Strong estimates for OOV words via combined subword and context clue embeddings.
Cognitive psychology insights reveal shape bias in DNNs.
problem Interpreting modern deep neural networks (DNNs).
method Applied developmental psychology's word learning theory to DNNs.
result DNNs exhibit a shape bias similar to human word learning.
End-to-end ASR model combines word and character representation for improved performance.
problem Difficulty in training with word-level supervision due to sparsity of examples.
method Multi-task learning framework combining word and character representations.
result Improved word-error rate (WER) by interpolating between word-level and character-level models.
Paper analyzes word embedding composition using tensor decomposition.
problem Given vector representations of two words, compute a vector for the entire phrase.
method Generative model with low rank Tucker decomposition of word embedding correlations.
result Word embeddings and a core tensor can be derived from the Tucker decomposition.
SWESA learns word embeddings with document labels for sentiment analysis.
problem Sentiment analysis using limited text data.
method SWESA uses supervised learning to optimize word embeddings and classifier performance.
result SWESA outperforms existing methods in sentiment analysis.
New concept SB-generation helps classify transformation groups.
problem Classifying transformation groups through quasi-isometry invariants.
method Identifying SB-generated groups in specific transformation groups.
result SB-generation provides robust extension of finite generation.