Paper proposes MMD-Sense-Analysis for detecting word sense shifts.
problem Detecting and interpreting shifts in word meanings over time.
method Leverages Maximum Mean Discrepancy (MMD) to identify and explain word sense changes.
result Demonstrates effectiveness of MMD-Sense-Analysis through empirical results.
Word embeddings reveal multiple senses, which can be recovered using sparse coding.
problem Understanding word senses in polysemous words.
method Sparse coding of word embeddings to recover multiple senses.
result Sparse coding can approximate multiple word senses, with each sense associated with a discourse atom.
Probabilistic FastText captures multiple word senses and sub-word structures.
problem Capturing multiple word senses and sub-word structures in word embeddings.
method Probabilistic FastText uses Gaussian mixture densities to represent words, sharing statistical strength across sub-word structures and capturing different word senses.
result Probabilistic FastText outperforms existing models on word-similarity benchmarks and discerning different meanings.
Model learns multi-sense word embeddings using bilingual and monolingual data.
problem Learning multi-sense word embeddings with crosslingual information.
method Discrete autoencoder with encoder and decoder components.
result Bilingual word representations outperform monolingual ones across tasks.
We describe our language-independent unsupervised word sense induction system. This system only uses topic features to cluster different word senses in their global context topic space. Using unlabeled data, this system trains a latent Dirichlet allocation (LDA) topic model then uses it to infer the topics distribution…
Approaches KL divergence for learning multi-sense word distributions.
problem Capturing the polysemy and uncertainty of words in word embeddings.
method Modeling words as multi-sense Gaussian mixtures and using KL divergence for learning.
result The proposed approach effectively captures word entailment and distribution similarity.
Proposes methods to model polysemous words using vector representations and geometric analysis.
problem Modeling words with multiple meanings in NLP.
method Three-fold approach: context representations, sense induction, lexeme representations.
result Disambiguation of word senses using low-rank subspaces and Grassmannian geometry.
A single BLSTM network tackles ambiguous words in text data.
problem Ambiguity in text data, especially in technical domains.
method Proposes a single Bidirectional LSTM network for all ambiguous words.
result Comparable performance to top WSD algorithms on SensEval-3 benchmark.
Humor in word embeddings reveals patterns across different groups.
problem Understanding humor in natural language processing.
method Analyzed humor ratings and word embeddings to identify humor features.
result Different demographic groups have distinct humor preferences.
There is rising interest in vector-space word embeddings and their use in NLP, especially given recent methods for their fast estimation at very large scale. Nearly all this work, however, assumes a single vector per word type ignoring polysemy and thus jeopardizing their usefulness for downstream tasks. We present an …
Paper proposes FOFE for efficient WSD.
problem Word sense disambiguation (WSD) problem.
method Fixed-size ordinally forgetting encoding (FOFE) combined with FFNN.
result FOFE-based FFNN achieves comparable performance to state-of-the-art at lower cost.
DKPCA improves WSD accuracy with scarce labeled data.
problem Word sense disambiguation in natural language processing.
method DKPCA combines Kernel PCA and Semantic Diffusion Kernel.
result DKPCA outperforms SVM and KPCA on SensEval data.
Visualizes text models with in-text and word-as-pixel highlighting.
problem Diagnosing and understanding text models, especially for slang and historical texts.
method In-text annotations and word-as-pixel graphics.
result Interconnected methods help diagnose and understand text models.
3-manifold groups' word problem solved in nearly linear time.
problem Solving the word problem for 3-manifold groups efficiently.
method Proof for admissible graphs of groups, leveraging Croke and Kleiner's work.
result Word problem solved in O(nlogn) time for 3-manifold groups. Enhanced word embeddings boost multiclass text classification accuracy.
problem Improving multiclass text classification accuracy using pre-trained embeddings.
method Proposed word-class embeddings (WCEs) to enhance pre-trained word embeddings.
result WCEs significantly improve multiclass text classification accuracy.
Improved embeddings by topic-sensitive attention on large corpora.
problem Capturing sense of words in limited corpora using pretrained embeddings.
method Topic-sensitive attention on large topic-rich corpora to correct sense drift in pretrained embeddings.
result Limited corpus augmentation is more effective than adapting pretrained embeddings.
Improved SGNS model resolves ambiguity in word vector learning.
problem Ambiguity in word vectors under SGNS model.
method Rectified SGNS model with quadratic regularization.
result Simple modification structures the solution effectively.
Let S be a closed surface of genus at least 2. We show that a finitely generated group G which is an extension of the fundamental group H of S is word hyperbolic if and only the orbit map of the quotient group G/H on the complex of curves is a quasi-isometric embedding.This in turn is equivalent to G/H being convex coc…
Direct acoustics-to-word models improve speech recognition without LMs.
problem Improving speech recognition without requiring a Language Model (LM).
method Direct acoustics-to-word CTC models trained on public benchmark tasks.
result CTC word model achieves 13.0%/18.8% word error rate compared to 9.6%/16.0% for phone-based CTC with a 4-gram LM.
Fruit fly brain network learns word embeddings using sparse binary codes.
problem Learning semantic word representations from text.
method Inspired by mushroom body neural network, sparse binary hash codes.
result Fruit fly network achieves comparable NLP performance with reduced resources.
Corpus poisoning can manipulate word meanings in word embeddings, affecting natural language processing tasks.
problem Controlling word meanings via corpus modifications.
method Developed an explicit expression over corpus features to control word embeddings.
result Demonstrated the ability to manipulate word meanings in word embeddings, affecting various downstream tasks.
The paper uses semantic analysis of economic news to forecast stock prices.
problem Forecasting stock prices using financial news.
method Semantic analysis of news, neural networks, MATLAB Simulink.
result Highly adequate financial time series forecasting models were developed.
Enhances topic models to better handle polysemous words.
problem Lack of polysemy handling in Gaussian latent Dirichlet allocation.
method Introduces a hierarchical structure to capture polysemy in Gaussian latent Dirichlet allocation.
result Significantly improves polysemy detection and provides more parsimonious topic representations.
GASC models semantic change in Ancient Greek texts using genre metadata.
problem Associating correct meanings in historical Ancient Greek texts.
method Develops a dynamic semantic change model leveraging genre metadata.
result Improves predictive performance on semantic change in Ancient Greek texts.
Study convex cocompact subgroups in real projective geometry.
problem Characterize discrete subgroups acting on real projective space.
method Define and characterize convex cocompactness, extend results from orthogonal groups.
result Equivalence of different convex cocompactness conditions for word hyperbolic groups.
DeepCodec learns to take undersampled measurements and recover signals using deep neural networks.
problem Signal recovery from undersampled data.
method Adaptive deep convolutional neural networks for sensing and recovery.
result DeepCodec outperforms traditional ℓ1-minimization in signal recovery. Models learn spatial templates from implicit language, predicting spatial arrangements with high accuracy.
problem Predicting spatial arrangements from implicit spatial language.
method Simple neural-based models leveraging annotated images and structured text.
result Models can predict spatial arrangements from implicit spatial language with high accuracy, even for unseen objects.
BERT captures linguistic features in separate semantic and syntactic subspaces.
problem Understanding how transformer models like BERT represent linguistic features internally.
method Qualitative and quantitative investigations of BERT's internal representations.
result Evidence of a fine-grained geometric representation of word senses and syntactic representations.
We study the restless bandit associated with an extremely simple scalar Kalman filter model in discrete time. Under certain assumptions, we prove that the problem is indexable in the sense that the Whittle index is a non-decreasing function of the relevant belief state. In spite of the long history of this problem, thi…
Team QCRI-MIT detects hyperpartisan news with 72.9% accuracy.
problem Detecting hyperpartisan news from biased political content.
method Logistic regression model using engineered features from propaganda detection.
result Significant performance improvements with better feature pre-processing.
Let G be a word hyperbolic group in the sense of Gromov and P its associated Rips complex. We prove that the fixed point set PH is contractible for every finite subgroups H of G. This is the main ingredient for proving that P is a finite model for the universal space e.g. of proper actions. As a corollar…
A method extracts binary features directly from CS measurements for compressive image classification.
problem Efficiently classify images using compressive sensing without reconstruction.
method DCT-based approach for binary feature extraction from CS measurements, feature fusion with CNN features.
result Fused features outperform state-of-the-art methods in image classification.
This paper is concerned with detecting when a closed braid and its axis are 'mutually braided' in the sense of Rudolph. It deals with closed braids which are fibred links, the simplest case being closed braids which present the unknot. The geometric condition for mutual braiding refers to the existence of a close contr…
Study shows no new Euclidean factors can appear in the limit of CAT(0) spaces.
problem Stability of Euclidean factors in CAT(0) spaces under convergence.
method GH-convergence of CAT(0) spaces with uniformly cocompact discrete groups of isometries.
result Dimension of the maximal Euclidean factor is the same for large j. In this paper we study geometric coincidence problems in the spirit of the following problems by B. Grünbaum: How many affine diameters of a convex body in Rn must have a common point? How many centers (in some sense) of hyperplane sections of a convex body in Rn must coincide? One possible approa…
We prove that the mirror map is trivial for the canonical formal families of Calabi-Yau varieties constructed by Gross and the second author. In other words, the natural coordinate in a canonical Calabi-Yau family is a canonical coordinate in the sense of Hodge theory. This implies that the higher weight periods direct…
The paper explores linear representations in language models using counterfactuals.
problem Understanding linear representations and geometric concepts in large language models.
method Formalized linear representation in output and input spaces, identified causal inner product.
result Unified understanding of linear representations and their connection to interpretation and control.
Paper learns identity-sensitive word embeddings from text corpora.
problem Lack of context-aware word embeddings.
method Constructs a heterogeneous network of words and identities, then embeds into a low-dimensional space.
result Identity-sensitive word embeddings capture different meanings of words.
A new method selects anchor words for better topic discovery in text corpora.
problem Selecting anchor words for improved topic modeling in text corpora.
method Proposes a new greedy method to find a minimum edge-weight anchor clique in a word similarity graph.
result The proposed method outperforms existing methods on topic quality and is faster.
CLDA improves topic modeling for large, dynamic datasets.
problem Dynamic topic modeling for large, diverse text streams.
method Data decomposition followed by topic modeling on segments, then clustering.
result Very fast runtime and insight into topic composition over time.
There are certain families of words and word sequences (words in the generators of a two-generator group) that arise frequently in the Teichm{ü}ller theory of hyperbolic three-manifolds and Kleinian and Fuchsian groups and in the discreteness problem for two generator matrix groups. We survey some of the families of su…
FRAGE learns word embeddings without frequency bias, improving performance across NLP tasks.
problem Word embeddings are biased towards word frequency, affecting performance for rare words.
method Adversarial training to learn Frequency-Agnostic word Embedding (FRAGE).
result FRAGE achieves higher performance than baselines in all four NLP tasks.
Geometrically transforms word embeddings into a common space for better comparison.
problem Comparing embeddings from different sources is challenging.
method Applies orthogonal rotations and Mahalanobis scaling to transform embeddings into a shared latent space.
result The method improves word similarity and analogy tasks.
We discuss a topological approach to words introduced by the author. Words on an arbitrary alphabet are approximated by Gauss words and then studied up to natural modifications inspired by the Reidemeister moves on knot diagrams. This leads us to a notion of homotopy for words. We introduce several homotopy invariants …
New word distributions capture multiple meanings and outperform existing methods.
problem Capturing semantic information for words with multiple meanings.
method Gaussian mixtures with an energy-based max-margin objective.
result Multimodal word distributions outperform word2vec and Gaussian embeddings.
Paper proposes a new word embedding method optimizing word similarity.
problem Optimizing word similarity in embedding space.
method Two-step random walks between words via topics to learn an optimal embedding simplex.
result Our method outperforms existing approaches in various queries.
The abstract explains how word and relation representations capture semantic meaning.
problem Understanding how word and relation representations capture semantic meaning.
method Theoretical justification and extension of geometric relationships between word embeddings and knowledge graph representations.
result The geometric relationships between word embeddings correspond to semantic relations between words and entities in knowledge graphs.
Paper finds better words for topic models by reranking top words.
problem Top words in topic models are not always representative.
method Reranking words by considering marginal probability over every topic.
result Reranked top words are more representative of topics.