Low-dimensional vectors improve semantic understanding of music and language.
problem Noise in shared semantics due to individual brain biases.
method Jointly model multiple brains to learn low-dimensional vector embeddings.
result These embeddings outperform high-dimensional fMRI data in music and language classification.
The paper evaluates different vector space models for text similarity.
problem Measuring semantic text similarity in natural language processing.
method Comparison of TFIDF, topic models, and neural models for patent-to-patent similarity.
result TFIDF performs well for longer, technical texts or finer distinctions.
CSTEM models document topics using VAE with semantic distance.
problem Inability of previous topic models to explain semantic relations correctly.
method Continuous semantic topic embedding model using variational autoencoder and Mahalanobis distance.
result Improves topic coherence and semantic relation explanation.
Based on the Aristotelian concept of potentiality vs. actuality allowing for the study of energy and dynamics in language, we propose a field approach to lexical analysis. Falling back on the distributional hypothesis to statistically model word meaning, we used evolving fields as a metaphor to express time-dependent c…
SOM-VQ tokenizes discrete models with semantic structure and navigable topology.
problem Lack of semantic structure in vector quantized representations limits interpretable human control.
method Combines vector quantization with Self-Organizing Maps to learn discrete codebooks with explicit topology.
result SOM-VQ produces more learnable token sequences and provides an explicit navigable geometry in code space.
Paper explores how adding high-dimensional vectors can memorize and solve set membership problems.
problem Set membership problem in high-dimensional vector spaces.
method Utilizes the almost orthogonal property of high-dimensional random vectors to add them efficiently.
result Efficient probabilistic solution to set membership problem.
Enhances word embedding by transferring external knowledge.
problem Low-frequency words in semantic space.
method Latent Semantic Imputation (LSI) integrating graph theory and spectral embeddings.
result LSI generates reliable embedding vectors for low-frequency words.
Top2Vec finds topic vectors from documents and words without needing stop words or custom settings.
problem Topic modeling weaknesses, including needing known topics, stop words, and custom settings.
method Joint document and word semantic embedding to find topic vectors automatically.
result Top2Vec finds more informative and representative topics than probabilistic models.
Word2Vec captures musical relationships in complex polyphonic music.
problem Capturing meaningful relationships in musical contexts.
method Skip-gram version of word2vec applied to music slices from a large corpus.
result Word2Vec embeddings reveal functional chord and harmonic associations.
Adaptive adjustment of semantic feature space improves zero-shot recognition.
problem Domain shift and hubness problems in zero-shot recognition.
method Adaptive adjustment of semantic feature space.
result Remarkable performance improvement compared to existing methods.
A new text representation model combines CNN and VAE for better semantic extraction.
problem Difficult to effectively extract semantic features and distinguish polysemy in text data.
method Integrates CNN for feature extraction and VAE for consistent Gaussian distribution.
result The model outperforms traditional classification algorithms in text classification tasks.
Paper proposes a new approach to unify and compare knowledge graph embedding methods.
problem Lack of understanding and comparison of existing knowledge graph embedding methods.
method Introduces a multi-embedding interaction mechanism to unify and generalize existing models.
result Proposes a new multi-embedding model based on quaternion algebra.
Generative model learns from multiple sources to embed words and relationships.
problem Lack of diverse training data in specialized domains.
method Integrates evidence from diverse data sources using affine transformations on semantic vector spaces.
result Outperforms recent models on link prediction tasks and partially observed data.
The paper improves semantic interpolation in latent spaces of implicit models.
problem Interpolating between latent points in implicit models requires careful distributional matching.
method Proposes modifying the prior code distribution to concentrate more probability mass near the origin.
result Linear interpolation paths are shortest and pass through high-density regions, improving sample quality and semantics.
Method compares sentences by cosine similarity of vector projections.
problem Measuring semantic similarity of sentences.
method Cosine similarity of vector projections of sentence groups.
result Advantages over existing methods in preserving word order and syntactic connections.
Proposes AMS-SFE to improve zero-shot learning by aligning semantic feature spaces.
problem Domain shift problem in zero-shot learning due to disjoint seen and unseen data.
method Expands semantic features using an autoencoder and aligns them with visual feature manifold.
result Remarkable performance improvement over existing methods.
New model creates code semantics vectors for better understanding.
problem Improving code understanding and embedding quality.
method Siamese recurrent neural network on Python source code.
result Model significantly outperforms bag-of-tokens embeddings.
W-RNN improves text classification by extracting serialized text semantics.
problem Semantic constraint in sparse representation classification methods.
method Weighted RNN using word vectors and recurrent neural networks.
result W-RNN outperforms other methods in precision, recall, F1, and loss values.
This paper explores sentence vector properties for automatic summarization.
problem Understanding the internal structure and properties of sentence vectors.
method Compositional sentence vector representations using artificial neural networks.
result Cosine similarity correlates with sentence importance and can identify gaps in summaries.
A new ZSL algorithm uses shared sparse representations for unseen classes.
problem Classifying images from unseen classes using only semantic information.
method Coupled dictionary learning to represent visual and semantic features in an intermediate space.
result The proposed method outperforms state-of-the-art ZSL algorithms on benchmark datasets.
A new neural network model extends word embedding vectors with MeSH concepts for biomedical semantic similarity.
problem Eliciting semantic similarity between biomedical concepts remains challenging.
method Proposes a MeSH-gram neural network model that extends skip-gram by using MeSH descriptors.
result MeSH-gram outperforms skip-gram and is comparable to best methods but requires more computation and external resources.
CADD improves generative quality by augmenting discrete diffusion with continuous latent space.
problem Loss of semantic information between denoising steps in discrete diffusion models.
method Introduces a framework that augments discrete state space with a continuous latent space, allowing for graded, informative masked tokens.
result CADD improves generative quality across text generation, image synthesis, and code modeling.
RCAV quantifies model sensitivity to semantic concepts, improving interpretability methods.
problem Lack of semantic interpretability in image classification models.
method RCAV calculates concept gradients and ascent steps to assess model sensitivity to semantic concepts.
result RCAV yields more accurate and robust interpretations of model behavior.
Ultra-fast search algorithm for trillion-scale corpora with semantic flexibility.
problem Efficiently searching over large natural language corpora with semantic variations.
method String matching based on suffix arrays, vector representation of words, dynamic corpus-aware pruning, fast exact lookup.
result Substantially lower search latency compared to existing methods on FineWeb-Edu corpus.
Method infers domain-specific models without domain semantic descriptors.
problem Poor performance of standard supervised learning methods in unseen domains.
method Introduces latent domain vectors and neural networks for optimization.
result Inference of appropriate domain-specific models without semantic descriptors.
This work shows cosine similarity is equivalent to Pearson correlation for word vectors, but not all vectors are suitable for cosine.
problem The use of cosine similarity for semantic textual similarity is often taken for granted, despite its limitations.
method Characterized cases where Pearson correlation is unfit and introduced rank correlation as an alternative.
result Pearson correlation is equivalent to cosine similarity for many word vectors but not all, and rank correlation can improve performance.
The study uses machine learning to model semantic drift in digital content.
problem Detecting and measuring semantic drift in evolving digital content.
method Employing machine learning algorithms on a dataset of Tate Galleries metadata.
result Semantic drift can be modeled using a metaphor of social mechanics.
Advances word embedding methods using Markov processes.
problem Improving semantic understanding of word representations.
method Ground embeddings in cognitive literature, unify and generalize metric recovery methods.
result New algorithms for metric recovery and embedding on graphs and manifolds.
ISDA augments deep networks by adding semantic transformations.
problem Improving deep network generalization through semantic data augmentation.
method ISDA augments deep feature space by estimating covariance and drawing random vectors.
result ISDA consistently improves deep model performance on various datasets.
Paper proposes structured semantic perturbations to improve adversarial attacks.
problem Vulnerability of deep neural networks to adversarial attacks.
method Manipulates semantic attributes via disentangled latent codes.
result Demonstrates the effectiveness of structured semantic perturbations.
Dynamic model tracks word meanings over time.
problem Capturing semantic evolution of words over time.
method Latent diffusion process, variational inference algorithms.
result Higher predictive likelihoods and interpretable word trajectories.
A new text clustering method using NMF and LSA improves stability and performance.
problem Text data's large, sparse term-document matrix makes clustering difficult.
method Proposes a new feature agglomeration method based on NMF and deterministic K-Means initialization.
result Significantly improves clustering performance and stability.
EmbNum learns numerical attribute representations without distributional assumptions.
problem Semantic labeling of numerical values with unknown distributions.
method Neural numerical embedding model (EmbNum) for deep metric learning.
result EmbNum significantly outperforms state-of-the-art methods for numerical attribute semantic labeling.
LPL optimizes embeddings to align local neighborhoods, improving cross-lingual word alignment.
problem Aligning embeddings across different datasets and languages.
method Locality Preserving Loss (LPL) optimizes model to project embeddings while maintaining local neighborhoods and aligning them.
result LPL-based alignment leads to better and consistent accuracy, especially in small training set settings.
Semantic word embeddings represent the meaning of a word via a vector, and are created by diverse methods. Many use nonlinear operations on co-occurrence statistics, and have hand-tuned hyperparameters and reweighting methods. This paper proposes a new generative model, a dynamic version of the log-linear topic model o…
Unsupervised method improves word vectors by suppressing high variance features.
problem Improving semantic information in word vectors.
method Using conceptors to suppress high variance features in word vectors.
result Post-processed word vectors outperform existing alternatives in lexical evaluation tasks.
Model converts code snippets into vectors for predicting method names.
problem Representing code as vectors for semantic analysis.
method Decomposes code into abstract syntax tree paths, learns atomic representations simultaneously with aggregation.
result Code vectors trained on 14M methods can predict method names from unseen files.
Import2vec creates embeddings for software libraries to improve learning tasks.
problem Developing semantic representations for software libraries.
method Applied word embedding techniques from NLP to library packages.
result Library vectors capture meaningful relationships among libraries.
New method visualizes tabular feature semantics for better model understanding.
problem Lack of feature interaction interpretation in tabular ML models.
method Feature Vectors method for global tabular dataset interpretability.
result Visualizes semantic relationships among tabular features.
New method reduces word embedding storage space by 100x.
problem Large space required for storing word embeddings.
method Inspired by quantum computing, proposes word2ket and word2ketXS methods.
result Achieves a hundred-fold reduction in space required for word embeddings.
Embeddings leak sensitive information about input data, which can be recovered or inferred.
problem Information leakage in embedding models.
method Developed three classes of attacks to study information leakage.
result Embeddings leak sensitive information about input data, which can be partially recovered or inferred.
Improved segmentation model adaptation for new domains.
problem Reduced performance of pre-trained models on new domains.
method Calculated soft-label prototypes and predicted closest to class probabilities.
result Significant performance improvements on synthetic-to-real segmentation.
Graph neural networks can perform approximate reasoning in latent space for mathematical statements.
problem Can neural networks perform steps of approximate reasoning in a fixed dimensional latent space?
method Design and conduct an experiment using graph neural networks to predict rewrite-success of mathematical statements in a latent space.
result Graph neural networks can make non-trivial predictions about rewrite-success of statements in latent space.
New method uses word subspaces and term-frequency to improve text classification.
problem Lack of semantic meaning in bag-of-words features.
method Proposes word subspaces and term-frequency weighted word subspaces for text classification.
result Improved text classification performance compared to state-of-the-art algorithms.
This work provides uncertainty intervals for semantic latent variables in disentangled latent spaces.
problem Challenges in providing meaningful uncertainty quantification for semantic information in disentangled latent spaces.
method Uses quantile regression to output heuristic uncertainty intervals, calibrates these intervals to contain true latent values, and propagates them through the generator.
result Reliably communicates semantically meaningful, principled, and instance-adaptive uncertainty in image super-resolution and image completion.
Corpus poisoning can manipulate word meanings in word embeddings, affecting natural language processing tasks.
problem Controlling word meanings via corpus modifications.
method Developed an explicit expression over corpus features to control word embeddings.
result Demonstrated the ability to manipulate word meanings in word embeddings, affecting various downstream tasks.
New graph embedding method uses domain-specific knowledge.
problem Lack of semantic information in existing graph embedding approaches.
method Domain-aware biased random walks to incorporate semantic information.
result Embeddings achieve equal or greater accuracy than domain-independent methods.
Sherlock uses deep learning to accurately detect data types from column headers.
problem Detecting accurate semantic types of data columns for data science tasks.
method Sherlock is a multi-input deep neural network trained on a corpus of 686,765 data columns.
result Sherlock achieves a support-weighted F1 score of 0.89, outperforming existing methods.