We create interpretable word embeddings through sparse coding.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
DiSeNE generates interpretable node embeddings without supervision.
POLAR framework interprets word embeddings using polar opposites.
Embeddings are ubiquitous in machine learning, appearing in recommender systems, NLP, and many other applications. Researchers and developers often need to explore the properties of a specific embedding, and one way to analyze embeddings is to visualize them. We present the Embedding Projector, a tool for interactive v…
New method interprets deep embeddings for diabetes patient clustering.
Word embeddings have demonstrated strong performance on NLP tasks. However, lack of interpretability and the unsupervised nature of word embeddings have limited their use within computational social science and digital humanities. We propose the use of informative priors to create interpretable and domain-informed dime…
New algorithm improves interpretability in sequence classification.
New model combines neural networks and embeddings for better choice modeling interpretability.
Unified interpretation of softmax cross-entropy and negative sampling for knowledge graph embedding.
Knowledge bases are employed in a variety of applications from natural language processing to semantic web search; alas, in practice their usefulness is hurt by their incompleteness. Embedding models attain state-of-the-art accuracy in knowledge base completion, but their predictions are notoriously hard to interpret. …
In this paper we propose and study the novel problem of explaining node embeddings by finding embedded human interpretable subspaces in already trained unsupervised node representation embeddings. We use an external knowledge base that is organized as a taxonomy of human-understandable concepts over entities as a guide…
CAMEL enhances manifold embedding and learning with curvature metrics.
Generating high-quality and interpretable adversarial examples in the text domain is a much more daunting task than it is in the image domain. This is due partly to the discrete nature of text, partly to the problem of ensuring that the adversarial examples are still probable and interpretable, and partly to the proble…
Minimal token perturbations reveal how Transformer models process information.
VICE embeds concepts in a vector space using human data.
A new HP model balances interpretability and flexibility for EHR event sequences.
TransINT embeds KGs by preserving implication rules, outperforming existing methods.
Document network embedding aims at learning representations for a structured text corpus i.e. when documents are linked to each other. Recent algorithms extend network embedding approaches by incorporating the text content associated with the nodes in their formulations. In most cases, it is hard to interpret the learn…
Topic modeling analyzes documents to learn meaningful patterns of words. However, existing topic models fail to learn interpretable topics when working with large and heavy-tailed vocabularies. To this end, we develop the Embedded Topic Model (ETM), a generative model of documents that marries traditional topic models …
DCR improves interpretability of concept-based models by using neural networks to build rule structures.
A novel GP architecture, Thin and Deep GP, learns lower-dimensional representations without losing interpretability.
Dimensionality reduction (DR) on the manifold includes effective methods which project the data from an implicit relational space onto a vectorial space. Regardless of the achievements in this area, these algorithms suffer from the lack of interpretation of the projection dimensions. Therefore, it is often difficult to…
Low dimensional embeddings that capture the main variations of interest in collections of data are important for many applications. One way to construct these embeddings is to acquire estimates of similarity from the crowd. However, similarity is a multi-dimensional concept that varies from individual to individual. Ex…
Following great success in the image processing field, the idea of adversarial training has been applied to tasks in the natural language processing (NLP) field. One promising approach directly applies adversarial training developed in the image processing field to the input word embedding space instead of the discrete…
AMES framework selects optimal embedding space for latent graph inference.
Node embedding is the task of extracting informative and descriptive features over the nodes of a graph. The importance of node embeddings for graph analytics, as well as learning tasks such as node classification, link prediction and community detection, has led to increased interest on the problem leading to a number…
Study geodesic properties of time series data using Wasserstein metric.
A framework for stable dynamic network embeddings using static methods.
We break down transformer embeddings into interpretable components revealing hidden geometric structures.
This work analyzes PPR-based node embeddings and their topological information.
In this paper we introduce score embedding, a neural network based model to learn interpretable vector representations for words. Score embedding is a supervised method that takes advantage of the labeled training data and the neural network architecture to learn interpretable representations for words. Health care has…
Combines OT and PCA for DR, preserving clusters.
OracleAD detects multivariate time series anomalies without labels.
A new method simplifies HLLE for better robustness.
Models use embeddings and attention for better claim severity prediction.
Hybrid model learns interpretable meal-level glycemic control.
Most existing word embedding methods can be categorized into Neural Embedding Models and Matrix Factorization (MF)-based methods. However some models are opaque to probabilistic interpretation, and MF-based methods, typically solved using Singular Value Decomposition (SVD), may incur loss of corpus information. In addi…
LOT framework embeds high-dimensional cell data into interpretable Euclidean space.
The interpretability of machine learning, particularly for deep neural networks, is crucial for decision making in real-world applications. One approach is replacing the un-interpretable machine learning model with a surrogate model, which has a simple structure for interpretation. Another approach is understanding the…
With the rising interest in graph representation learning, a variety of approaches have been proposed to effectively capture a graph's properties. While these approaches have improved performance in graph machine learning tasks compared to traditional graph techniques, they are still perceived as techniques with limite…
By representing words with probability densities rather than point vectors, probabilistic word embeddings can capture rich and interpretable semantic information and uncertainty. The uncertainty information can be particularly meaningful in capturing entailment relationships -- whereby general words such as "entity" co…
Graph embedding has become a key component of many data mining and analysis systems. Current graph embedding approaches either sample a large number of node pairs from a graph to learn node embeddings via stochastic optimization or factorize a high-order proximity/adjacency matrix of the graph via computationally expen…
Protein Thoughts interprets protein interactions with clear reasoning, improving prediction accuracy.
The transition amplitudes between coherent states on a coherent state manifold are expressed in terms of the embedding of the coherent state manifold into a projective Hilbert space. Consequences for the dimension of projective Hilbert space and a simple geometric interpretation of Calabi's diastasis follows.
MPVAE learns latent embeddings and label correlations for multi-label classification.
LIFE framework improves model accuracy and interpretability.
We present a novel spectral embedding of graphs that incorporates weights assigned to the nodes, quantifying their relative importance. This spectral embedding is based on the first eigenvectors of some properly normalized version of the Laplacian. We prove that these eigenvectors correspond to the configurations of lo…
Study evaluates interpretability of time series foundation models' latent spaces.