New methods identify concepts in trained embeddings reliably without human labels.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
CA connects visual concepts to neural network representations.
A framework visualizes embedding spaces of neural survival analysis models using anchor directions.
Embedding large and high dimensional data into low dimensional vector spaces is a necessary task to computationally cope with contemporary data sets. Superseding latent semantic analysis recent approaches like word2vec or node2vec are well established tools in this realm. In the present paper we add to this line of res…
Representation learning methods that transform encoded data (e.g., diagnosis and drug codes) into continuous vector spaces (i.e., vector embeddings) are critical for the application of deep learning in healthcare. Initial work in this area explored the use of variants of the word2vec algorithm to learn embeddings for m…
Proposes ACCA for better alignment of multiple data perspectives.
In this paper we propose and study the novel problem of explaining node embeddings by finding embedded human interpretable subspaces in already trained unsupervised node representation embeddings. We use an external knowledge base that is organized as a taxonomy of human-understandable concepts over entities as a guide…
This work presents the concept of kernel mean embedding and kernel probabilistic programming in the context of stochastic systems. We propose formulations to represent, compare, and propagate uncertainties for fairly general stochastic dynamics in a distribution-free manner. The new tools enjoy sound theory rooted in f…
VICE embeds concepts in a vector space using human data.
Word embeddings are a popular approach to unsupervised learning of word relationships that are widely used in natural language processing. In this article, we present a new set of embeddings for medical concepts learned using an extremely large collection of multimodal medical data. Leaning on recent theoretical insigh…
Curve shortening flow shrinks curves to points.
Enhancing spectral embedding for low-dimensional embeddings in rare disease cohorts
DCR improves interpretability of concept-based models by using neural networks to build rule structures.
Introduces Hurewicz fibrations for embedding maps of orbifold charts.
A method for concept-based learning using probabilistic inference and expert rules.
We present a baseline approach for cross-modal knowledge fusion. Different basic fusion methods are evaluated on existing embedding approaches to show the potential of joining knowledge about certain concepts across modalities in a fused concept representation.
Prob2Vec embeds problems for adaptive tutoring, achieving high similarity accuracy.
Word representation is fundamental in NLP tasks, because it is precisely from the coding of semantic closeness between words that it is possible to think of teaching a machine to understand text. Despite the spread of word embedding concepts, still few are the achievements in linguistic contexts other than English. In …
In an effort to understand the meaning of the intermediate representations captured by deep networks, recent papers have tried to associate specific semantic concepts to individual neural network filter responses, where interesting correlations are often found, largely by focusing on extremal filter responses. In this …
Develops tensor calculus for submanifolds of arbitrary codimension.
PACE explains ViTs by modeling patch-level concept distributions, surpassing existing methods.
After learning a concept, humans are also able to continually generalize their learned concepts to new domains by observing only a few labeled instances without any interference with the past learned knowledge. In contrast, learning concepts efficiently in a continual learning setting remains an open challenge for curr…
Service robots benefit from encoding information in semantically meaningful ways to enable more robust task execution. Prior work has shown multi-relational embeddings can encode semantic knowledge graphs to promote generalizability and scalability, but only within a batched learning paradigm. We present Incremental Se…
IITK wins FinSim 2020 task on financial hypernym detection.
This is a survey of our research on geometric structures of projective embeddings and includes some topics of our talks in several symposia during 1990-99. We clarify our main problem, which is to construct a kind of geometric composition series of projective embeddings. The concept of "geometric composition series" is…
New models handle survival analysis with concept-based learning.
In Classical Knot Theory and in the new Theory of Quantum Invariants substantial effort was directed toward the search for unknotting moves on links. We solve, in this note, several classical problems concerning unknotting moves. Our approach uses a new concept, Burnside groups of links, which establishes unexpected re…
FONDUE identifies ambiguous nodes in networks for better analysis.
CONDA-PM framework helps analyze concept drift in business processes.
Paper generalizes wrinkled embedding concept to jet spaces.
Introduces Kähler duality between domains in complex space.
LCBM model improves image classification without human supervision.
New framework quantifies and reduces concept-based models' leakage.
Word embeddings have demonstrated strong performance on NLP tasks. However, lack of interpretability and the unsupervised nature of word embeddings have limited their use within computational social science and digital humanities. We propose the use of informative priors to create interpretable and domain-informed dime…
Proposes a method to align language and image data.
Electronic health record (EHR) systems are used extensively throughout the healthcare domain. However, data interchangeability between EHR systems is limited due to the use of different coding standards across systems. Existing methods of mapping coding standards based on manual human experts mapping, dictionary mappin…
Risk adjustment has become an increasingly important tool in healthcare. It has been extensively applied to payment adjustment for health plans to reflect the expected cost of providing coverage for members. Risk adjustment models are typically estimated using linear regression, which does not fully exploit the informa…
Embedding methods which enforce a partial order or lattice structure over the concept space, such as Order Embeddings (OE) (Vendrov et al., 2016), are a natural way to model transitive relational data (e.g. entailment graphs). However, OE learns a deterministic knowledge base, limiting expressiveness of queries and the…
Automated extraction of concepts from patient clinical records is an essential facilitator of clinical research. For this reason, the 2010 i2b2/VA Natural Language Processing Challenges for Clinical Records introduced a concept extraction task aimed at identifying and classifying concepts into predefined categories (i.…
Let M and N be smooth manifolds without boundary. Immersion theory suggests that an understanding of the space of smooth embeddings emb(M,N) should come from an analysis of the cofunctor V |--> emb(V,N) from the poset O of open subsets of M to spaces. We therefore abstract some of the properties of this cofunctor, and …
SketchEmbedNet learns image representations from sketches, useful for few-shot learning.
The scientific literature is a rich source of information for data mining with conceptual knowledge graphs; the open science movement has enriched this literature with complementary source code that implements scientific models. To exploit this new resource, we construct a knowledge graph using unsupervised learning me…
A new framework CL embeds features and labels for multi-label classification.
Debias concept-based explanations by removing confounding information.
Topological data analysis is an emerging mathematical concept for characterizing shapes in multi-scale data. In this field, persistence diagrams are widely used as a descriptor of the input data, and can distinguish robust and noisy topological properties. Nowadays, it is highly desired to develop a statistical framewo…
A neural network learns phase space properties for time series analysis.
New method for simplifying complex 4D shapes with boundaries.
Text corpora are widely used resources for measuring societal biases and stereotypes. The common approach to measuring such biases using a corpus is by calculating the similarities between the embedding vector of a word (like nurse) and the vectors of the representative words of the concepts of interest (such as gender…