A new method for higher-order co-occurrences in hypergraphs.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper proposes a new method for accurate data labeling using pairwise co-occurrences.
We present a new algorithm for identifying the transition and emission probabilities of a hidden Markov model (HMM) from the emitted data. Expectation-maximization becomes computationally prohibitive for long observation records, which are often required for identification. The new algorithm is particularly suitable fo…
Study on learning distributions from multiple data providers with restricted samples.
Develops a measure-theoretic framework for complex co-occurrence data.
Advances citation and subject label recommendation using multi-modal adversarial autoencoders.
Matrix Chernoff bound for Markov chains applied to co-occurrence matrices.
Matrix factorization is a simple and effective solution to the recommendation problem. It has been extensively employed in the industry and has attracted much attention from the academia. However, it is unclear what the low-dimensional matrices represent. We show that matrix factorization can actually be seen as simult…
Representing a word by its co-occurrences with other words in context is an effective way to capture the meaning of the word. However, the theory behind remains a challenge. In this work, taking the example of a word classification task, we give a theoretical analysis of the approaches that represent a word X by a func…
The paper models and predicts co-occurrence counts using Gamma regression.
Anomaly detection plays an important role in modern data-driven security applications, such as detecting suspicious access to a socket from a process. In many cases, such events can be described as a collection of categorical values that are considered as entities of different types, which we call heterogeneous categor…
Analyzes word vectors and co-occurrence statistics in NLP models.
Framework uncovers symmetric and asymmetric species associations from data.
We present a novel method for hierarchical topic detection where topics are obtained by clustering documents in multiple ways. Specifically, we model document collections using a class of graphical models called hierarchical latent tree models (HLTMs). The variables at the bottom level of an HLTM are observed binary va…
Chromatic Learning reduces feature dimensions for sparse datasets.
Person re-identification (re-id), an emerging problem in visual surveillance, deals with maintaining entities of individuals whilst they traverse various locations surveilled by a camera network. From a visual perspective re-id is challenging due to significant changes in visual appearance of individuals in cameras wit…
LLMs learn new tasks from unstructured data, but it depends on word co-occurrence and positional information.
Multi-view clustering has received much attention recently. Most of the existing multi-view clustering methods only focus on one-sided clustering. As the co-occurring data elements involve the counts of sample-feature co-occurrences, it is more efficient to conduct two-sided clustering along the samples and features si…
Probabilistic Latent Semantic Analysis is a novel statistical technique for the analysis of two-mode and co-occurrence data, which has applications in information retrieval and filtering, natural language processing, machine learning from text, and in related areas. Compared to standard Latent Semantic Analysis which s…
Text corpora are widely used resources for measuring societal biases and stereotypes. The common approach to measuring such biases using a corpus is by calculating the similarities between the embedding vector of a word (like nurse) and the vectors of the representative words of the concepts of interest (such as gender…
An important problem in multi-label classification is to capture label patterns or underlying structures that have an impact on such patterns. This paper addresses one such problem, namely how to exploit hierarchical structures over labels. We present a novel method to learn vector representations of a label space give…
Transformers learn topic structure through embedding and attention mechanisms.
This paper focuses on a traditional relation extraction task in the context of limited annotated data and a narrow knowledge domain. We explore this task with a clinical corpus consisting of 200 breast cancer follow-up treatment letters in which 16 distinct types of relations are annotated. We experiment with an approa…
DenseHMM improves HMMs by learning dense representations that enable gradient-based optimization.
Separable Non-negative Matrix Factorization (SNMF) is an important method for topic modeling, where "separable" assumes every topic contains at least one anchor word, defined as a word that has non-zero probability only on that topic. SNMF focuses on the word co-occurrence patterns to reveal topics by two steps: anchor…
A new test statistic counts tree co-occurrences to detect edge correlation between networks.
Latent topic models have been successfully applied as an unsupervised topic discovery technique in large document collections. With the proliferation of hypertext document collection such as the Internet, there has also been great interest in extending these approaches to hypertext [6, 9]. These approaches typically mo…
This paper analyzes DeepWalk and node2vec for community detection in large networks.
BGNN improves GNN by modeling interactions between neighbor nodes.
We develop necessary and sufficient conditions and a novel provably consistent and efficient algorithm for discovering topics (latent factors) from observations (documents) that are realized from a probabilistic mixture of shared latent factors that have certain properties. Our focus is on the class of topic models in …
Develops a non-parametric Dirichlet process method for probabilistic biclustering.
Unified framework for comparing clusterings from information-theoretic and pair-counting perspectives.
In this paper, we propose a new kernel-based co-occurrence measure that can be applied to sparse linguistic expressions (e.g., sentences) with a very short learning time, as an alternative to pointwise mutual information (PMI). As well as deriving PMI from mutual information, we derive this new measure from the Hilbert…
The paper classifies trades into types based on proximity and measures their impact on stock prices.
To understand the relationship between news sentiment and company stock price movements, and to better understand connectivity among companies, we define an algorithm for measuring sentiment-based network risk. The algorithm ranks companies in networks of co-occurrences, and measures sentiment-based risk, by calculatin…
Most popular word embedding techniques involve implicit or explicit factorization of a word co-occurrence based matrix into low rank factors. In this paper, we aim to generalize this trend by using numerical methods to factor higher-order word co-occurrence based arrays, or \textit{tensors}. We present four word embedd…
We study the problem of ranking from crowdsourced pairwise comparisons. Answers to pairwise tasks are known to be affected by the position of items on the screen, however, previous models for aggregation of pairwise comparisons do not focus on modeling such kind of biases. We introduce a new aggregation model factorBT …
IMPACC improves consensus clustering for bioinformatics data.
In this study, a pairwise comparison matrix is generalized to the case when coefficients create Lie group , non necessarily abelian. A necessary and sufficient criterion for pairwise comparisons matrices to be consistent is provided. Basic criteria for finding a nearest consistent pairwise comparisons matrix (extend…
Paper introduces differential pairwise privacy for secure metric learning.
ROVAE uses noisy pairwise comparisons to disentangle factors in VAEs.
Cross-entropy loss linked to metric learning, outperforming complex pairwise losses.
Develops a statistical framework to measure uncertainty in model rankings based on human preferences.
Improves labeling quality in machine learning with pairwise feedback.
As one of the most important types of (weaker) supervised information in machine learning and pattern recognition, pairwise constraint, which specifies whether a pair of data points occur together, has recently received significant attention, especially the problem of pairwise constraint propagation. At least two reaso…
Paper proposes Pcomp classification for binary classification with pairwise confidence comparisons.
Study on pairwise counter-monotonicity, a type of negative dependence.
Paper models spatio-temporal extremes using conditional variational autoencoders.