Develops a measure-theoretic framework for complex co-occurrence data.
problem Modeling and interpreting complex co-occurrences in high-dimensional data.
method Introduces measure-theoretic probability and conditional probability, investigates E-integrals.
result Establishes a rigorous measure-theoretic foundation for co-occurrence modeling.
Analyzes word vectors and co-occurrence statistics in NLP models.
problem Understanding biases in NLP models through co-occurrence statistics.
method Developed an analytic model of statistics learned by Word2Vec and GloVe, derived the first solution to Word2Vec's algorithm, and analyzed independence in co-occurrence models.
result Demonstrated a universal property of word vectors that can reveal biases in data before they are absorbed by DL models.
Matrix Chernoff bound for Markov chains applied to co-occurrence matrices.
problem Analyzing the behavior of co-occurrence statistics in sequential data.
method Proved a matrix Chernoff-type bound for sums of matrix-valued random variables sampled via a regular Markov chain.
result Achieved exponentially fast convergence rate and sample complexity analysis for co-occurrence matrices.
Probabilistic Latent Semantic Analysis is a novel statistical technique for the analysis of two-mode and co-occurrence data, which has applications in information retrieval and filtering, natural language processing, machine learning from text, and in related areas. Compared to standard Latent Semantic Analysis which s…
Study proposes an alternative method to measure societal biases using smoothed co-occurrence relations.
problem Measuring societal biases using word embeddings can introduce irrelevant concepts.
method Proposes an alternative approach using smoothed first-order co-occurrence relations.
result First-order approach shows higher correlations with actual gender bias statistics.
A new test statistic counts tree co-occurrences to detect edge correlation between networks.
problem Detecting edge correlation between networks using latent vertex correspondence.
method The test statistic is based on counting co-occurrences of signed trees for a family of non-isomorphic trees.
result The test runs in n2+o(1) time and succeeds with high probability for large n. A new method for higher-order co-occurrences in hypergraphs.
problem Computing higher-order co-occurrences in hypergraphs.
method Face-splitting product or transpose Khatri-Rao product for higher order tuple co-occurrences.
result Demonstrates the utility of the higher order co-occurrence tensor in NLP and hypergraph models.
Paper proposes a new method for accurate data labeling using pairwise co-occurrences.
problem Accurate data labeling via crowdsourcing with limited data.
method Pairwise co-occurrences framework and algebraic/identifiability-enhanced algorithms.
result The approach can identify the Dawid-Skene model under realistic conditions.
Matrix factorization simplifies user-item co-occurrence analysis.
problem Understanding the meaning of low-dimensional matrices in matrix factorization.
method Showed matrix factorization equals calculating eigenvectors of co-occurrence matrices, using RMT insights.
result Low-dimension matrices represent a reduced noise user and item co-occurrence space.
Analyzes how word meaning is captured by co-occurrence features.
problem Understanding the theoretical basis of word representation by co-occurrences.
method Theoretical analysis of word representation methods using co-occurrences.
result Using multiple context features improves word prediction scores.
Advances citation and subject label recommendation using multi-modal adversarial autoencoders.
problem Improving recommendation systems for citations and subject labels.
method Multi-modal adversarial autoencoders with adversarial regularization, sparsity, and input modality analysis.
result Adversarial regularization consistently improves recommendation performance.
The paper models and predicts co-occurrence counts using Gamma regression.
problem Predicting relevance between items or users from high-dimensional sparse co-occurrence count data.
method Shared parameter alternating zero-inflated Gamma regression models (SA-ZIG) with Fisher scoring and learning rate adjustment.
result SA-ZIG with learning rate adjustment performs satisfactorily in predicting relevance.
This paper justifies and improves lexicon-based classification without labeled data.
problem Lack of justification for lexicon-based classification and its lower accuracy compared to supervised methods.
method Derives probabilistic justification and learns weights from co-occurrence statistics.
result Lexicon-based classification can be improved without labeled data, offering higher accuracy.
Proposes a method for two-sided clustering of co-occurrence data.
problem Efficient clustering of co-occurrence data in multi-view settings.
method Information-theoretic multi-view co-clustering (MV-ITCC).
result Demonstrates superior performance on text and image datasets.
New algorithm learns HMM from pairwise co-occurrences, improving topic modeling.
problem Identifying hidden Markov models from limited pairwise co-occurrence data.
method Uses pairwise co-occurrence data to uniquely identify HMMs, even if higher-order probabilities are unknown.
result Shows improved topic modeling quality with HMMs compared to bag-of-words models.
We present a novel method for hierarchical topic detection where topics are obtained by clustering documents in multiple ways. Specifically, we model document collections using a class of graphical models called hierarchical latent tree models (HLTMs). The variables at the bottom level of an HLTM are observed binary va…
Chromatic Learning reduces feature dimensions for sparse datasets.
problem Sparse, high-dimensional data challenges traditional learning methods.
method Graph coloring over co-occurrence graph to create dense feature representation.
result Compresses sparse datasets significantly while maintaining model accuracy.
A new method selects anchor words for better topic discovery in text corpora.
problem Selecting anchor words for improved topic modeling in text corpora.
method Proposes a new greedy method to find a minimum edge-weight anchor clique in a word similarity graph.
result The proposed method outperforms existing methods on topic quality and is faster.
Paper improves relation extraction in clinical texts with limited data.
problem Relation extraction in narrow knowledge domains with scarce annotated data.
method Introduces a bag-of-concepts (BoC) model and compares it with window-bounded co-occurrence (WBC).
result BoC model outperforms baseline and other complex methods on small dataset.
Person re-identification (re-id), an emerging problem in visual surveillance, deals with maintaining entities of individuals whilst they traverse various locations surveilled by a camera network. From a visual perspective re-id is challenging due to significant changes in visual appearance of individuals in cameras wit…
LLMs learn new tasks from unstructured data, but it depends on word co-occurrence and positional information.
problem Understanding how LLMs can learn new tasks from unstructured data without explicit training.
method Examined the capabilities of LLMs trained on unstructured data, focusing on sequence model requirements and training data structure.
result Many ICL capabilities can emerge from word co-occurrence in unstructured data, but positional information is crucial for certain tasks.
Investor clusters analyzed in Helsinki Stock Exchange IPOs.
problem Lack of research on investor behavior in IPOs.
method Statistically validated network method to infer investor links based on trade timing.
result Large network structures form in IPO and mature companies, with evidence of institutional herding.
Semantic word embeddings represent the meaning of a word via a vector, and are created by diverse methods. Many use nonlinear operations on co-occurrence statistics, and have hand-tuned hyperparameters and reweighting methods. This paper proposes a new generative model, a dynamic version of the log-linear topic model o…
Study on learning distributions from multiple data providers with restricted samples.
problem Learning an unknown distribution from conditional samples given a fixed family of queryable sets.
method Studied a stylized model of distribution learning from restricted conditional samples, focusing on the co-occurrence graph associated with the queryable sets.
result The optimal sample complexity of PAC learning ranges from nearly linear to quadratic, depending on the structure of the queryable sets.
New method tests weighted networks without thresholding, improving accuracy.
problem Testing and anomaly detection on weighted network data.
method Hierarchical Bayesian hypothesis testing framework for weighted networks.
result Method shows lower Type I error and higher statistical power compared to alternatives.
A fast kernel-based measure for sparse linguistic expressions.
problem Efficiently measuring co-occurrence in sparse linguistic data.
method Derives PHSIC from HSIC, estimates it linearly, and uses various kernels.
result Empirically, PHSIC outperforms PMI in accuracy and learning speed.
An important problem in multi-label classification is to capture label patterns or underlying structures that have an impact on such patterns. This paper addresses one such problem, namely how to exploit hierarchical structures over labels. We present a novel method to learn vector representations of a label space give…
We introduce a factor analysis model that summarizes the dependencies between observed variable groups, instead of dependencies between individual variables as standard factor analysis does. A group may correspond to one view of the same set of objects, one of many data sets tied by co-occurrence, or a set of alternati…
New method finds robust clusters with statistical guarantees.
problem Clustering solutions are unstable and lack robustness guarantees.
method Quantifies cluster instability and finds robust clusters (core clusters).
result Core clusters are more stable and robust to changes in data.
The paper introduces tensor factorization for word embeddings.
problem Creating embeddings for words with multiple meanings.
method Tensor factorization of higher-order co-occurrence arrays.
result Tensor-based embeddings can discern polysemous words' meanings.
Deep learning solves jigsaw puzzles by classifying fragment positions.
problem Automated reconstruction of archaeological fragments from jigsaw puzzles.
method Classifies relative positions of fragments using deep neural networks and local feature co-occurrences.
result Our method outperforms state-of-the-art by 25%.
Study analyzes fintech terms in news and blogs, revealing specialized attributes of fintech companies.
problem Understanding specialized attributes of fintech companies through term analysis.
method Large scale analysis of fintech terms in news and blogs, using complex networks and statistically validated networks.
result Companies with fintech terms have over-expressions of specific attributes related to geography and economy.
DenseHMM improves HMMs by learning dense representations that enable gradient-based optimization.
problem Learning dense representations for hidden states and observables in HMMs.
method DenseHMM uses kernelized transition probabilities and two optimization schemes.
result DenseHMM achieves superior performance and expressiveness compared to standard HMMs.
Latent topic models have been successfully applied as an unsupervised topic discovery technique in large document collections. With the proliferation of hypertext document collection such as the Internet, there has also been great interest in extending these approaches to hypertext [6, 9]. These approaches typically mo…
This paper analyzes DeepWalk and node2vec for community detection in large networks.
problem Community detection in large, sparse networks.
method Low-dimensional network embedding algorithms (DeepWalk and node2vec) applied to random walk segments.
result The performance of DeepWalk and node2vec in recovering communities depends on the length of random walk segments and sparsity of the network.
We develop necessary and sufficient conditions and a novel provably consistent and efficient algorithm for discovering topics (latent factors) from observations (documents) that are realized from a probabilistic mixture of shared latent factors that have certain properties. Our focus is on the class of topic models in …
Develops a non-parametric Dirichlet process method for probabilistic biclustering.
problem Challenges in finding biclusters with strong co-occurrence in rows and columns.
method Dual Dirichlet process mixture models for row and column clustering, with cluster number determined by data.
result Improves bicluster extraction in text mining and gene expression analysis.
A new model generates summaries by conditioning on input text and latent topics.
problem Improving abstractive summarization quality.
method Conditioning decoder output on both input text and latent topics identified by LDA.
result Strongly improved ROUGE scores on CNN/Daily Mail and WikiHow datasets.
The paper classifies trades into types based on proximity and measures their impact on stock prices.
problem Understanding the impact of high-frequency trades on stock prices and their predictability.
method Classifies trades into five types based on proximity, measures conditional order imbalance (COI), and develops trading strategies.
result Strong positive correlations between contemporaneous returns and COIs, and positive associations with future returns for isolated trades.
GEMRank embeds users and items using co-occurrence relations for better collaborative filtering.
problem Lack of textual data for entity embedding in recommender systems.
method Uses profile co-occurrence for entity relations and factorization for embedding. Feeds embeddings into a neural network for predictions.
result Significantly outperforms baseline algorithms in various data sets.
New topic modeling method scales to large datasets.
problem Large-scale topic modeling with high co-occurrence data.
method Introduced Full Dependence Mixture (FDM) model for direct topic learning.
result FDM model performs comparably or better than benchmarks on large datasets.
Algorithm measures sentiment-based network risk in companies.
problem Understanding the relationship between news sentiment and stock price movements.
method Algorithm ranks companies based on news sentiment and co-occurrences, calculating individual and aggregated risks.
result The highest quarterly risk value correlates with a higher chance of stock price decline up to 70 days later.
IMPACC improves consensus clustering for bioinformatics data.
problem Consensus clustering's inefficiency and lack of interpretability for large-scale data.
method Ensemble minipatch co-occurrences, adaptive sampling of observations and features.
result Significantly improved accuracy and interpretability with substantial computational savings.
New methods for unsupervised learning of word and entity representations.
problem Learning distributed representations of words and entities from text and knowledge bases.
method MVLSA for words and NVSE for entities, both unsupervised learning methods.
result MVLSA and NVSE outperform state-of-the-art models in word and entity representation learning.
New DTMs model text evolution with Gaussian processes and scalable inference.
problem Challenges in modeling text evolution with continuous stochastic processes.
method Extended tractable priors to Gaussian processes and developed scalable inference methods.
result Found interesting patterns in large-scale datasets not accessible before.
The paper analyzes how news sentiment of companies can affect market movements.
problem Understanding how news sentiment impacts market performance and volatility.
method Applied NLP techniques to analyze news sentiment of 87 companies over 7 years.
result Strong media sentiment towards one company can indicate significant changes in sentiment towards related companies.
A hierarchical community detection method using recursive partitioning.
problem Finding interpretable and accurate community structures in networks.
method Top-down recursive partitioning starting with spectral clustering.
result The algorithm correctly recovers community trees under mild assumptions.
BBM models short texts using biterms to improve coherence.
problem Challenges in analyzing short texts from social media.
method Bag of Biterms (BoB) for document representation and simple statistical models.
result BBM enhances coherence and performance over traditional models.