Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

59118176235 · Jun 202019922001200920182026
48 results for co-occurrence statistics

Develops a measure-theoretic framework for complex co-occurrence data.

problem Modeling and interpreting complex co-occurrences in high-dimensional data.
method Introduces measure-theoretic probability and conditional probability, investigates E-integrals.
result Establishes a rigorous measure-theoretic foundation for co-occurrence modeling.

Analyzes word vectors and co-occurrence statistics in NLP models.

problem Understanding biases in NLP models through co-occurrence statistics.
method Developed an analytic model of statistics learned by Word2Vec and GloVe, derived the first solution to Word2Vec's algorithm, and analyzed independence in co-occurrence models.
result Demonstrated a universal property of word vectors that can reveal biases in data before they are absorbed by DL models.

Matrix Chernoff bound for Markov chains applied to co-occurrence matrices.

problem Analyzing the behavior of co-occurrence statistics in sequential data.
method Proved a matrix Chernoff-type bound for sums of matrix-valued random variables sampled via a regular Markov chain.
result Achieved exponentially fast convergence rate and sample complexity analysis for co-occurrence matrices.

Probabilistic Latent Semantic Analysis is a novel statistical technique for the analysis of two-mode and co-occurrence data, which has applications in information retrieval and filtering, natural language processing, machine learning from text, and in related areas. Compared to standard Latent Semantic Analysis which s…

2013-01-23abs ↗pdf ↗

Study proposes an alternative method to measure societal biases using smoothed co-occurrence relations.

problem Measuring societal biases using word embeddings can introduce irrelevant concepts.
method Proposes an alternative approach using smoothed first-order co-occurrence relations.
result First-order approach shows higher correlations with actual gender bias statistics.

A new test statistic counts tree co-occurrences to detect edge correlation between networks.

problem Detecting edge correlation between networks using latent vertex correspondence.
method The test statistic is based on counting co-occurrences of signed trees for a family of non-isomorphic trees.
result The test runs in n2+o(1)n^{2+o(1)} time and succeeds with high probability for large nn.

Paper proposes a new method for accurate data labeling using pairwise co-occurrences.

problem Accurate data labeling via crowdsourcing with limited data.
method Pairwise co-occurrences framework and algebraic/identifiability-enhanced algorithms.
result The approach can identify the Dawid-Skene model under realistic conditions.

Matrix factorization simplifies user-item co-occurrence analysis.

problem Understanding the meaning of low-dimensional matrices in matrix factorization.
method Showed matrix factorization equals calculating eigenvectors of co-occurrence matrices, using RMT insights.
result Low-dimension matrices represent a reduced noise user and item co-occurrence space.

Advances citation and subject label recommendation using multi-modal adversarial autoencoders.

problem Improving recommendation systems for citations and subject labels.
method Multi-modal adversarial autoencoders with adversarial regularization, sparsity, and input modality analysis.
result Adversarial regularization consistently improves recommendation performance.

The paper models and predicts co-occurrence counts using Gamma regression.

problem Predicting relevance between items or users from high-dimensional sparse co-occurrence count data.
method Shared parameter alternating zero-inflated Gamma regression models (SA-ZIG) with Fisher scoring and learning rate adjustment.
result SA-ZIG with learning rate adjustment performs satisfactorily in predicting relevance.

This paper justifies and improves lexicon-based classification without labeled data.

problem Lack of justification for lexicon-based classification and its lower accuracy compared to supervised methods.
method Derives probabilistic justification and learns weights from co-occurrence statistics.
result Lexicon-based classification can be improved without labeled data, offering higher accuracy.

New algorithm learns HMM from pairwise co-occurrences, improving topic modeling.

problem Identifying hidden Markov models from limited pairwise co-occurrence data.
method Uses pairwise co-occurrence data to uniquely identify HMMs, even if higher-order probabilities are unknown.
result Shows improved topic modeling quality with HMMs compared to bag-of-words models.

We present a novel method for hierarchical topic detection where topics are obtained by clustering documents in multiple ways. Specifically, we model document collections using a class of graphical models called hierarchical latent tree models (HLTMs). The variables at the bottom level of an HLTM are observed binary va…

2016-05-21abs ↗pdf ↗

A new method selects anchor words for better topic discovery in text corpora.

problem Selecting anchor words for improved topic modeling in text corpora.
method Proposes a new greedy method to find a minimum edge-weight anchor clique in a word similarity graph.
result The proposed method outperforms existing methods on topic quality and is faster.

Paper improves relation extraction in clinical texts with limited data.

problem Relation extraction in narrow knowledge domains with scarce annotated data.
method Introduces a bag-of-concepts (BoC) model and compares it with window-bounded co-occurrence (WBC).
result BoC model outperforms baseline and other complex methods on small dataset.

Person re-identification (re-id), an emerging problem in visual surveillance, deals with maintaining entities of individuals whilst they traverse various locations surveilled by a camera network. From a visual perspective re-id is challenging due to significant changes in visual appearance of individuals in cameras wit…

2014-06-13abs ↗pdf ↗

LLMs learn new tasks from unstructured data, but it depends on word co-occurrence and positional information.

problem Understanding how LLMs can learn new tasks from unstructured data without explicit training.
method Examined the capabilities of LLMs trained on unstructured data, focusing on sequence model requirements and training data structure.
result Many ICL capabilities can emerge from word co-occurrence in unstructured data, but positional information is crucial for certain tasks.

Semantic word embeddings represent the meaning of a word via a vector, and are created by diverse methods. Many use nonlinear operations on co-occurrence statistics, and have hand-tuned hyperparameters and reweighting methods. This paper proposes a new generative model, a dynamic version of the log-linear topic model o…

2015-02-12abs ↗pdf ↗

Study on learning distributions from multiple data providers with restricted samples.

problem Learning an unknown distribution from conditional samples given a fixed family of queryable sets.
method Studied a stylized model of distribution learning from restricted conditional samples, focusing on the co-occurrence graph associated with the queryable sets.
result The optimal sample complexity of PAC learning ranges from nearly linear to quadratic, depending on the structure of the queryable sets.

New method tests weighted networks without thresholding, improving accuracy.

problem Testing and anomaly detection on weighted network data.
method Hierarchical Bayesian hypothesis testing framework for weighted networks.
result Method shows lower Type I error and higher statistical power compared to alternatives.

An important problem in multi-label classification is to capture label patterns or underlying structures that have an impact on such patterns. This paper addresses one such problem, namely how to exploit hierarchical structures over labels. We present a novel method to learn vector representations of a label space give…

2014-12-22abs ↗pdf ↗

We introduce a factor analysis model that summarizes the dependencies between observed variable groups, instead of dependencies between individual variables as standard factor analysis does. A group may correspond to one view of the same set of objects, one of many data sets tied by co-occurrence, or a set of alternati…

2011-10-14abs ↗pdf ↗

Deep learning solves jigsaw puzzles by classifying fragment positions.

problem Automated reconstruction of archaeological fragments from jigsaw puzzles.
method Classifies relative positions of fragments using deep neural networks and local feature co-occurrences.
result Our method outperforms state-of-the-art by 25%.

Study analyzes fintech terms in news and blogs, revealing specialized attributes of fintech companies.

problem Understanding specialized attributes of fintech companies through term analysis.
method Large scale analysis of fintech terms in news and blogs, using complex networks and statistically validated networks.
result Companies with fintech terms have over-expressions of specific attributes related to geography and economy.

DenseHMM improves HMMs by learning dense representations that enable gradient-based optimization.

problem Learning dense representations for hidden states and observables in HMMs.
method DenseHMM uses kernelized transition probabilities and two optimization schemes.
result DenseHMM achieves superior performance and expressiveness compared to standard HMMs.

Latent topic models have been successfully applied as an unsupervised topic discovery technique in large document collections. With the proliferation of hypertext document collection such as the Internet, there has also been great interest in extending these approaches to hypertext [6, 9]. These approaches typically mo…

2012-06-13abs ↗pdf ↗

This paper analyzes DeepWalk and node2vec for community detection in large networks.

problem Community detection in large, sparse networks.
method Low-dimensional network embedding algorithms (DeepWalk and node2vec) applied to random walk segments.
result The performance of DeepWalk and node2vec in recovering communities depends on the length of random walk segments and sparsity of the network.

The paper classifies trades into types based on proximity and measures their impact on stock prices.

problem Understanding the impact of high-frequency trades on stock prices and their predictability.
method Classifies trades into five types based on proximity, measures conditional order imbalance (COI), and develops trading strategies.
result Strong positive correlations between contemporaneous returns and COIs, and positive associations with future returns for isolated trades.

GEMRank embeds users and items using co-occurrence relations for better collaborative filtering.

problem Lack of textual data for entity embedding in recommender systems.
method Uses profile co-occurrence for entity relations and factorization for embedding. Feeds embeddings into a neural network for predictions.
result Significantly outperforms baseline algorithms in various data sets.

Algorithm measures sentiment-based network risk in companies.

problem Understanding the relationship between news sentiment and stock price movements.
method Algorithm ranks companies based on news sentiment and co-occurrences, calculating individual and aggregated risks.
result The highest quarterly risk value correlates with a higher chance of stock price decline up to 70 days later.

IMPACC improves consensus clustering for bioinformatics data.

problem Consensus clustering's inefficiency and lack of interpretability for large-scale data.
method Ensemble minipatch co-occurrences, adaptive sampling of observations and features.
result Significantly improved accuracy and interpretability with substantial computational savings.

New methods for unsupervised learning of word and entity representations.

problem Learning distributed representations of words and entities from text and knowledge bases.
method MVLSA for words and NVSE for entities, both unsupervised learning methods.
result MVLSA and NVSE outperform state-of-the-art models in word and entity representation learning.

New DTMs model text evolution with Gaussian processes and scalable inference.

problem Challenges in modeling text evolution with continuous stochastic processes.
method Extended tractable priors to Gaussian processes and developed scalable inference methods.
result Found interesting patterns in large-scale datasets not accessible before.

The paper analyzes how news sentiment of companies can affect market movements.

problem Understanding how news sentiment impacts market performance and volatility.
method Applied NLP techniques to analyze news sentiment of 87 companies over 7 years.
result Strong media sentiment towards one company can indicate significant changes in sentiment towards related companies.