GraphSTONE uses topic models to capture graph structures, improving GCN performance.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
While generative models such as Latent Dirichlet Allocation (LDA) have proven fruitful in topic modeling, they often require detailed assumptions and careful specification of hyperparameters. Such model complexity issues only compound when trying to generalize generative models to incorporate human input. We introduce …
We reduce a broad class of machine learning problems, usually addressed by EM or sampling, to the problem of finding the extremal rays spanning the conical hull of a data point set. These "anchors" lead to a global solution and a more interpretable model that can even outperform EM and sampling on generalizatio…
Learning node embeddings that capture a node's position within the broader graph structure is crucial for many prediction tasks on graphs. However, existing Graph Neural Network (GNN) architectures have limited power in capturing the position/location of a given node with respect to all other nodes of the graph. Here w…
Separable Non-negative Matrix Factorization (SNMF) is an important method for topic modeling, where "separable" assumes every topic contains at least one anchor word, defined as a word that has non-zero probability only on that topic. SNMF focuses on the word co-occurrence patterns to reveal topics by two steps: anchor…
Originally designed to model text, topic modeling has become a powerful tool for uncovering latent structure in domains including medicine, finance, and vision. The goals for the model vary depending on the application: in some cases, the discovered topics may be used for prediction or some other downstream task. In ot…
Deep neural networks, while generalize well, are known to be sensitive to small adversarial perturbations. This phenomenon poses severe security threat and calls for in-depth investigation of the robustness of deep learning models. With the emergence of neural networks for graph structured data, similar investigations …
Language is dynamic, constantly evolving and adapting with respect to time, domain or topic. The adaptability of language is an active research area, where researchers discover social, cultural and domain-specific changes in language using distributional tools such as word embeddings. In this paper, we introduce the gl…
LargeMvC-Net improves scalability of multi-view clustering.
GraphReach improves GNN performance by incorporating node positions.
A new framework for graph representation learning.
A framework evaluates the impact of different modules in graph contrastive learning.
Sparse subspace clustering (SSC) is one of the current state-of-the-art methods for partitioning data points into the union of subspaces, with strong theoretical guarantees. However, it is not practical for large data sets as it requires solving a LASSO problem for each data point, where the number of variables in each…
Can neural networks learn to compare graphs without feature engineering? In this paper, we show that it is possible to learn representations for graph similarity with neither domain knowledge nor supervision (i.e.\ feature engineering or labeled graphs). We propose Deep Divergence Graph Kernels, an unsupervised method …
Reinterprets DNF as a deep generative LDA model for complex data.
Despite many years of research into latent Dirichlet allocation (LDA), applying LDA to collections of non-categorical items is still challenging. Yet many problems with much richer data share a similar structure and could benefit from the vast literature on LDA. We propose logistic LDA, a novel discriminative variant o…
In latent Dirichlet allocation (LDA), topics are multinomial distributions over the entire vocabulary. However, the vocabulary usually contains many words that are not relevant in forming the topics. We adopt a variable selection method widely used in statistical modeling as a dimension reduction tool and combine it wi…
Hyper-parameters play a major role in the learning and inference process of latent Dirichlet allocation (LDA). In order to begin the LDA latent variables learning process, these hyper-parameters values need to be pre-determined. We propose an extension for LDA that we call 'Latent Dirichlet allocation Gibbs Newton' (LD…
Simple Deep LDA models achieve accuracy competitive with softmax baselines.
This tutorial explains Linear Discriminant Analysis (LDA) and Quadratic Discriminant Analysis (QDA) as two fundamental classification methods in statistical and probabilistic learning. We start with the optimization of decision boundary on which the posteriors are equal. Then, LDA and QDA are derived for binary and mul…
E-LDA offers faster, interpretable LDA topic models.
Adaptive anchor methods improve multi-modal learning by balancing intra-modal and inter-modal information.
LDA-GO improves LDA for high-dimensional data via gradient optimization.
Novel framework improves GNN uncertainty estimates under distribution shifts.
Latent Dirichlet Allocation (LDA) is a popular topic modeling technique for discovery of hidden semantic architecture of text datasets, and plays a fundamental role in many machine learning applications. However, like many other machine learning algorithms, the process of training a LDA model may leak the sensitive inf…
LDA improves image classification accuracy with fewer features.
A new LDA variant improves multi-label classification performance.
Private anchors affect how information is communicated and can improve or distort transmission.
MILDA uses unlabelled data to compute LDA projections.
For organizing large text corpora topic modeling provides useful tools. A widely used method is Latent Dirichlet Allocation (LDA), a generative probabilistic model which models single texts in a collection of texts as mixtures of latent topics. The assignments of words to topics rely on initial values such that general…
Anchors explain text model decisions by highlighting key words.
Derives VMP for LDA, simplifying inference for topic modeling.
We introduce a novel approach for estimating Latent Dirichlet Allocation (LDA) parameters from collapsed Gibbs samples (CGS), by leveraging the full conditional distributions over the latent variable assignments to efficiently average over multiple samples, for little more computational cost than drawing a single addit…
Anchoring is a term used in psychology to describe the common human tendency to rely too heavily (anchor) on one piece of information when making decisions. A trading algorithm inspired by biological motors, introduced by L. Gil\cite{Gil}, is suggested as a testing ground for anchoring in financial markets. An exact so…
LDA-XGB1 balances fairness and accuracy in lending models.
Δ-UQ uses anchoring to estimate uncertainty in models.
Paper explores anchoring for vision models, improving generalization and safety.
LDTA expands LDA's topic modeling capacity with tree-structured priors.
This paper defines less discriminatory algorithms and explores their feasibility.
The sBIC outperforms other model selection criteria in LDA topic modeling.
DNLL loss improves deep LDA accuracy and consistency.
Improved topic modeling captures temporal relationships in speech.
A new method improves label propagation for unsupervised domain adaptation.
Market stability depends on a fundamental value anchor, not price crashes.
SWRLDA improves LDA for multi-class classification with edge classes.
A new LDA model with covariates for mixed-membership clusters.
This work merges 3-anchored bundles into 3-Lie algebroids.
A new noise model for preferential Bayesian optimization using user anchors.