Developed DLCM for more accurate clustering of categorical data.
problem Restrictive conditional independence assumption in traditional LCMs.
method Bayesian Dependent Latent Class Model (DLCM) that allows conditional dependence.
result DLCMs are effective in applications with time series, overlapping items, and structural zeroes.
Detects and explains anomalies in non-i.i.d data with latent-class dependencies.
problem Detecting and explaining anomalies in data with latent-class dependencies.
method Derives SVDD method to handle latent-class dependency structure, provides probabilistic interpretation.
result Demonstrates effectiveness on real-world offshore data.
New algorithms for latent class analysis using regularized spectral clustering.
problem Identifying latent classes within populations from categorical data.
method Developed two new algorithms using a regularized Laplacian matrix to estimate latent classes.
result Our algorithms provide consistent latent class analysis under mild conditions and can accurately infer the number of latent classes.
A new model for latent class analysis with weighted responses.
problem Limitation of latent class model for real-world data with continuous or negative responses.
method Proposed a novel generative model, the weighted latent class model (WLCM).
result The proposed WLCM is more realistic and general than the latent class model.
Bayesian model identifies three types of travelers adapting to feedback.
problem Capturing adaptive, feedback-driven travel behavior in heterogeneous individuals.
method Latent Class Reinforcement Learning (LCRL) model with Variational Bayes estimation.
result Three distinct traveler classes identified: context-dependent, persistent exploitative, and exploratory.
The paper proposes a test to determine the number of latent classes in ordinal categorical data.
problem Determining the correct number of latent classes in latent class models with ordinal categorical data.
method The test statistic centers the largest singular value of a normalized residual matrix by a simple sample-size adjustment.
result The test statistic converges to zero under the null hypothesis and exceeds a fixed positive constant under an under-fitted alternative.
New bounds for contrastive learning handle domain shifts and generalization.
problem Domain shifts and generalization challenges in downstream tasks.
method Novel generalization bounds accounting for both domain shift and generalization.
result Performance of contrastively learned representations depends on statistical discrepancy between pretraining and downstream distributions.
New model for multi-layer categorical data improves latent class analysis.
problem Traditional latent class analysis for single-layer categorical data is insufficient for multi-layer data.
method Developed a multi-layer latent class model (multi-layer LCM) and three spectral methods for estimation.
result The debiased sum of Gram matrices method performs best in estimating latent classes.
ESRLCM clusters similar responses, more broadly than traditional models.
problem Clustering multivariate categorical data with common response patterns.
method Bayesian Equivalence Set Restricted Latent Class Model (ESRLCM).
result ESRLCM identifies clusters with similar item response probabilities.
New method classifies mixtures without labels, recovering latent classes.
problem Classification without reliable instance-level labels.
method Posterior simplex geometry for multiclass learning.
result Classifier trained on mixture identities recovers latent classes and their proportions.
This paper uses Factored Latent Analysis (FLA) to learn a factorized, segmental representation for observations of tracked objects over time. Factored Latent Analysis is latent class analysis in which the observation space is subdivided and each aspect of the original space is represented by a separate latent class mod…
StepMix estimates mixture models with covariates for social science applications.
problem Estimating latent classes with covariates in social science models.
method Pseudo-likelihood estimation using one-, two-, and three-step approaches.
result Unified framework for expectation-maximization subroutines.
Study improves choice model accuracy and heterogeneity representation using mixture models.
problem Improving prediction accuracy and heterogeneity representation in choice models.
method Semi-nonparametric Latent Class Choice Model with mixture models and EM algorithm.
result Mixture models enhance prediction accuracy and heterogeneity representation without sacrificing interpretability.
New method detects changes in complex models using hierarchical latent-class models.
problem Detecting abrupt transitions in high-dimensional or heterogeneous models.
method Hierarchical latent-class model with CRP and EM algorithm for continual learning.
result The method reliably infers the number of latent classes and performs CPD.
Improves labeling quality in machine learning with pairwise feedback.
problem Scalability and quality of labeled datasets in machine learning.
method Incorporates pairwise feedback into the programmatic creation of labeled datasets.
result Even a small number of pairwise feedback sources can substantially improve label quality.
The paper analyzes unsupervised learning using contrastive methods and introduces a theoretical framework.
problem Learning useful feature representations from unlabeled data.
method Introduces latent classes and contrasts similar vs. non-similar data points.
result Proves guarantees on the performance of learned representations on downstream tasks.
New optimization methods improve convergence of latent class model estimators.
problem Slow convergence of the EM algorithm in latent class model estimation.
method Transformed likelihood-based approach into constrained nonlinear optimization problem and applied quasi-Newton type methods.
result Proposed methods converge in fewer iterations and produce more accurate estimators.
VBphenoR uses variational Bayes for EHR-based patient phenotyping.
problem Phenotyping patients from EHR data for targeted treatments.
method Variational Bayes Gaussian Mixture Model (GMM) and logistic regression.
result Closed-form inference for efficient patient phenotype determination.
Model complexity is an important factor to consider when selecting among graphical models. When all variables are observed, the complexity of a model can be measured by its standard dimension, i.e. the number of independent parameters. When hidden variables are present, however, standard dimension might no longer be ap…
Study uses LCA to identify ARDS sub-phenotypes improving predictive models.
problem Complex and heterogeneous nature of ARDS makes early recognition difficult.
method Applied latent class analysis to identify sub-groups, then built predictive models.
result Significantly improved prediction performance for two sub-phenotypes of ARDS.
This research quantifies cross-sectoral inequalities using latent class analysis.
problem Addressing multiple and intersecting forms of inequality in various sectors.
method Innovative latent class analysis approach to quantify discrepancies.
result Significant discrepancies found among minority ethnic groups and between them and non-minority groups.
The main aim of this paper is to inspect the properties of survey based on households inflation expectations, conducted by Reserve Bank of India. It is theorized that the respondents answers are exaggerated by extreme response bias. Latent class analysis has been hailed as a promising technique for studying measurement…
In many countries information on expectations collected through consumer confidence surveys are used in macroeconomic policy formulation. Unfortunately, before doing so, the consistency of responses is often not taken into account, leading to biases creeping in and affecting the reliability of the indices hence created…
A new model estimates mixed memberships for categorical data with weighted responses.
problem Limited applicability of existing GoM model to weighted categorical data.
method Proposes Weighted Grade of Membership (WGoM) model, relaxing distribution constraints.
result WGoM can describe any response matrix with finite distinct elements.
Bayesian model enhances phenotype discovery in asthma EHRs.
problem Lack of interpretability in unsupervised learning phenotyping of EHR data.
method Operationalized a Bayesian latent class framework with clinical knowledge priors.
result Identified an asthma sub-phenotype with elevated eosinophil levels and allergy markers.
A large amount of observational data has been accumulated in various fields in recent times, and there is a growing need to estimate the generating processes of these data. A linear non-Gaussian acyclic model (LiNGAM) based on the non-Gaussianity of external influences has been proposed to estimate the data-generating …
We develop a personalized real time risk scoring algorithm that provides timely and granular assessments for the clinical acuity of ward patients based on their (temporal) lab tests and vital signs. Heterogeneity of the patients population is captured via a hierarchical latent class model. The proposed algorithm aims t…
We present a mixed multinomial logit (MNL) model, which leverages the truncated stick-breaking process representation of the Dirichlet process as a flexible nonparametric mixing distribution. The proposed model is a Dirichlet process mixture model and accommodates discrete representations of heterogeneity, like a laten…
In record linkage (RL), or exact file matching, the goal is to identify the links between entities with information on two or more files. RL is an important activity in areas including counting the population, enhancing survey frames and data, and conducting epidemiological and follow-up studies. RL is challenging when…
Contrastive learning performance doesn't degrade with more negative samples.
problem Theoretical and empirical evidence of negative samples hurting performance in contrastive learning.
method Simple theoretical setting and empirical support on CIFAR-10 and CIFAR-100 datasets.
result Contrastive learning performance does not degrade with the number of negative samples.
Combines ocean surface and interior data to study ocean dynamics.
problem Modeling local ocean currents and global climate patterns.
method Observation-driven framework using Latent-class regression.
result Improved prediction of vertical ocean temperature.
Prediction of the future trajectory of a disease is an important challenge for personalized medicine and population health management. However, many complex chronic diseases exhibit large degrees of heterogeneity, and furthermore there is not always a single readily available biomarker to quantify disease severity. Eve…
PCMC-Net uses neural networks to estimate transition rates in choice models, improving accuracy over traditional methods.
problem Inference limitations of traditional PCMC models when examples are scarce or new alternatives are observed.
method Amortized inference approach embedding PCMC definition into a neural network.
result Neural network outperforms feature engineered and machine learning models in airline booking prediction.
Probabilistic Latent Semantic Analysis is a novel statistical technique for the analysis of two-mode and co-occurrence data, which has applications in information retrieval and filtering, natural language processing, machine learning from text, and in related areas. Compared to standard Latent Semantic Analysis which s…
Modeling disease progression using irregular time intervals in EHRs.
problem Challenges in analyzing temporal data from EHRs.
method Developed a Markovian generative model using EHR data.
result Model accurately recovers underlying disease progression patterns from irregular time intervals.
Study compares clustering methods for mixed-type data.
problem Challenges in clustering mixed-type data.
method Distance-based (k-prototypes, PDQ, convex k-means), probabilistic (KAY-means, MBNs, LCM).
result KAMILA, LCM, and k-prototypes perform best.
Latent tree models are graphical models defined on trees, in which only a subset of variables is observed. They were first discussed by Judea Pearl as tree-decomposable distributions to generalise star-decomposable distributions such as the latent class model. Latent tree models, or their submodels, are widely used in:…
Unified framework for clustering with auxiliary data.
problem Clustering with datasets reflecting similar but different latent structures.
method Adaptive Transfer Clustering (ATC) algorithm that optimizes bias-variance decomposition.
result ATC proves optimal under Gaussian mixture model and shows transfer benefits.
We develop a Bayesian nonparametric approach to a general family of latent class problems in which individuals can belong simultaneously to multiple classes and where each class can be exhibited multiple times by an individual. We introduce a combinatorial stochastic process known as the negative binomial process (NBP)…
New method resolves causal heterogeneity by defining a resolution profile.
problem Causal subgroup analyses often oversimplify heterogeneity into a small number of groups.
method Introduces a resolution profile as a functional of the causal feature law, using Bayesian-bootstrap inference.
result Shows that the resolution profile is a continuous path with discontinuities at knots, providing integer-valued subgroup numbers.
More than two thirds of mental health problems have their onset during childhood or adolescence. Identifying children at risk for mental illness later in life and predicting the type of illness is not easy. We set out to develop a platform to define subtypes of childhood social-emotional development using longitudinal,…
Proposes a new method to handle noisy labels without needing accurate noise transition estimation.
problem Learning with noisy labels in the presence of class-conditional noise.
method Introduces a Latent Class-Conditional Noise (LCCN) model that embeds noise transition in a Bayesian framework and iteratively infers latent labels.
result Demonstrates superior performance compared to state-of-the-art methods on various noisy label datasets.
We face network data from various sources, such as protein interactions and online social networks. A critical problem is to model network interactions and identify latent groups of network nodes. This problem is challenging due to many reasons. For example, the network nodes are interdependent instead of independent o…
Paper reviews and compares NMF, PLSA, LBA, EMA, and LCA models.
problem Identifiability of latent models.
method Comparison and proof of identifiability.
result Identifiability of LBA, EMA, LCA, PLSA is unique if and only if NMF is unique.
We consider the problem of segmenting a large population of customers into non-overlapping groups with similar preferences, using diverse preference observations such as purchases, ratings, clicks, etc. over subsets of items. We focus on the setting where the universe of items is large (ranging from thousands to millio…
Paper adapts Bayesian Hui-Walter method for unlabeled data.
problem Lack of labeled data in machine learning.
method Adapted Hui-Walter paradigm for online, unlabeled data.
result Estimates performance metrics without labeled data.
Unified framework for variable selection in model-based clustering with missing data.
problem Challenges in identifying relevant variables and handling missing data in model-based clustering.
method Unified framework incorporating a data-driven penalty matrix and a mechanism for missingness modeling.
result Achieves both asymptotic consistency and selection consistency in the presence of missing data.
PCA outperforms random projections in retaining second order signals from latent groups.
problem Preserving second order structure in latent groups under unsupervised linear projections.
method Theoretical framework and quasi-exhaustive enumeration of projections.
result PCA outperforms random projections in retaining second order signals across a broad range of data-generating parameters.