Proposes a new model for clustering with heavier tails.
problem Clustering with heavy-tailed data.
method Finite mixture of skewed sub-Gaussian stable distributions, maximum likelihood estimation, EM algorithm.
result The proposed model can robustly handle heavy-tailed data.
New algorithms improve spectral clustering for finite mixture models.
problem Issues with EM algorithm in spectral clustering.
method Spectral decomposition and non-parametric bootstrap sampling.
result Improved convergence and avoidance of poor solutions.
New method estimates mixture model components efficiently.
problem Estimating the number of components in finite mixture models.
method Group-Sort-Fuse (GSF) procedure for simultaneous estimation of order and mixing measure.
result GSF achieves consistent estimation of true mixture order and n−1/2 convergence rate. TAMD prevents degeneracy in finite mixtures, offering strong guarantees but modest practical improvements.
problem Degeneracy in maximum likelihood estimation of finite mixtures.
method Transcendental regularization with analytic barrier functions.
result Strong theoretical guarantees (identifiability, consistency, robustness) but modest practical improvements.
Study singularity structures in finite mixtures affecting parameter estimation rates.
problem Understanding how singularity structures impact parameter estimation in finite mixtures.
method Developed a general framework to identify singularity structures in finite mixtures and studied their effects on convergence rates and minimax lower bounds.
result Established convergence rates for finite mixtures of skew-normal distributions, revealing complex asymptotic behaviors.
The paper tightens the upper bound on likelihood of finite mixtures.
problem Non-convexity and local maxima in finite mixture models.
method Convex optimization to find a tight upper bound on likelihood.
result A tight convex upper bound on the likelihood of a finite mixture can be computed.
Bayesian model tackles mixed-type data challenges.
problem Challenges in EM algorithm sensitivity, biomarker LOD, and variable importance.
method Bayesian finite mixture model with variable selection, LOD handling, and spike-and-slab prior.
result Improved parameter estimates and variable importance.
In this paper we present a new Bayesian network model for classification that combines the naive-Bayes (NB) classifier and the finite-mixture (FM) classifier. The resulting classifier aims at relaxing the strong assumptions on which the two component models are based, in an attempt to improve on their classification pe…
Paper detects gradual changes in cluster structure using MC fusion.
problem Detecting gradual changes in cluster structure over time.
method MC fusion for multiple mixture numbers, examining MC transition.
result Accurately captures cluster structure during transitional periods.
Spatially constrained Gaussian mixture models reduce covariance complexity.
problem High dimensionality in finite mixture models for spatial data.
method Spatial covariance constraint with only four free parameters.
result Improves clustering of multi-way spatial data and inference of spatial patterns.
Paper presents a more accurate method for nonparametric density estimation using FMMPL and SIR.
problem Improving nonparametric density estimation for complex datasets.
method Finite mixture model of nonparametric density estimation using sampling importance resampling.
result FMMPL provides more accurate results with less space complexity.
Unified clustering framework using optimal transport with regularization.
problem Inferring Finite Mixture Models from discrete data.
method Formulates optimal transport problem with entropic regularization, unifying hard and soft clustering.
result Shows benefits of using λ>1 for improved inference performance and λo0 for better classification. Study on Dirichlet process mixtures for clustering consistency.
problem Consistency of clustering with Dirichlet process mixtures.
method Analysis of posterior distribution as sample size increases, focusing on consistency for the number of clusters.
result Consistency for the number of clusters can be achieved with a properly adapted concentration parameter in a Bayesian setting.
DFMR improves robustness of learning finite mixture models in distributed settings.
problem Learning finite mixture models in distributed settings with Byzantine failures.
method DFMR leverages pairwise L2 distances to filter and retain local estimates, ensuring robust aggregation.
result DFMR achieves optimal convergence rate and asymptotic equivalence to global maximum likelihood estimate.
Proposes a robust FMR model for handling sample heterogeneity.
problem Handling sample heterogeneity with a single regression model.
method Clusters samples and jointly models multiple incomplete mixed-type targets.
result Achieves state-of-the-art performance on synthetic and real-world data.
A fast Modal EM algorithm for Gaussian mixtures.
problem Clustering with Gaussian mixtures.
method Modal EM algorithm for Gaussian mixtures.
result High flexibility in various clustering contexts.
The paper proposes a method to calibrate evidential clustering using bootstrapped finite mixture models.
problem Representing uncertainty in cluster membership using Dempster-Shafer mass functions.
method Constructing Dempster-Shafer mass functions by bootstrapping finite mixture models, computing confidence intervals, and calibrating the evidential partition.
result The proposed method calibrates the evidential partition such that the belief and plausibility degrees approximate the true probabilities with high confidence.
Finite mixtures of regression models offer a flexible framework for investigating heterogeneity in data with functional dependencies. These models can be conveniently used for unsupervised learning on data with clear regression relationships. We extend such models by imposing an eigen-decomposition on the multivariate …
New method selects FMM components via variational Bayes.
problem Selecting the correct number of components in finite mixture models.
method Variational Bayes with mean-field approximation.
result Consistency of model selection based on ELBO maximization.
New model clusters discrete time series data.
problem Handling discreteness and time series properties in data.
method Finite mixture model with INAR type models.
result Demonstrated clustering on real data.
FMM fails to accurately determine the number of components even with consistent posterior.
problem Determining the number of subpopulations in a data set using FMM.
method Analysis of FMM component-count posterior under model misspecification.
result FMM component-count posterior diverges under model misspecification, contrary to intuition.
A new sampler speeds up Bayesian mixture models.
problem Sampling from Bayesian finite mixture models is slow and hard.
method Introduces a non-reversible sampling scheme for Bayesian finite mixture models.
result The new sampler outperforms classical samplers in many scenarios, especially during convergence.
SSLfmm package improves semi-supervised learning by incorporating informative missingness in finite mixture models.
problem Improving semi-supervised learning with informative missingness in datasets.
method Estimates Bayes' classifier under a finite mixture model with MCAR and MAR missingness mechanisms.
result The classifier trained on partially labelled data can achieve lower misclassification rates than supervised methods.
We propose a kernel method to identify finite mixtures of nonparametric product distributions. It is based on a Hilbert space embedding of the joint distribution. The rank of the constructed tensor is equal to the number of mixture components. We present an algorithm to recover the components by partitioning the data p…
Improved convergence rates for MLE in mixture models using penalized log-likelihood.
problem Convergence rates for MLE in finite mixture models.
method Penalizing log-likelihood to discourage vanishing mixing weights, using Wasserstein distance and new loss functions.
result Improved convergence rates for some mixture components, faster than traditional methods.
NMDR estimates complex mixtures of distributions efficiently.
problem Estimating complex finite mixtures of distributions in high-dimensional settings.
method Flexible additive predictors, neural networks, and deep learning optimizers.
result Competitive performance in complex scenarios compared to existing approaches.
A new method uses Mean Field Games to optimize mixture models of Bernoulli and categorical distributions.
problem Optimizing parameters of finite mixture models of Bernoulli and categorical distributions.
method Mean Field Games theory applied to multi-population systems.
result The Mean Field Games approach provides a method to compute mixture model parameters.
In many applications, a finite mixture is a natural model, but it can be difficult to choose an appropriate number of components. To circumvent this choice, investigators are increasingly turning to Dirichlet process mixtures (DPMs), and Pitman-Yor process mixtures (PYMs), more generally. While these models may be well…
A new clustering model for sublinearly growing cluster sizes.
problem Clustering models assume clusters grow linearly with data size, but this is not always desirable.
method Defines microclustering property and introduces a new model.
result The new model yields clusters whose sizes grow sublinearly with data size.
New method uses dendrograms for better mixture model selection and clustering.
problem Selecting the correct number of components in finite mixture models.
method Hierarchical clustering tree derived from overfitted latent mixing measures.
result Consistently selects the true number of mixing components and optimal convergence rate for parameter estimation.
Model-based clustering imposes a finite mixture modelling structure on data for clustering. Finite mixture models assume that the population is a convex combination of a finite number of densities, the distribution within each population is a basic assumption of each particular model. Among all distributions that have …
New method reduces mixture model evaluation cost for large models.
problem Computational infeasibility of evaluating all mixture components.
method Combining EM and Metropolis-Hastings for stochastic sampling.
result Significantly reduced computational cost for large models.
Improved VB algorithm for NIG mixtures outperforms Gaussian mixtures for non-Gaussian data.
problem Clustering non-Gaussian data, especially heavy-tailed and asymmetric.
method Proposed an improved VB algorithm for NIG mixture models and extended Dirichlet process mixture models.
result Outperforms Gaussian mixtures and existing NIG mixture models, especially for highly non-normative data.
The paper explores strong identifiability and parameter learning in regression models with heterogeneous responses.
problem Understanding heterogeneity in data populations through conditional distributions of a response variable.
method Investigation of strong identifiability, convergence rates, and posterior contraction behavior in finite mixture of regression models.
result Theoretical findings on conditions for strong identifiability and rates of convergence in regression mixture models.
The paper proposes a parallelizable clustering method for multivariate data.
problem The standard model-based clustering method assumes the same number of clusters per margin, which is often unrealistic.
method Developed a finite mixture model per margin with different numbers of clusters, and used a game-inspired algorithm to cluster multivariate data.
result The proposed method shows good performance in various scenarios and real datasets.
New EM algorithms for weighted-data clustering improve audio-visual scene analysis.
problem Improving clustering of weighted data in heterogeneous environments.
method Proposed weighted-data Gaussian mixture model and two EM algorithms.
result Validation shows improved clustering in audio-visual scenes.
A new theory explains why ungrammatical sentences are read faster.
problem Why do ungrammatical sentences sometimes sound grammatical?
method Used hierarchical Bayesian mixture models to analyze reading times from 10 studies.
result Feature overwriting explains faster reading times better than other theories.
Study finds the minimum number of finite Gaussian mixtures for best approximation.
problem Finding the minimum number of finite Gaussian mixtures for best approximation.
method Local moment matching for upper bound and spectral analysis for lower bound.
result Corrects a previous lower bound in the case of Gaussian mixing distributions.
This article proposes a method to quantify the structure of a bipartite graph using a network entropy per link. The network entropy of a bipartite graph with random links is calculated both numerically and theoretically. As an application of the proposed method to analyze collective behavior, the affairs in which parti…
Method estimates mixture models without normalization.
problem Estimating mixture models with intractable normalization.
method Extends noise contrastive estimation (NCE) for mixture models.
result Probabilistic clustering using deep representations.
Model captures context-dependent neural correlations using Poisson mixtures.
problem Capturing context-dependent noise correlations in neural populations.
method Conditional finite mixtures of Poisson distributions, cross-validation for dimensionality, EM algorithm.
result Model successfully captures stimulus-dependent correlations in V1 neuron responses.
How to forecast next year's portfolio-wide credit default rate based on last year's default observations and the current score distribution? A classical approach to this problem consists of fitting a mixture of the conditional score distributions observed last year to the current score distribution. This is a special (…
Proposes FMC for fair clustering with independent parameters.
problem Finding clusters with balanced sensitive attribute proportions.
method Model-based clustering using finite mixture model with mini-batch learning.
result FMC scales up easily and can handle non-metric data.
The paper approximates CARMA models for option pricing.
problem Approximating the transition density of CARMA(p, q) models.
method Using Gauss-Laguerre quadrature and time changed Brownian Motion.
result Provides an analytical formula for option prices.
Context-aware learning boosts generative model performance.
problem Improving generative model performance with contextual information.
method Extended finite mixture models with embedded context variables, using expectation-maximization (EM) for maximum-likelihood estimation.
result Contextual assistance improves estimation precision, standard errors, and classification accuracy.
Identifying components and estimating mixing weights in unlabeled finite mixtures under marginal independence.
problem Identifying components and estimating mixing weights in unlabeled finite mixtures.
method Proving structural results and extending them to observable mixtures.
result Identifying components and estimating mixing weights under marginal independence.
We propose a method based on finite mixture models for classifying a set of observations into number of different categories. In order to demonstrate the method, we show how the component densities for the mixture model can be derived by using the maximum entropy method in conjunction with conservation of Pythagorean m…
Bayesian approach learns nonparametric mixture components from heterogeneous data.
problem Realistic modeling of heterogeneous data populations with nonparametric mixture components.
method Bayesian nonparametric modeling using Dirichlet process mixture priors.
result Posterior contraction rates for component densities are nearly polynomial, improving over deconvolution methods.