We derive relations between theoretical properties of restricted Boltzmann machines (RBMs), popular machine learning models which form the building blocks of deep learning models, and several natural notions from discrete mathematics and convex geometry. We give implications and equivalences relating RBM-representable …
New insights into identifying mixtures of product distributions using Hadamard extensions.
problem Identifying mixtures of product distributions on binary variables.
method Analysis of Hadamard extensions of matrix products.
result Conditions for full column rank of Hadamard extensions.
Study learns mixtures of smooth product distributions from samples.
problem Learning mixtures of non-parametric product distributions.
method Two-stage approach using identifiability properties of tensor decomposition and signal processing techniques.
result Recovery of component distributions under a smoothness condition.
Algorithm identifies sources in product distributions with improved complexity.
problem Identifying sources in mixtures of product distributions.
method Approximate multilinear moments input, 2^{O(k^2)} n^{O(k)} operations.
result First explicit bound on computational complexity of source identification.
Algorithm completes symmetric tensors from few entries, learns product mixtures.
problem Learning product mixtures over the hypercube from incomplete data.
method Tensor completion algorithm applied to matrix completion for adversarially missing entries.
result Recover distributions with many centers in polynomial/quasi-polynomial time.
We propose a kernel method to identify finite mixtures of nonparametric product distributions. It is based on a Hilbert space embedding of the joint distribution. The rank of the constructed tensor is equal to the number of mixture components. We present an algorithm to recover the components by partitioning the data p…
New method reduces mixture model evaluation cost for large models.
problem Computational infeasibility of evaluating all mixture components.
method Combining EM and Metropolis-Hastings for stochastic sampling.
result Significantly reduced computational cost for large models.
Two models of binary variables are shown to represent the same distributions.
problem Comparing two models of binary variables.
method Semi-algebraic description and maximum likelihood estimates.
result Two models represent the same set of distributions.
We study the problem of learning a distribution from samples, when the underlying distribution is a mixture of product distributions over discrete domains. This problem is motivated by several practical applications such as crowd-sourcing, recommendation systems, and learning Boolean functions. The existing solutions e…
Improved sample and time complexity for identifying mixtures of product distributions.
problem Identifying a mixture of k product distributions from statistics. method Combining robust tensor decomposition and Hadamard extensions to bound the condition number of key matrices.
result Achieved sample complexity and run-time complexity of (1/ζ)O(k) for n≥2k−1. A new sampling method accelerates inference in discrete probabilistic models.
problem Slow convergence of Markov chain Monte Carlo algorithms in discrete probabilistic models.
method Proposes a mixture of product distributions using semigradient information to accelerate convergence.
result Combining the new sampler with existing ones improves inference in various models.
Identifying components and estimating mixing weights in unlabeled finite mixtures under marginal independence.
problem Identifying components and estimating mixing weights in unlabeled finite mixtures.
method Proving structural results and extending them to observable mixtures.
result Identifying components and estimating mixing weights under marginal independence.
A new model uses low-rank DPPs to improve product recommendation.
problem Scalability issues in DPP models for large datasets.
method Low-rank DPP mixture model with MCMC learning.
result Substantially better predictive performance than single DPP models.
New criteria judge combinatorial equivalence of polytopes to products of simplices.
problem Determining combinatorial equivalence of polytopes to products of simplices.
method Combination of combinatorial, geometric, and topological conditions inspired by toric topology.
result New criteria for judging combinatorial equivalence of polytopes to products of simplices.
Feature selection improves learning from mixtures of discrete variables.
problem Learning mixtures of discrete random variables, especially in unreliable crowdsourcing.
method Algorithm based on mutual information to rank workers and induce a low-order statistical model.
result Improvement in real data sets can be substantial.
Proposes a new prior for deep generative models to capture latent properties.
problem Complex non-linear relationships between data and latent properties.
method Factorial mixture prior with Gaussian mixture models for quantization.
result Empirically evaluated method for learning discrete properties in unsupervised or semi-supervised settings.
New method for summarizing Bayesian mixture models using sliced Wasserstein distances.
problem Estimating the mixing measure in nonparametric Bayesian mixture models.
method Decision-theoretic approach using sliced Wasserstein distances for Gaussian mixtures.
result Effective estimation of the mixing measure and mixture density.
PIMA autoencoders discover shared features in multimodal scientific data.
problem Discovering shared information in high-throughput scientific datasets.
method Physics-informed multimodal autoencoders (PIMA) with Gaussian mixture prior and product of experts formulation.
result Accurate cross-modal inference between images and mechanical stress-strain response in lattice metamaterials.
Deep sum-product networks learn faster than shallow models.
problem The speed of parameter optimization in sum-product networks.
method Theoretical analysis and empirical experiments on overparameterized sum-product networks.
result Gradient-based optimization in deep sum-product networks is equivalent to gradient ascent with adaptive and time-varying learning rates and additional momentum terms.
Adversarial MoE learns category-specific models for product search.
problem Variations in product features and importance across categories.
method Mixture of Experts with adversarial regularization and soft gating constraints.
result Improved clustering of gate output vectors and shared experts among similar categories.
Novel approach for estimating joint probability densities using tensor decompositions and dictionaries.
problem Estimating joint probability densities of mixed discrete and continuous variables.
method Low-rank tensor decomposition combined with dictionary learning.
result Better classification and lower error rates compared to existing methods.
Bayesian networks with hidden variables help identify causal relationships obscured by confounding.
problem Identifying causal relationships obscured by unobserved confounders.
method Use finite k-mixtures of Bayesian networks with hidden variables to recover the joint probability distribution and identify causal relationships. result First algorithm to learn mixtures of non-empty DAGs, recovering identifiable causal relationships.
Characterizes exchangeable feature allocations with specific probability functions.
problem Tackles the characterization of exchangeable feature allocations with product-form probability functions.
method Characterizes the class of exchangeable feature allocations using a countable matrix, sequences of weights, and a consistency condition.
result Provides a characterization of the Indian Buffet Process and Beta--Bernoulli model as the only consistent exchangeable feature allocations with product form.
We introduce RNADE, a new model for joint density estimation of real-valued vectors. Our model calculates the density of a datapoint as the product of one-dimensional conditionals modeled using mixture density networks with shared parameters. RNADE learns a distributed representation of the data, while having a tractab…
PSD models simplify probability density estimation.
problem Effective modeling of probability densities for inference.
method Positive semi-definite (PSD) models for non-negative functions.
result PSD models efficiently support product and sum rules.
In this paper we show that very large mixtures of Gaussians are efficiently learnable in high dimension. More precisely, we prove that a mixture with known identical covariance matrices whose number of components is a polynomial of any fixed degree in the dimension n is polynomially learnable as long as a certain non-d…
This paper uses SPNs with GPs to efficiently model complex data.
problem Inference cost and memory issues in Gaussian processes.
method Integrating Gaussian processes into sum-product networks.
result The model efficiently learns input-dependent parameters and hyper-parameters.
New EM algorithm learns from experts in stagewise fashion.
problem Learning mixtures of product distributions from noisy data.
method Stagewise learning approach, starting from a single mixture class, developing 'experts' in a mutual information criterion-driven fashion.
result Stagewise EM outperforms other initialization techniques for crowdsourcing and neurosciences applications.
Enhances neural forecasting for hierarchically organized time series data.
problem Probabilistic coherent forecasting of time series data across different levels of aggregation.
method Proposes a coherent multivariate mixture output for neural forecasting architectures, optimizing with a composite likelihood objective.
result 13.2% average accuracy improvements on most datasets compared to state-of-the-art baselines.
Corrected whitening restores orthogonality in high-dimensional spherical Gaussian mixtures.
problem In high-dimensional data, standard whitening fails to preserve orthogonality of mixture means.
method Derived exact limits for whitened means dot products using random matrix theory, constructed a corrected whitening matrix.
result Corrected whitening allows for improved estimation of spherical Gaussian mixtures in the large-dimensional regime.
BN^2MF identifies unknown exposure patterns in environmental mixtures.
problem Identifying unknown exposure patterns in environmental mixtures.
method Bayesian non-parametric non-negative matrix factorization (BN^2MF) with non-negative continuous priors and a non-parametric sparse prior.
result Estimates patterns of chemical exposures without specifying the number of patterns.
The paper studies multi-view representation learning with generalization guarantees and a new regularizer.
problem Distributed multi-view representation learning with correct estimation at a decoder.
method Generalization bounds using relative entropy and MDL, data-dependent Gaussian mixture priors.
result Data-dependent Gaussian mixture priors lead to good performance and outperform existing methods.
Method estimates joint probability density from samples using low-rank decomposition and random projections.
problem Estimating joint probability density from limited samples.
method Low-rank tensor decomposition, dictionaries, and Radon transforms.
result Algorithm outperforms previous methods in estimating synthetic probability densities.
Hydra boosts efficiency for long-context reasoning in resource-constrained settings.
problem Quadratic complexity of transformers limits long-context reasoning in resource-constrained systems.
method Hydra uses a modular architecture with adaptive routing between sparse global attention, mixture-of-experts, and dual memories.
result Hydra achieves significant throughput and accuracy improvements for long-context reasoning.
New method selects FMM components via variational Bayes.
problem Selecting the correct number of components in finite mixture models.
method Variational Bayes with mean-field approximation.
result Consistency of model selection based on ELBO maximization.
Algorithm learns data structure in real-time with outliers and change points.
problem Sequential online prediction in the presence of outliers and change points.
method INTEL algorithm using WGPs and POE model for real-time structure learning.
result Significantly better performance than benchmarks in real datasets.
Method proposed for pricing insurance products covering both foreseeable and unforeseeable risks.
problem Pricing insurance products that include unforeseeable risks.
method Mixed Poisson process with Bayesian setup and linear exponential family distributions.
result Bayesian premiums are more reactive to claim trends than traditional ones.
The paper corrects for node degree in spectral clustering using random walk Laplacian.
problem Node degree heterogeneity in spectral clustering.
method Graph spectral embedding using the random walk Laplacian.
result The embedding provides uniformly consistent estimates of degree-corrected latent positions.
MFVI mode collapse explained; RoVI proposed to mitigate.
problem Mode collapse in MFVI for mixture distributions.
method Introducing ε-separateness, deriving bounds, proposing RoVI.
result MFVI optimizers collapse to a single component when components are ε-separated.
A new method splits data into independent mechanisms for better generative models.
problem Training a single model to capture the overall distribution of data.
method A competitive training procedure using mixtures of independent deep generative models and discriminators.
result The approach splits the training distribution in a sensible way and improves the quality of generated samples.
We prove a central limit theorem for the components of the largest eigenvectors of the adjacency matrix of a finite-dimensional random dot product graph whose true latent positions are unknown. In particular, we follow the methodology outlined in \citet{sussman2012universally} to construct consistent estimates for the …
Tensorial Mixture Models combine tractable structure with rich distribution representation.
problem Lack of tractable marginalization in generative models.
method Derived from tensor analysis, TMMs use simple convolutional networks and leverage theoretical analyses.
result Tensorial Mixture Models deliver state-of-the-art accuracies in classification tasks with missing data.
New algorithm learns mixtures of subcubes efficiently, with applications to decision trees.
problem Learning mixtures of subcubes over binary space.
method Higher-order multilinear moments for nO(logk)-time learning. result Polynomial dependence on 1/ε for decision trees with stochastic transitions. Modeling the Drosophila connectome using semiparametric spectral methods.
problem Understanding the structure and function of the Drosophila mushroom body network.
method Semiparametric spectral modeling, latent structure model (LSM), Gaussian mixture modeling (GMM), adjacency spectral embedding (ASE).
result Captures latent connectome structure and elucidates neuronal properties.
Efficient algorithms learn high-dimensional distributions robustly, independent of dimensionality.
problem Learning high-dimensional distributions in the presence of adversarial corruption.
method Developed computationally efficient algorithms with dimension-independent error guarantees.
result Achieved error independent of dimension and nearly-linear in corrupted samples fraction.
ARMDN forecasts retail demand by modeling associative factors and trends.
problem Accurately forecasting demand for e-retailers with many associative factors and non-stationary shifts.
method ARMDN combines feature embeddings, MLP, LSTM, and mixture of Gaussian distributions to model demand.
result ARMDN outperforms existing methods in forecasting accuracy.
Paper proves identifiability and consistency of hub model for network inference.
problem Identifying network structure from group behavior.
method Hub model and variants, proving identifiability and consistency under mild conditions.
result Identifiability and estimation consistency of hub model and its variants proved.
XSPNs combine SPNs and MEVMs for efficient inference in data with repeated parts.
problem Efficient inference in data with repeated interchangeable parts.
method Introducing Exchangeability-Aware Sum-Product Networks (XSPNs) that combine SPNs and MEVMs.
result XSPNs can be more accurate than conventional SPNs when data contains repeated parts.