We derive relations between theoretical properties of restricted Boltzmann machines (RBMs), popular machine learning models which form the building blocks of deep learning models, and several natural notions from discrete mathematics and convex geometry. We give implications and equivalences relating RBM-representable …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New insights into identifying mixtures of product distributions using Hadamard extensions.
We study the problem of learning a mixture model of non-parametric product distributions. The problem of learning a mixture model is that of finding the component distributions along with the mixing weights using observed samples generated from the mixture. The problem is well-studied in the parametric setting, i.e., w…
Algorithm identifies sources in product distributions with improved complexity.
We propose a kernel method to identify finite mixtures of nonparametric product distributions. It is based on a Hilbert space embedding of the joint distribution. The rank of the constructed tensor is equal to the number of mixture components. We present an algorithm to recover the components by partitioning the data p…
Improved sample and time complexity for identifying mixtures of product distributions.
We study the problem of learning a distribution from samples, when the underlying distribution is a mixture of product distributions over discrete domains. This problem is motivated by several practical applications such as crowd-sourcing, recommendation systems, and learning Boolean functions. The existing solutions e…
We give an algorithm for completing an order- symmetric low-rank tensor from its multilinear entries in time roughly proportional to the number of tensor entries. We apply our tensor completion algorithm to the problem of learning mixtures of product distributions over the hypercube, obtaining new algorithmic result…
We compare two statistical models of three binary random variables. One is a mixture model and the other is a product of mixtures model called a restricted Boltzmann machine. Although the two models we study look different from their parametrizations, we show that they represent the same set of distributions on the int…
New method reduces mixture model evaluation cost for large models.
We consider the problem of inference in discrete probabilistic models, that is, distributions over subsets of a finite ground set. These encompass a range of well-known models in machine learning, such as determinantal point processes and Ising models. Locally-moving Markov chain Monte Carlo algorithms, such as the Gib…
Novel approach for estimating joint probability densities using tensor decompositions and dictionaries.
Bayesian networks with hidden variables help identify causal relationships obscured by confounding.
We study high-dimensional distribution learning in an agnostic setting where an adversary is allowed to arbitrarily corrupt an -fraction of the samples. Such questions have a rich history spanning statistics, machine learning and theoretical computer science. Even in the most basic settings, the only known…
Method proposed for pricing insurance products covering both foreseeable and unforeseeable risks.
Feature selection can facilitate the learning of mixtures of discrete random variables as they arise, e.g. in crowdsourcing tasks. Intuitively, not all workers are equally reliable but, if the less reliable ones could be eliminated, then learning should be more robust. By analogy with Gaussian mixture models, we seek a…
Identifying components and estimating mixing weights in unlabeled finite mixtures under marginal independence.
Method estimates joint probability density from samples using low-rank decomposition and random projections.
New distributions allow greedy arm selection in sparse bandit problems.
Study calculates tail risk for various mixture distributions.
We introduce RNADE, a new model for joint density estimation of real-valued vectors. Our model calculates the density of a datapoint as the product of one-dimensional conditionals modeled using mixture density networks with shared parameters. RNADE learns a distributed representation of the data, while having a tractab…
A common assumption in causal modeling posits that the data is generated by a set of independent mechanisms, and algorithms should aim to recover this structure. Standard unsupervised learning, however, is often concerned with training a single model to capture the overall distribution or aspects thereof. Inspired by c…
MFVI mode collapse explained; RoVI proposed to mitigate.
Adversarial MoE learns category-specific models for product search.
Stochastic gradient descent converges to universal limits in high dimensions.
The paper studies multi-view representation learning with generalization guarantees and a new regularizer.
HessFormer enables distributed Hessian computation for large models.
The performance of EM in learning mixtures of product distributions often depends on the initialization. This can be problematic in crowdsourcing and other applications, e.g. when a small number of 'experts' are diluted by a large number of noisy, unreliable participants. We develop a new EM algorithm that is driven by…
Understanding proper distance measures between distributions is at the core of several learning tasks such as generative models, domain adaptation, clustering, etc. In this work, we focus on mixture distributions that arise naturally in several application domains where the data contains different sub-populations. For …
Sum-Product Networks (SPNs) can be regarded as a form of deep graphical models that compactly represent deeply factored and mixed distributions. An SPN is a rooted directed acyclic graph (DAG) consisting of a set of leaves (corresponding to base distributions), a set of sum nodes (which represent mixtures of their chil…
We characterize the class of exchangeable feature allocations assigning probability to a feature allocation of individuals, displaying features with counts for these features. Each element of this class is parametrized by a countable matrix …
We study the mixtures of factorizing probability distributions represented as visible marginal distributions in stochastic layered networks. We take the perspective of kernel transitions of distributions, which gives a unified picture of distributed representations arising from Deep Belief Networks (DBN) and other netw…
Heavy-tailed distributions are widely used in robust mixture modelling due to possessing thick tails. As a computationally tractable subclass of the stable distributions, sub-Gaussian -stable distribution received much interest in the literature. Here, we introduce a type of expectation maximization algorithm that e…
New method selects FMM components via variational Bayes.
In this paper we show that very large mixtures of Gaussians are efficiently learnable in high dimension. More precisely, we prove that a mixture with known identical covariance matrices whose number of components is a polynomial of any fixed degree in the dimension n is polynomially learnable as long as a certain non-d…
Consistent estimator for mixtures of nonparametric elliptical distributions helps cluster analysis.
Paper proposes a new Wasserstein distance for mixtures of radially contoured distributions.
NMDR estimates complex mixtures of distributions efficiently.
Deep neural networks converge to Gaussian mixtures as layer width increases.
The paper examines risk aggregation under mixtures of marginals, finding that more homogeneous distributions lead to larger uncertainty.
Optimal transport for vector Gaussian mixtures improves efficiency and structure preservation.
A new method uses Mean Field Games to optimize mixture models of Bernoulli and categorical distributions.
Mixture modelling involves explaining some observed evidence using a combination of probability distributions. The crux of the problem is the inference of an optimal number of mixture components and their corresponding parameters. This paper discusses unsupervised learning of mixture models using the Bayesian Minimum M…
We study the problem of partitioning a small sample of individuals from a mixture of product distributions over a Boolean cube according to their distributions. Each distribution is described by a vector of allele frequencies in . Given two distributions, we use to denote the average $\el…
Proposes a new model for clustering with heavier tails.
We give several new criteria to judge whether a simple convex polytope in a Euclidean space is combinatorially equivalent to a product of simplices. These criteria are mixtures of combinatorial, geometrical and topological conditions that are inspired by the ideas from toric topology.
We introduce the problem of learning mixtures of subcubes over , which contains many classic learning theory problems as a special case (and is itself a special case of others). We give a surprising -time learning algorithm based on higher-order multilinear moments. It is not possible to l…
Langevin Dynamics fails to sample from mixture distributions efficiently.