Tensorial Mixture Models combine tractable structure with rich distribution representation.
problem Lack of tractable marginalization in generative models.
method Derived from tensor analysis, TMMs use simple convolutional networks and leverage theoretical analyses.
result Tensorial Mixture Models deliver state-of-the-art accuracies in classification tasks with missing data.
Unified tractability conditions for various compositional inference queries.
problem Analyzing tractability of probabilistic and causal inference queries.
method Algebraic perspective on circuits, focusing on semiring operators.
result Unified sufficient conditions for tractable composition of operators.
Paper introduces md-vtrees for efficient probabilistic and causal inference.
problem Efficient inference in complex probabilistic models.
method Introduces md-vtrees to generalize tractability conditions for advanced inference queries.
result Derives first polytime algorithms for causal inference queries.
The aim of this short note is to draw attention to a method by which the partition function and marginal probabilities for a certain class of random fields on complete graphs can be computed in polynomial time. This class includes Ising models with homogeneous pairwise potentials but arbitrary (inhomogeneous) unary pot…
MDMA provides closed-form marginals and conditionals for deep networks.
problem Lack of closed-form marginals and conditionals in deep neural models.
method MDMA architecture combining deep scalar representations and hierarchical tensor decompositions.
result MDMA outperforms state-of-the-art models in tasks requiring marginalization and conditional inference.
New tractable density models from squaring neural networks.
problem Flexible models for probability distributions in machine learning.
method Squared Neural Family (SNEFY) models formed by squaring neural network outputs and normalizing.
result SNEFYs are fully tractable with closed form normalizing constants in many cases.
Deep Gaussian processes provide a flexible approach to probabilistic modelling of data using either supervised or unsupervised learning. For tractable inference approximations to the marginal likelihood of the model must be made. The original approach to approximate inference in these models used variational compressio…
Automatically improves Monte Carlo estimators in probabilistic programs.
problem Reducing variance in Monte Carlo estimators for probabilistic programs.
method Dynamic mechanism using conjugate priors and affine transformations.
result Automatic Rao-Blackwellization and locally-optimal proposals.
Hybrid model combines continuous and tractable probabilistic models.
problem Intractable probabilistic inference in continuous latent-space models.
method Continuous mixtures of tractable probabilistic models with finite integration points.
result Hybrid models achieve state-of-the-art performance in density estimation.
Proposes MFSWB for marginal fairness in SWB, improving efficiency and performance.
problem Achieving marginal fairness in SWB averaging.
method Defining MFSWB as a constrained SWB problem, proposing two surrogate problems and a new slicing distribution.
result Surrogate MFSWB problems effectively minimize distances to marginals and encourage marginal fairness.
Infinite-dimensional polynomial diffusions preserve tractability of finite-dimensional counterparts.
problem Modeling and analyzing infinite-dimensional probability measure-valued diffusions.
method Introduced polynomial diffusions, transferred properties from finite to infinite dimensions, and proved well-posedness of martingale problems.
result Tractability of finite-dimensional polynomial processes is preserved in the infinite-dimensional setting.
Bayesian approach learns invariances from data alone, but last layer approximation is not always sufficient.
problem Learning invariances in neural networks using only training data.
method Bayesian marginal likelihood for last layer, custom optimisation routine, new lower bound.
result Partial success on standard benchmarks and medical imaging dataset, failure on CIFAR10.
The paper proposes a parallelizable clustering method for multivariate data.
problem The standard model-based clustering method assumes the same number of clusters per margin, which is often unrealistic.
method Developed a finite mixture model per margin with different numbers of clusters, and used a game-inspired algorithm to cluster multivariate data.
result The proposed method shows good performance in various scenarios and real datasets.
Proposes logistic-beta process for modeling dependent probabilities with beta marginals.
problem Limited work on flexible and computationally convenient stochastic process extensions for dependent random probabilities.
method Introduces logistic-beta process with logistic transformation and beta marginals, capable of modeling dependence in discrete and continuous domains.
result Logistic-beta processes enable effective posterior inference and design of computationally tractable dependent Bayesian nonparametric models.
Introduces joint exclusivity (JE), a new form of negative dependence.
problem Negative dependence structures in probability distributions.
method Defines JE by exclusion of the interior of the non-negative orthant, establishes necessary and sufficient conditions for existence, proposes a canonical construction.
result Sharp necessary and sufficient condition for existence of JE random vectors with prescribed marginals.
DBKs enable scalable GPs with tractable inference for large datasets.
problem Scaling Gaussian processes to large and complex datasets while maintaining tractable inference.
method DBKs constructed from neural-network-parameterized basis functions with explicit low-rank structure, enabling linear-complexity inference.
result DBKs provide a unified perspective and improve predictive accuracy, uncertainty quantification, and computational efficiency.
New method bounds causal effects using local consistency of marginals.
problem Bounding causal effects due to unmeasured confounding.
method Enforces compatibility between marginals of causal models and data.
result Explicit algorithm and implementation of causal marginal polytope.
ACVAEs improve on CVAEs by learning more flexible latent correlations.
problem Learning latent representations with correlated structure.
method Adaptive prior distribution and belief propagation.
result ACVAEs outperform CVAEs in link prediction and hierarchical clustering.
AIS uses a suboptimal extended target distribution, which this paper improves using SGM.
problem Improving the efficiency of Annealed Importance Sampling for marginal likelihood estimation.
method Leveraging score-based generative modeling to approximate the optimal extended target distribution.
result Demonstrated novel, differentiable AIS procedures on synthetic and real-world data.
DIET tests conditional independence using marginal dependence measures of residual information.
problem Computational intractability of conditional randomization tests (CRTs).
method DIET avoids fitting large models by leveraging marginal independence statistics of information residuals.
result DIET achieves higher power than other tractable CRTs on synthetic and real benchmarks.
New method for Bayesian SBM inference is scalable, accurate, and concise.
problem Bayesian inference in the stochastic block model for graph clustering.
method Combining variational approximation and Laplace's method, a tractable algorithm is derived.
result The method solves cluster assignment and model selection tasks concurrently.
This work proposes a new method to estimate joint probability from pairwise marginals, reducing sample complexity.
problem Direct nonparametric estimation of high-dimensional joint probability is infeasible due to the curse of dimensionality.
method Developed a coupled nonnegative matrix factorization (CNMF) framework using only pairwise marginals.
result The method provably recovers the joint probability mass function up to bounded error in finite iterations under reasonable conditions.
This paper proposes a new method to improve VI approximations by capturing dependence between blocks using vector copulas.
problem Improving variational inference accuracy for complex models with challenging posteriors.
method Using vector copulas to model dependence between multivariate blocks, with learnable transport maps for flexible marginals.
result The proposed method produces more accurate posterior approximations than existing methods at limited computational cost.
New method estimates high-dimensional copulas for time series data.
problem Estimating copulas for high-dimensional, discrete-margined time series data.
method Variational Bayes estimator for tractable augmented posterior.
result Faster than previous likelihood-based methods, estimating up to 792 dimensions.
A new method trains and samples from energy-based models using diffusion recovery likelihood.
problem Training and sampling high-dimensional datasets with energy-based models is challenging.
method Trains EBMs with a diffusion recovery likelihood method, maximizing conditional probabilities of data at different noise levels.
result Generates high-fidelity images with low FID and inception scores, and accurately estimates normalized data density.
We consider the problem of covariance matrix estimation in the presence of latent variables. Under suitable conditions, it is possible to learn the marginal covariance matrix of the observed variables via a tractable convex program, where the concentration matrix of the observed variables is decomposed into a sparse ma…
In this paper we study output coding for multi-label prediction. For a multi-label output coding to be discriminative, it is important that codewords for different label vectors are significantly different from each other. In the meantime, unlike in traditional coding theory, codewords in output coding are to be predic…
The mean field methods, which entail approximating intractable probability distributions variationally with distributions from a tractable family, enjoy high efficiency, guaranteed convergence, and provide lower bounds on the true likelihood. But due to requirement for model-specific derivation of the optimization equa…
Factorized information criterion (FIC) is a recently developed approximation technique for the marginal log-likelihood, which provides an automatic model selection framework for a few latent variable models (LVMs) with tractable inference algorithms. This paper reconsiders FIC and fills theoretical gaps of previous FIC…
The key limiting factor in graphical model inference and learning is the complexity of the partition function. We thus ask the question: what are general conditions under which the partition function is tractable? The answer leads to a new kind of deep architecture, which we call sum-product networks (SPNs). SPNs are d…
Deep Gaussian Processes are reinterpreted as deep trigonometric networks for tractable inference.
problem Challenging inference in DGPs due to intractable marginalization in latent function space.
method Viewing DGPs as deep trigonometric networks with Bochner's theorem, and using the wide limit with a bottleneck to translate DGPs into deep trigonometric networks.
result The weight space view yields the same effective covariance functions as obtained in function space, and varying prior distributions over network parameters is equivalent to employing different kernels.
Quantum computing speeds up option pricing for multiple assets.
problem High-dimensional integration bottleneck in option pricing.
method Calibrated marginal distributions, Gaussian copula, QAMC with QAE.
result QAMC reduces integration queries by 10-100 times for similar precision.
Develops active learning method for linear optimization with margin-based criterion.
problem Optimizing decisions in linear optimization problems with limited labeled data.
method Smart Predict-then-Optimize (SPO) loss and margin-based active learning algorithm.
result Algorithm achieves significantly fewer labels than naive supervised learning, especially for minimizing SPO loss.
Proposes a new portfolio optimization method considering reward, dispersion, and asymmetry.
problem Capturing fat-tails and asymmetry in asset return distributions.
method Market model with tempered stable distribution; extended mean-variance optimization.
result Closed-form solutions for VaR and CVaR; efficient frontier extended to three dimensions.
Paper proposes Monarch matrices for scalable probabilistic circuits.
problem Improving scalability of probabilistic circuits.
method Sparse Monarch matrices for sum blocks in PCs.
result Significantly reduces memory and computation costs, enabling unprecedented scaling.
A new method scores contextual Markov networks without assuming chordality.
problem Learning structure in contextual Markov networks is hard due to many possible structures.
method Marginal pseudo-likelihood as a consistent structure estimator.
result Marginal pseudo-likelihood yields a consistent structure estimator.
Dynamic trees are mixtures of tree structured belief networks. They solve some of the problems of fixed tree networks at the cost of making exact inference intractable. For this reason approximate methods such as sampling or mean field approaches have been used. However, mean field approximations assume a factorized di…
A method to compute divergences between decomposable models, useful in supervised learning.
problem Computing exact divergences between high-dimensional distributions is intractable.
method Proposes an approach to compute exact alpha-beta divergences between marginal and conditional distributions of decomposable models.
result Tractable computation of marginal and conditional alpha-beta divergences.
MIM learns joint distributions with mutual information and low divergence.
problem Learning joint distributions over observations and latent variables.
method Probabilistic auto-encoder with three design principles: low divergence, high mutual information, and low marginal entropy.
result MIM learns representations with high mutual information, consistent encoding and decoding distributions, effective latent clustering, and comparable data log likelihood to VAE.
GPRF approximates large-scale Gaussian processes for efficient modeling.
problem High computational complexity of Gaussian processes for large datasets.
method Local Gaussian processes coupled via pairwise potentials, with a tractable likelihood.
result GPRF enables latent variable modeling and hyperparameter selection on large datasets.
New method for fitting graphical models with latent variables using regularized conditional likelihood.
problem Graphical modeling with latent variables and confounding dependencies.
method Regularized conditional likelihood for exponential family graphical models.
result Framework applicable to broader settings without knowing latent variables' distribution.
Novel convex surrogate for submodular losses with tractable computation.
problem Learning with non-modular losses for set prediction.
method Proposed Lovász hinge loss function for submodular losses.
result First tractable convex surrogates for submodular losses.
Amortized VI for DGPs learns efficient inference.
problem Expressive limitations in GP approximations.
method Amortized variational inference for DGPs.
result Improved expressive prior and posterior for DGPs.
New methods optimize sums of bivariate functions on finite domains.
problem Optimizing functions with multiple arguments that are sums of bivariate functions.
method Measure-valued extensions, ℓ2-approximation, entropy-regularization, linear programming, coordinate ascent. result Tractable problem formulations solvable with various methods.
We analyze the generalized Mallows model, a popular exponential model over rankings. Estimating the central (or consensus) ranking from data is NP-hard. We obtain the following new results: (1) We show that search methods can estimate both the central ranking pi0 and the model parameters theta exactly. The search is n!…
One of the fundamental problems in machine learning is the estimation of a probability distribution from data. Many techniques have been proposed to study the structure of data, most often building around the assumption that observations lie on a lower-dimensional manifold of high probability. It has been more difficul…
Proposes a new model to better handle correlation risk in credit risk calculations.
problem Empirical evidence shows correlation risk is significant in credit risk models.
method Introduces a stochastic correlation extension of the Vasicek model using circular diffusion.
result Demonstrates how correlation volatility and persistence affect joint default and survival probabilities.
DiPhon generates scalable graphs via diffusion on graphons.
problem Scaling diffusion models to large graphs.
method Formulated a continuous diffusion process on graphon space via Jacobi SDE, discretized for finite graphs.
result DiPhon matches the first moment of graphon dynamics and approximates the second moment.