Simplifies RCA for anomalies with separable likelihoods.
problem Challenges in accurate and friendly root cause analysis.
method Bayesian framework with separable likelihoods under certain restrictions.
result Framework successfully applied to web server error logs.
Separable losses are inconsistent for structured prediction models.
problem Inconsistency of separable losses in structured prediction models.
method Analysis of separable negative log-likelihood losses for structured prediction.
result Separable losses are not Bayes consistent and may not predict the most probable structure.
Consistent estimator for mixtures of nonparametric elliptical distributions helps cluster analysis.
problem Consistency of maximum likelihood estimator for mixtures of nonparametric elliptical distributions.
method Maximum likelihood estimation for mixtures of elliptically-symmetric distributions under nonparametric P. result Components of the estimator correspond to well-separated components of the underlying distribution P. Prob-PIT improves speech separation by considering output-label permutations as random variables.
problem Overconfident output-label assignment in PIT leads to unreliable speech separation.
method Prob-PIT treats output-label permutations as a discrete latent random variable with a uniform prior distribution and maximizes the log-likelihood function.
result Prob-PIT significantly outperforms PIT in terms of Signal to Distortion Ratio and Signal to Interference Ratio.
Algorithm finds frequencies, amplitudes, and phases of sinusoids in noisy data.
problem Finding frequencies, amplitudes, and phases of sinusoids in noisy data.
method Maximum likelihood approach to estimate tone parameters from contaminated observations. Successively estimates frequencies and jointly optimizes amplitudes and phases.
result Near-linear computational complexity (O(N)) for estimating M number of sinusoidal sources. GANs outperform traditional methods in speech source separation.
problem Improving speech source separation accuracy.
method Used a multi-layer perceptron trained with a Wasserstein-GAN formulation to compare against traditional methods.
result Multi-layer perceptron with Wasserstein-GAN outperformed other methods in source to distortion ratio.
New decision-theoretic characterization separates belief and decision posteriors.
problem Understanding the conditions under which loss-based updating coincides with Bayesian updating.
method Decision-theoretic approach to distinguish belief and decision posteriors.
result Generalized Bayes coincides with ordinary Bayesian updating only if the loss is proportional to negative log-likelihood.
Study uniform rates for estimating Gaussian mixtures without separation assumption.
problem Estimating parameters in two-component Gaussian mixtures without separation.
method Uniform convergence rates derived using minimax lower bounds and careful analysis of polynomial equalities.
result Phase transition in optimal estimation rate based on mixture balance.
Study reveals structure of local minima in GMMs, identifying key cluster centers.
problem Identifying optimal cluster centers in non-convex GMM landscapes.
method Analyzing the negative log-likelihood function of GMMs in the population limit.
result Local minima share a common structure that partially identifies true cluster centers.
Study shows LDA topic models converge at rate n^-1/4 without strict topic separability.
problem Convergence rates of Latent Dirichlet Allocation (LDA) topic models.
method Maximum likelihood estimator, Wasserstein's distance metric, without separability or non-degeneracy assumptions.
result Maximum likelihood estimator converges at rate n^-1/4, optimal in worst case.
New ICA method for sources with mixed spectra.
problem Inaccurate separation of sources with temporal autocorrelations and mixed spectra.
method Estimates spectral density functions and line spectra using cubic splines and indicator functions, then maximizes the Whittle likelihood function.
result Outperforms existing ICA methods in simulations and EEG data applications.
New model handles mixed data types better.
problem Handling data with different types of attributes.
method Mixed likelihood Gaussian process latent variable model with separate likelihoods for each dimension.
result Better predictive performance for real-world data with mixed attributes.
DDSME outperforms SME in estimating multimodal distributions.
problem Efficiency of score matching in multimodal distributions.
method Diffusion-based denoising score matching (DDSME) compared to vanilla score matching (SME).
result DDSME avoids the error bound deterioration of SME with increasing mode separation.
Simple Deep LDA models achieve accuracy competitive with softmax baselines.
problem Training Deep LDA models by maximum likelihood estimation leads to overlapping or collapsed class clusters.
method Proposed a constrained Deep LDA formulation with geometric constraints to fix class means and covariance.
result MLE becomes stable under geometric constraints, yielding well-separated class clusters.
Picard-O improves ICA for faster, robust separation of signals.
problem Efficiently separating signals in multi-channel data.
method Preconditioned L-BFGS over orthogonal matrices.
result Picard-O outperforms FastICA in speed and robustness.
The study examines MCMC methods for arbitrary objectives and finds likelihood sharpness impacts performance and regularization.
problem Limitations of MCMC methods for arbitrary objective functions.
method Two-block MCMC framework with Metropolis-Hastings and Gibbs sampling, exploring likelihood curvature and sharpness.
result Likelihood sharpness governs in-sample performance and regularization inferred by training data.
Detecting and recovering labels in binomial logistic mixtures is challenging due to an information gap.
problem Detecting and recovering labels in binomial logistic mixtures
method Propose two feasibility-aware inference procedures
result Avoid misleading component selections and improve label probability calibration
DNLL loss improves deep LDA accuracy and consistency.
problem Pathological solutions in unconstrained Deep LDA.
method Introducing Discriminative Negative Log-Likelihood (DNLL) loss.
result Deep LDA trained with DNLL produces clean latent spaces and better calibrated probabilities.
We derive a statistical model for estimation of a dendrogram from single linkage hierarchical clustering (SLHC) that takes account of uncertainty through noise or corruption in the measurements of separation of data. Our focus is on just the estimation of the hierarchy of partitions afforded by the dendrogram, rather t…
Unified framework for unlearning in diffusion models using KL divergence and likelihood constraints.
problem Removing undesirable data or concepts while preserving utility of pretrained models.
method Constrained optimization framework based on reverse and forward KL divergences, and likelihood constraints.
result Our KL-constrained approach achieves superior retention-unlearning tradeoffs compared to weight-based baselines.
Improved likelihood-free inference for high-dimensional models.
problem Challenges in likelihood-free inference for high-dimensional parameter spaces.
method Bayesian optimization-based approach with misspecification-robust characterisation.
result Efficient inference in 100-dimensional space with real data application.
Study of asymmetric rank-one tensor models with non-Gaussian noise.
problem Analyzing maximum-likelihood estimators for asymmetric rank-one tensor models.
method Spectrally separated branch analysis, resolvent methods, cumulant expansions, Efron-Stein-type variance bounds.
result Asymptotic singular value and mode-wise alignments are robust to non-Gaussian noise.
Reservoir computing's success depends on mapping different input time series to separable states.
problem Quantifying the ability of random linear reservoirs to map different input time series.
method Mathematical framework using spectral properties of the connectivity matrix.
result Separation capacity is fully characterized by the spectral properties of the connectivity matrix.
A new method improves Bayesian inference for multimodal posteriors.
problem Insensitivity to well-separated modes in multimodal posteriors.
method Weighted Kernel Stein Discrepancy method.
result Significantly improved mode sensitivity compared to standard KSD-Bayes.
DISCoVeR learns disentangled representations by separating shared and condition-specific factors.
problem Learning disentangled representations for multi-condition data.
method Dual-latent architecture, parallel reconstructions, max-min objective.
result DISCoVeR achieves improved disentanglement on various datasets.
A new method extends Bayesian optimization to more models and utilities.
problem Extending Bayesian optimization to a broader class of models and utilities.
method Likelihood-free Bayesian Optimization (LFBO) which directly models the acquisition function without separate inference.
result LFBO outperforms state-of-the-art black-box optimization methods on real-world problems.
We provide two fundamental results on the population (infinite-sample) likelihood function of Gaussian mixture models with M≥3 components. Our first main result shows that the population likelihood function has bad local maxima even in the special case of equally-weighted mixtures of well-separated and spherical…
High temporal resolution measurements of human brain activity can be performed by recording the electric potentials on the scalp surface (electroencephalography, EEG), or by recording the magnetic fields near the surface of the head (magnetoencephalography, MEG). The analysis of the data is problematic due to the fact …
This article establishes the performance of stochastic blockmodels in addressing the co-clustering problem of partitioning a binary array into subsets, assuming only that the data are generated by a nonparametric process satisfying the condition of separate exchangeability. We provide oracle inequalities with rate of c…
Maximum likelihood estimator performance in logistic regression analyzed.
problem Performance of maximum likelihood estimator in logistic regression.
method Sharp non-asymptotic guarantees for existence and excess logistic risk.
result Sharp guarantees for the existence and excess risk of MLE in logistic regression.
A new DDPM for link prediction using sub-graph likelihood estimation.
problem Link prediction in graph domains.
method Sub-graph based diffusion model with DDPMs, decomposing likelihood estimation.
result Our model achieves superior performance in link prediction across various datasets.
We define and discuss the first sparse coding algorithm based on closed-form EM updates and continuous latent variables. The underlying generative model consists of a standard `spike-and-slab' prior and a Gaussian noise model. Closed-form solutions for E- and M-step equations are derived by generalizing probabilistic P…
The paper analyzes condition numbers for logistic regression to understand first-order methods' performance.
problem Understanding the performance of first-order methods in logistic regression.
method Introducing condition numbers to measure non-separability and separability of data.
result Condition numbers inform the properties and convergence guarantees of first-order methods.
A new EM-based algorithm improves deep generative model training.
problem Training deep generative models with maximum likelihood is challenging.
method The paper proposes reweighted expectation maximization (REM), a new algorithm that directly maximizes the log marginal likelihood of the data.
result REM learns better generative models than the IWAE, leading to significantly better performance in density estimation benchmarks.
A novel kernel-based test detects equality versus singularity of two probability measures.
problem Detecting equality versus singularity of two probability distributions.
method Combines kernel mean and kernel covariance embeddings to construct a likelihood ratio test statistic.
result The test statistic satisfies a '0/\infty' law, vanishing under the null and diverging under the alternative.
A new method for clustering heterogeneous data using likelihood-adjusted SDP.
problem Clustering heterogeneous data with different cluster shapes and sizes.
method Iterative likelihood-adjusted semidefinite programming (iLA-SDP) method.
result iLA-SDP achieves lower mis-clustering errors compared to other methods.
BS-VAE separates decoder variance and beta to improve VAE performance.
problem Blurriness in VAE outputs and difficulty in analyzing model performance.
method Explicitly separates beta and decoder variance in Beta-Sigma VAE.
result Superior performance in natural image synthesis and controllable parameters.
Method for factor analysis in short panels without assuming sphericity or Gaussianity.
problem Factor analysis in short panels without assuming sphericity or Gaussianity.
method Pseudo maximum likelihood method and asymptotically uniformly most powerful invariant test.
result Systematic risk explains a large part of cross-sectional total variance in bear markets but is not spanned by observed factors.
Enhanced FastMNMF for better speech separation.
problem Improving blind source separation for speech.
method Gaussian scale mixture (GSM) for heavy-tailed distributions.
result GSM-FastMNMF outperforms existing methods in speech enhancement.
Alternative sampling method for autoregressive models using Langevin dynamics.
problem Efficiently sampling from autoregressive models.
method Initialize sequences with white noise and follow Langevin dynamics on global log-likelihood.
result Parallelizes and generalizes sampling process for autoregressive models.
In this article supervised learning problems are solved using soft rule ensembles. We first review the importance sampling learning ensembles (ISLE) approach that is useful for generating hard rules. The soft rules are then obtained with logistic regression from the corresponding hard rules. In order to deal with the p…
The paper investigates topic models, ensuring their statistical identifiability and accuracy.
problem Lack of formal theoretical investigation of topic model identifiability and estimation accuracy.
method Proposes a maximum likelihood estimator (MLE) based on integrated likelihood, introducing new geometric identifiability conditions.
result Introduces weaker conditions for topic model identifiability, allowing a broader investigation.
M-flows learn data manifolds and densities, improving manifold learning and inference.
problem Representing datasets with manifold structure more faithfully.
method Combining normalizing flows, GANs, autoencoders, and energy-based models, with a new training algorithm.
result M-flows learn data manifolds better than standard flows and provide handles for dimensionality reduction.
The paper tackles hypothesis testing for likelihood-free inference with a new kernel-based approach.
problem Testing hypotheses with limited labeled data in likelihood-free inference.
method Kernel-based tests using maximum mean discrepancy (MMD) for non-parametric density comparison.
result Existence of an asymmetric trade-off between labeled and unlabeled data samples.
Transformer autoencoder learns musical style from performances.
problem Learning high-level controls over symbolic music generation.
method Aggregates encodings of input data across time to obtain global style representation.
result Improves control over performance style and melody in music generation tasks.
VJE learns latent representations without contrastive learning, providing probabilistic semantics.
problem Learning latent representations without contrastive signals.
method VJE maximizes a symmetric conditional evidence lower bound (ELBO) on paired encoder embeddings, using a Student-t distribution on a polar representation.
result VJE outperforms standard non-contrastive baselines in ImageNet-1K, CIFAR-10/100, and STL-10.
Deep generative models tackle anomaly detection in imbalanced datasets.
problem Anomaly detection in imbalanced datasets, especially in medical image analysis.
method Formulate anomaly detection as a generative model problem, train generative models on negative data, estimate likelihood of unseen data.
result Generative models can separate positive and negative samples but may struggle with complex data.
Many models of interest in the natural and social sciences have no closed-form likelihood function, which means that they cannot be treated using the usual techniques of statistical inference. In the case where such models can be efficiently simulated, Bayesian inference is still possible thanks to the Approximate Baye…