Adversarial learning of probabilistic models has recently emerged as a promising alternative to maximum likelihood. Implicit models such as generative adversarial networks (GAN) often generate better samples compared to explicit models trained by maximum likelihood. Yet, GANs sidestep the characterization of an explici…
This paper improves normalizing flows by combining MLE and sliced-Wasserstein distance for better data fidelity.
problem Normalizing flows struggle with generating realistic data and detecting out-of-distribution data.
method Proposes a hybrid objective function combining MLE and sliced-Wasserstein distance.
result Shows better generative abilities and lower likelihood of out-of-distribution data.
We propose a robust estimator to improve maximum likelihood in probabilistic models.
problem Overfitting and sensitivity to noise in maximum likelihood estimation.
method Distributionally robust maximum likelihood estimator that minimizes worst-case expected log-loss.
result The robust estimator is statistically consistent and performs well in regression and classification tasks.
Maximum likelihood estimation fails to be well-posed in Gaussian process regression.
problem Establishing well-posedness of maximum likelihood estimation in Gaussian process regression.
method Analyzing the conditions under which maximum likelihood estimation is not Lipschitz in the data with respect to the Hellinger distance.
result Maximum likelihood estimation is not well-posed in the noiseless data setting for any Gaussian process with a stationary covariance function whose lengthscale parameter is estimated using maximum likelihood.
A new method improves text generation quality and diversity.
problem Exposure bias in Maximum Likelihood Estimation for text generation.
method ψ-MLE, a new training scheme based on density ratio estimation.
result ψ-MLE outperforms Maximum Likelihood Estimation and other models in text generation quality and diversity.
Geometric approach solves maximum likelihood for Cauchy-like distributions.
problem Estimating center and scatter robustly from heavy-tailed data.
method Geodesic convexity and symmetry spaces of noncompact type.
result Efficient numerical solution for robust estimates of location and spread.
This work improves likelihood of score-based diffusion ODEs using high-order denoising score matching.
problem The gap between maximum likelihood and score matching objectives for score-based diffusion ODEs.
method High-order denoising score matching to maximize likelihood.
result Score-based diffusion ODEs achieve better likelihood on synthetic and CIFAR-10 data.
Paper proposes efficient training for normalizing flows in Boltzmann generators.
problem Training normalizing flows for Boltzmann generators is computationally challenging and unstable.
method Regression Training of Normalizing Flows (RegFlow) using ℓ 2 \ell_2 ℓ 2 -regression. result RegFlow enables efficient and stable training of normalizing flows for Boltzmann generators.
We consider two connected aspects of maximum likelihood estimation of the parameter for high-dimensional discrete graphical models: the existence of the maximum likelihood estimate (mle) and its computation. When the data is sparse, there are many zeros in the contingency table and the maximum likelihood estimate of th…
Maximum likelihood training improves the performance of score-based diffusion models.
problem Training score-based diffusion models with maximum likelihood.
method Trained by minimizing a weighted combination of score matching losses, with a specific weighting scheme that bounds negative log-likelihood.
result Maximum likelihood training improves the log-likelihood of score-based diffusion models across multiple datasets.
Paper presents a method for estimating Hawkes process parameters.
problem Estimating parameters of Hawkes processes with self-excitation or inhibition.
method Maximum likelihood estimation for Hawkes processes with self-excitation or inhibition.
result The proposed estimator provides more accurate estimations in the inhibition context.
New method for efficient maximum likelihood estimation of p p p -generalized probit regression.
problem Efficient estimation of p p p -generalized probit regression models. method Combining sketching techniques with importance subsampling to obtain a coreset.
result Maximum likelihood estimator can be approximated efficiently up to a factor of ( 1 + ε ) (1+\varepsilon) ( 1 + ε ) on large data. Improved autoregressive models generate higher quality images and are more robust to noise.
problem Generating high-quality images from autoregressive models.
method Noise conditional maximum likelihood estimation (MLE) with score-based sampling.
result Models trained with noise conditional MLE achieve better test likelihoods and generate higher quality images.
Estimates GLMs robustly against label corruptions.
problem Learning GLMs under adversarial label corruptions.
method Iterative trimmed maximum likelihood estimator.
result Achieves minimax near-optimal risk.
Corrects pseudo log-likelihood method issues in various applications.
problem Log-likelihood function unbounded issues in pseudo log-likelihood methods.
method Provided a counterexample and corrected algorithms in previous literature.
result Ensured well-definedness of maximum pseudo log-likelihood estimation.
Unified detector calibration and simulation using MLE from generative models.
problem Combining detector calibration and simulation using traditional methods.
method Maximum likelihood estimation from conditional generative models.
result Prior-independent and non-Gaussian resolutions possible.
Study on likelihood functions, associative equations, and Frobenius manifolds.
problem Maximum likelihood estimation and associativity equations in statistical models.
method Analyzes the cone of concentration matrices, log-likelihood function, and Frobenius manifolds.
result Maximum likelihood degree is indexed by components of Frobenius residuals.
EBMs trained with ML are shown to behave like GANs with a self-adversarial loss.
problem Training EBMs with ML is intractable due to intractable unnormalized distributions.
method Replaced MCMC with deterministic gradient descent ODE solutions to study density induced by dynamics.
result EBM training is effectively a self-adversarial procedure rather than ML estimation.
We present a new statistical learning paradigm for Boltzmann machines based on a new inference principle we have proposed: the latent maximum entropy principle (LME). LME is different both from Jaynes maximum entropy principle and from standard maximum likelihood estimation.We demonstrate the LME principle BY deriving …
New method resolves nonidentifiability in mixture models.
problem Nonidentifiability in marginal models of mixture models.
method Introducing an effective temperature to generalize the marginal likelihood.
result Maximization of the generalized likelihood leads to unique results.
Efficiently estimates GEV distribution parameters using neural networks.
problem Computational intensity of maximum likelihood estimation for GEV distribution.
method Neural network-based likelihood-free estimation method.
result Comparable accuracy to maximum likelihood method with significant speedup.
CMLE reduces spurious correlations in deep models.
problem Spurious correlations in deep learning models.
method Counterfactual Maximum Likelihood Estimation (CMLE) on interventional distribution.
result CMLE outperforms regular MLE in out-of-domain generalization and spurious correlation reduction.
New approach resolves ambiguity in PPCA model's maximum likelihood estimation.
problem Ambiguity in maximum likelihood estimation of PPCA model due to rotational symmetry.
method Using quotient topological spaces, the approach resolves ambiguity and shows consistency of the maximum likelihood solution.
result Maximum likelihood solution is consistent in an appropriate quotient Euclidean space.
A new method optimizes neural sequence models for better task performance.
problem Training neural sequence models with maximum likelihood estimation ignores task losses.
method Maximum likelihood guided parameter search (MGS) in the parameter space.
result MGS optimizes sequence-level losses, reducing repetition and non-termination.
Graphical lasso may fail to fit models when data points are insufficient.
problem When does graphical lasso fail to select and fit a graphical model?
method Computational experiments with graphical lasso.
result Graphical lasso may fail when the number of data points is less than the maximum likelihood threshold.
A new method speeds up quantum state estimation.
problem Exponential growth in sample size and dimension for quantum state tomography.
method Stochastic mirror descent with Burg entropy.
result Optimization error vanishes at a O ( ( 1 / t ) d log t ) O (\sqrt{ ( 1 / t ) d \log t }) O ( ( 1/ t ) d log t ) rate. We develop a maximum penalized quasi-likelihood estimator for estimating in a nonparametric way the diffusion function of a diffusion process, as an alternative to more traditional kernel-based estimators. After developing a numerical scheme for computing the maximizer of the penalized maximum quasi-likelihood function…
Invertibility conditions for observation-driven time series models often fail to be guaranteed in empirical applications. As a result, the asymptotic theory of maximum likelihood and quasi-maximum likelihood estimators may be compromised. We derive considerably weaker conditions that can be used in practice to ensure t…
Proposes a new approach to approximate maximum likelihood for complex models.
problem Intractable likelihood functions in complex parametric models.
method Simulation-based constrained approximation to the structural model.
result Estimators nearly as efficient as maximum likelihood, feasible in many cases.
New algorithms improve linear bandit performance with low computation.
problem Optimizing reward in linear stochastic bandits.
method Reward-biased maximum likelihood method modified for linear and generalized linear bandits.
result New policies achieve order-optimality and competitive empirical performance.
Machine learning should incorporate maximum likelihood for better estimation.
problem Lack of rigorous foundational theory in machine learning.
method Integrate maximum likelihood estimation into machine learning models.
result Foundationally rigorous machine learning models have greater practical impact.
Generative Adversarial Networks (GANs) enjoy great success at image generation, but have proven difficult to train in the domain of natural language. Challenges with gradient estimation, optimization instability, and mode collapse have lead practitioners to resort to maximum likelihood pre-training, followed by small a…
Unified view of KL-divergence and IPMs via DRE, with new DRM metrics.
problem Unified understanding of KL-divergence and IPMs.
method Unified representation via maximum likelihood density-ratio estimation (DRE).
result Unified form of IPMs and novel DRM metrics.
Paper proposes energy objective for training normalizing flows without determinants.
problem Challenges in training normalizing flows due to Jacobian determinants.
method Introduces energy objective based on proper scoring rules, determinant-free.
result Energy objective supports novel model families and competitive performance.
Proposes a guaranteed regularization method for maximum likelihood estimation using gauge symmetry in Kullback-Leibler divergence.
problem Overfitting in maximum likelihood estimation.
method Introduces a regularization approach based on gauge symmetry in Kullback-Leibler divergence.
result The method provides a theoretically guaranteed optimal model without frequent hyperparameter tuning.
Quantum ML predicts data with improved speed and accuracy.
problem Predicting data using maximum likelihood in a quantum setting.
method Quantum states embedding and minimization of quantum relative entropy.
result Unified framework for classical and quantum LLMs with performance guarantees.
New particle algorithms optimize latent variable models.
problem Optimizing latent variable models for maximum likelihood estimation.
method Identify gradient flows associated with free energy functional and discretize them to create particle-based algorithms.
result Novel particle algorithms scale to high-dimensional settings and perform well in experiments.
This work investigates training infinite mixtures with maximum likelihood for improved uncertainty quantification.
problem Improving uncertainty quantification in neural networks.
method Investigates training infinite mixtures with maximum likelihood instead of variational inference.
result The proposed method leads to stochastic networks with increased predictive variance, improved robustness, and higher entropy on out-of-distribution data.
We present an approximated maximum likelihood method for the multifractal random walk processes of [E. Bacry et al., Phys. Rev. E 64, 026103 (2001)]. The likelihood is computed using a Laplace approximation and a truncation in the dependency structure for the latent volatility. The procedure is implemented as a package…
Generalizes moment-matching for exponential families with conditioning or hidden data.
problem Generalizing moment-matching conditions for exponential families with conditioning or hidden data.
method First-principles explanation and self-contained derivation of generalized moment-matching conditions.
result Derives generalized moment-matching conditions for conditional exponential families and hidden data.
Maximum likelihood estimator performance in logistic regression analyzed.
problem Performance of maximum likelihood estimator in logistic regression.
method Sharp non-asymptotic guarantees for existence and excess logistic risk.
result Sharp guarantees for the existence and excess risk of MLE in logistic regression.
New analysis shows entropy term cancels out in likelihood-based OOD detection.
problem Curious likelihood values for out-of-distribution data.
method Decomposed average likelihood into KL divergence and entropy terms.
result Entropy term explains OOD behaviour and cancels out in expectation.
New method estimates latent gene expression factors without overlap with known confounders.
problem Estimating latent variance components in gene expression data with known confounders.
method Restricted maximum-likelihood method maximizing likelihood on orthogonal subspace.
result Method reduces runtime and attains greater likelihood values than gradient-based optimizers.
Optimizes MMD learning for generative models with theoretical guarantees.
problem Theoretical guarantees for optimizing non-convex MMD objectives.
method Analyzes MMD optimization landscape for specific distributions.
result Gradient-based methods globally minimize MMD objective for certain distributions.
MLE works best for covariate shift without modifications.
problem OOD generalization under covariate shift.
method Maximum Likelihood Estimation (MLE) without modifications.
result MLE achieves minimax optimality for covariate shift under well-specified setting.
Bayesian networks with latent variables are characterized and their likelihoods compared.
problem Characterizing and comparing likelihoods of Bayesian networks with latent variables.
method Characterized likelihood function and empirical Bayesian network. Proved dominance of global maximum likelihood from empirical model.
result The global maximum likelihood of the original Bayesian network is attained if and only if parameters are consistent with empirical model.
Chow and Liu (1968) studied the problem of learning a maximumlikelihood Markov tree. We generalize their work to more complexMarkov networks by considering the problem of learning a maximumlikelihood Markov network of bounded complexity. We discuss howtree-width is in many ways the appropriate measure of complexity and…
A new method learns latent variable updates directly, not approximating the posterior.
problem Intractable maximum-likelihood learning for complex latent-variable models.
method Amortised learning using wake-sleep Monte-Carlo strategy.
result Demonstrated effectiveness on various complex models.