The normalized maximized likelihood (NML) provides the minimax regret solution in universal data compression, gambling, and prediction, and it plays an essential role in the minimum description length (MDL) method of statistical modeling and estimation. Here we show that the normalized maximum likelihood has a Bayes-li…
This paper improves normalizing flows by combining MLE and sliced-Wasserstein distance for better data fidelity.
problem Normalizing flows struggle with generating realistic data and detecting out-of-distribution data.
method Proposes a hybrid objective function combining MLE and sliced-Wasserstein distance.
result Shows better generative abilities and lower likelihood of out-of-distribution data.
Paper proposes efficient training for normalizing flows in Boltzmann generators.
problem Training normalizing flows for Boltzmann generators is computationally challenging and unstable.
method Regression Training of Normalizing Flows (RegFlow) using ℓ2-regression. result RegFlow enables efficient and stable training of normalizing flows for Boltzmann generators.
Maximum likelihood training improves the performance of score-based diffusion models.
problem Training score-based diffusion models with maximum likelihood.
method Trained by minimizing a weighted combination of score matching losses, with a specific weighting scheme that bounds negative log-likelihood.
result Maximum likelihood training improves the log-likelihood of score-based diffusion models across multiple datasets.
This paper proposes a new RV prediction model using neural distributional transformation and co-training.
problem Predicting skewed and fat-tailed realized volatility (RV) is challenging.
method The paper uses a neural distributional transformation and co-training to predict RV. It jointly trains the transformation and prediction model using a maximum-likelihood objective function.
result The proposed method significantly outperforms other methods on a dataset of 100 stocks.
Operational risk models commonly employ maximum likelihood estimation (MLE) to fit loss data to heavy-tailed distributions. Yet several desirable properties of MLE (e.g. asymptotic normality) are generally valid only for large sample-sizes, a situation rarely encountered in operational risk. In this paper, we study how…
Paper proposes energy objective for training normalizing flows without determinants.
problem Challenges in training normalizing flows due to Jacobian determinants.
method Introduces energy objective based on proper scoring rules, determinant-free.
result Energy objective supports novel model families and competitive performance.
New method for efficient maximum likelihood estimation of p-generalized probit regression.
problem Efficient estimation of p-generalized probit regression models. method Combining sketching techniques with importance subsampling to obtain a coreset.
result Maximum likelihood estimator can be approximated efficiently up to a factor of (1+ε) on large data. A new path gradient estimator speeds up normalizing flows without sacrificing accuracy.
problem High computational cost and limited scalability of path gradient estimators for normalizing flows.
method Proposed a fast path gradient estimator that improves computational efficiency and scalability.
result The new estimator achieves superior performance and reduced variance across various applications.
The paper strengthens the classical result of MLE convergence to a Gaussian distribution.
problem The classical result of MLE convergence to a Gaussian distribution.
method Sub-Gaussian concentration and entropic normality of the normalized MLE.
result Entropic central limit theorem for a smoothed version of the estimator.
A method for converting NIW parameters for better estimation.
problem Estimating parameters of multivariate normal distribution.
method Convergent procedure for converting mean parameters to natural parameters in NIW family.
result Maximum likelihood estimation of natural parameters from observed statistics.
Stochastic normalizing flows use SDEs for efficient training and sampling.
problem Efficient maximum likelihood estimation and variational inference.
method Continuous normalizing flows extended with stochastic differential equations (SDEs) and rough path theory.
result Stochastic normalizing flows enable efficient training and sampling from complex distributions.
Paper proves method for calculating NML code length works for continuous models.
problem Uncertainty in calculating NML code length for continuous models.
method Introduced a novel decomposition approach based on the coarea formula to prove correctness for continuous cases.
result Method accurately calculates NML code length for continuous models.
A new layer, funnel, reduces dimensionality in flows for better performance.
problem Training high-dimensional models efficiently and accurately.
method Constructing dimension-reducing surjective flows using the funnel layer.
result The funnel layer improves model performance with a smaller latent space.
A new method normalizes EBM training by introducing a learnable parameter.
problem Training energy-based models with maximum likelihood is challenging due to intractable normalisation constants.
method Proposes a self-normalised log-likelihood (SNL) objective that introduces a learnable parameter representing the normalisation constant.
result The SNL objective is a lower bound of the log-likelihood and can be directly optimised using stochastic gradient techniques.
Paper proposes efficient method to calculate Fisher-Bingham distribution normalizing constant.
problem Efficiently calculating the normalizing constant of Fisher-Bingham distributions.
method Numerical integration with continuous Euler transform to Fourier-type integral representation.
result The method is fast and accurate, applicable to high-dimensional distributions.
Maximum likelihood estimation fails to be well-posed in Gaussian process regression.
problem Establishing well-posedness of maximum likelihood estimation in Gaussian process regression.
method Analyzing the conditions under which maximum likelihood estimation is not Lipschitz in the data with respect to the Hellinger distance.
result Maximum likelihood estimation is not well-posed in the noiseless data setting for any Gaussian process with a stationary covariance function whose lengthscale parameter is estimated using maximum likelihood.
We study asymptotic properties of maximum likelihood estimators of drift parameters for a jump-type Heston model based on continuous time observations, where the jump process can be any purely non-Gaussian Lévy process of not necessarily bounded variation with a Lévy measure concentrated on (−1,∞). We prove stro…
The paper explores how over-parameterized linear regression models generalize without violating learning theory principles.
problem Understanding how over-parameterized linear regression models generalize without violating learning theory principles.
method The paper uses the predictive normalized maximum likelihood (pNML) learner to investigate the minimum norm solution of over-parameterized linear regression models.
result The model generalizes well when the test sample lies in a subspace spanned by eigenvectors associated with large eigenvalues of the training data.
The Predictive Normalized Maximum Likelihood (pNML) scheme has been recently suggested for universal learning in the individual setting, where both the training and test samples are individual data. The goal of universal learning is to compete with a ``genie'' or reference learner that knows the data values, but is res…
Residual flows are shown to approximate MMD well.
problem Lack of theoretical understanding of normalizing flows' expressiveness.
method Proved residual flows are universal approximators in MMD.
result Residual flows can approximate MMD with a bounded number of blocks.
We study asymptotic properties of maximum likelihood estimators for Heston models based on continuous time observations of the log-price process. We distinguish three cases: subcritical (also called ergodic), critical and supercritical. In the subcritical case, asymptotic normality is proved for all the parameters, whi…
ACNML method improves uncertainty estimation for deep networks.
problem Uncertainty estimation and calibration for deep neural networks under distribution shift.
method Approximate Bayesian inference to approximate CNML distribution.
result ACNML compares favorably to prior techniques for uncertainty estimation.
Score matching fails to train VAEs robustly, revealing autoencoding loss insights.
problem Catastrophic failure of variational score matching on VAE models.
method Analysis of existing variational score matching objectives and their equivalence to autoencoding losses.
result Score matching methods fail to produce robust VAE models, predicting poor performance.
We improve maximum likelihood for location estimation in finite samples.
problem Estimating a parameter from samples with unknown or varying distribution.
method Use smoothed Fisher information for finite sample size and varying distributions.
result Recover optimal estimation theory for finite n and arbitrary f. Unified detector calibration and simulation using MLE from generative models.
problem Combining detector calibration and simulation using traditional methods.
method Maximum likelihood estimation from conditional generative models.
result Prior-independent and non-Gaussian resolutions possible.
Paper proposes CoopFlow, a two-flow generator for energy-based models.
problem Training energy-based models with Langevin flow and normalizing flow.
method CoopFlow trains an energy-based model using a normalizing flow initialization and a short-run Langevin flow revision.
result CoopFlow converges to a moment matching estimator and synthesizes realistic images.
We consider a stable Cox--Ingersoll--Ross process driven by a standard Wiener process and a spectrally positive strictly stable Lévy process, and we study asymptotic properties of the maximum likelihood estimator (MLE) for its growth rate based on continuous time observations. We distinguish three cases: subcritical, c…
An efficient LDP protocol for QMLE with improved practicality and theoretical guarantees.
problem Difficult implementation of existing LDP QMLE for large-scale surveys.
method Developed an alternative LDP protocol without long waiting time, high communication cost, and derivative boundedness assumptions.
result Sufficient conditions for consistency and asymptotic normality of the protocol.
Likelihood from a generative model is a natural statistic for detecting out-of-distribution (OoD) samples. However, generative models have been shown to assign higher likelihood to OoD samples compared to ones from the training distribution, preventing simple threshold-based detection rules. We demonstrate that OoD det…
Pre-training improves model coverage, crucial for downstream performance.
problem Understanding why pre-training enhances model performance.
method Coverage principle, focusing on next-token prediction and model quality.
result Coverage generalizes faster than cross-entropy, improving downstream performance.
There are many models, often called unnormalized models, whose normalizing constants are not calculated in closed form. Maximum likelihood estimation is not directly applicable to unnormalized models. Score matching, contrastive divergence method, pseudo-likelihood, Monte Carlo maximum likelihood, and noise contrastive…
Normalizing flows improve density estimation from noisy data.
problem Estimating underlying density from noisy samples.
method Use normalizing flows for density estimation with arbitrary noise distributions, using amortized variational inference.
result Normalizing flows can outperform Gaussian mixtures for density deconvolution.
Deep learning model estimates uncertainty in complex regression tasks.
problem Uncertainty quantification in probabilistic regression predictions.
method Combines statistical and deep learning transformation models using gradient descent.
result State-of-the-art performance on small datasets and complex image data.
Quantum ML predicts data with improved speed and accuracy.
problem Predicting data using maximum likelihood in a quantum setting.
method Quantum states embedding and minimization of quantum relative entropy.
result Unified framework for classical and quantum LLMs with performance guarantees.
Machine learning should incorporate maximum likelihood for better estimation.
problem Lack of rigorous foundational theory in machine learning.
method Integrate maximum likelihood estimation into machine learning models.
result Foundationally rigorous machine learning models have greater practical impact.
New method trains any neural network as a generative model.
problem Constrained design of normalizing flows due to analytical invertibility.
method Efficient gradient estimator for non-analytically invertible networks.
result Any dimension-preserving neural network can be used as a generative model.
Paper extends LME models to allow sign constraints on coefficients with SDTN random effects.
problem Inference with sign constraints on random effects in LME models.
method Proposes SDTN distribution for random effects and develops likelihood-based approaches for estimation.
result Proposed constrained model improves real-world interpretations and achieves satisfactory performance.
Optimal downsampling improves GLM performance in imbalanced classification.
problem Improving GLM performance in imbalanced classification.
method Proposed a pseudo maximum likelihood estimator for optimal downsampling.
result The introduced estimator outperforms existing alternatives in both synthetic and empirical data.
Paper proposes a simple estimator for DPP correlation kernels.
problem Estimating the correlation kernel matrix of DPPs.
method Closed-form estimator for correlation kernel, easy to implement.
result Consistency and asymptotic normality of the estimator proved.
Improved Gaussian Neural Processes for efficient multi-dimensional predictions.
problem Inability to model dependencies in outputs limits CNPs and NPs applicability.
method Proposes a new approach to model output dependencies using latent variables for maximum likelihood training, scalable to 2D and 3D data.
result Proposed models show good performance in synthetic experiments.
We propose a robust estimator to improve maximum likelihood in probabilistic models.
problem Overfitting and sensitivity to noise in maximum likelihood estimation.
method Distributionally robust maximum likelihood estimator that minimizes worst-case expected log-loss.
result The robust estimator is statistically consistent and performs well in regression and classification tasks.
Paper proposes an alternative to MLE for GLMs with non-canonical link functions.
problem Challenges in MLE for GLMs with non-canonical link functions.
method Variational Inequality (VI) estimation framework.
result Established finite-sample error bounds and asymptotic normality for VI estimator.
Label shift refers to the phenomenon where the prior class probability p(y) changes between the training and test distributions, while the conditional probability p(x|y) stays fixed. Label shift arises in settings like medical diagnosis, where a classifier trained to predict disease given symptoms must be adapted to sc…
Supervised topic models utilize document's side information for discovering predictive low dimensional representations of documents. Existing models apply the likelihood-based estimation. In this paper, we present a general framework of max-margin supervised topic models for both continuous and categorical response var…
Proposes LFGP for likelihood-free Gaussian process regression.
problem Inability to set likelihood functions in unknown probability models.
method Clusters and approximates likelihood using asymptotic normality.
result Reduces assumptions and computational costs for scalable problems.
New method for robust distribution alignment using log-likelihood ratio and normalizing flows.
problem Distribution alignment challenges in deep learning.
method Log-likelihood ratio statistic and normalizing flows.
result Minimizing the proposed objective yields robust domain alignment.
This work investigates training infinite mixtures with maximum likelihood for improved uncertainty quantification.
problem Improving uncertainty quantification in neural networks.
method Investigates training infinite mixtures with maximum likelihood instead of variational inference.
result The proposed method leads to stochastic networks with increased predictive variance, improved robustness, and higher entropy on out-of-distribution data.