Paper proposes a faster SPIDER-EM variant for large-scale nonconvex optimization.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
sEM uses optimal transport to improve EM algorithm for better convergence and avoiding local optima.
Improves EM algorithm for better local optima in mixture models.
Develops a high-dimensional differentially-private EM algorithm with near-optimal statistical guarantees.
EM algorithm achieves optimal sample complexity for learning two-component mixed linear regression.
The EM algorithm is a novel numerical method to obtain maximum likelihood estimates and is often used for practical calculations. However, many of maximum likelihood estimation problems are nonconvex, and it is known that the EM algorithm fails to give the optimal estimate by being trapped by local optima. In order to …
Paper proves EM algorithm convergence for mixtures of discrete and continuous parameters.
We take a new look at parameter estimation for Gaussian Mixture Models (GMMs). In particular, we propose using \emph{Riemannian manifold optimization} as a powerful counterpart to Expectation Maximization (EM). An out-of-the-box invocation of manifold optimization, however, fails spectacularly: it converges to the same…
New method optimizes clustering with better log-likelihood landscape.
We consider maximum likelihood estimation for Gaussian Mixture Models (Gmms). This task is almost invariably solved (in theory and practice) via the Expectation Maximization (EM) algorithm. EM owes its success to various factors, of which is its ability to fulfill positive definiteness constraints in closed form is of …
Gradient EM converges exponentially to optimal solution in agnostic mixtures.
A new EM gradient algorithm for mixture models with skewed components.
This paper uses dynamical systems to analyze and ensure convergence of the Bayesian EM algorithm.
Two-Timescale EM Methods improve EM for nonconvex models.
Latent class model (LCM), which is a finite mixture of different categorical distributions, is one of the most widely used models in statistics and machine learning fields. Because of its non-continuous nature and the flexibility in shape, researchers in practice areas such as marketing and social sciences also frequen…
Generalising the idea of the classical EM algorithm that is widely used for computing maximum likelihood estimates, we propose an EM-Control (EM-C) algorithm for solving multi-period finite time horizon stochastic control problems. The new algorithm sequentially updates the control policies in each time period using Mo…
Gradient descent on LSE objectives implicitly performs EM, leading to collapse without volume control.
Differentiable EM for Gaussian Mixture Models improves model integration.
The speed of convergence of the Expectation Maximization (EM) algorithm for Gaussian mixture model fitting is known to be dependent on the amount of overlap among the mixture components. In this paper, we study the impact of mixing coefficients on the convergence of EM. We show that when the mixture components exhibit …
We investigate the ergodic problem of growth-rate maximization under a class of risk constraints in the context of incomplete, Itô-process models of financial markets with random ergodic coefficients. Including {\em value-at-risk} (VaR), {\em tail-value-at-risk} (TVaR), and {\em limited expected loss} (LEL), these cons…
Various bias-correction methods such as EXTRA, gradient tracking methods, and exact diffusion have been proposed recently to solve distributed {\em deterministic} optimization problems. These methods employ constant step-sizes and converge linearly to the {\em exact} solution under proper conditions. However, their per…
To measure the quality of a set of vector quantization points a means of measuring the distance between a random point and its quantization is required. Common metrics such as the {\em Hamming} and {\em Euclidean} metrics, while mathematically simple, are inappropriate for comparing natural signals such as speech or im…
MLE and CVE are equivalent under exponential families, leading to faster and more stable EM algorithms.
EDML is a recently proposed algorithm for learning MAP parameters in Bayesian networks. In this paper, we present a number of new advances and insights on the EDML algorithm. First, we provide the multivalued extension of EDML, originally proposed for Bayesian networks over binary variables. Next, we identify a simplif…
Maximum likelihood estimation (MLE) is one of the most important methods in machine learning, and the expectation-maximization (EM) algorithm is often used to obtain maximum likelihood estimates. However, EM heavily depends on initial configurations and fails to find the global optimum. On the other hand, in the field …
EM algorithm converges linearly and achieves sharp rate in estimating mixtures of pairwise differences.
Enhances large language models' reasoning through simpler off-policy reinforcement learning.
Proposes new methods for Markov chain choice models with panel data.
Mirror descent method improved RL algorithms.
FIEM accelerates EM for large datasets with nonasymptotic convergence bounds.
New algorithm improves on EM for streaming data, outperforming existing methods.
In this paper, we propose a dynamical systems perspective of the Expectation-Maximization (EM) algorithm. More precisely, we can analyze the EM algorithm as a nonlinear state-space dynamical system. The EM algorithm is widely adopted for data clustering and density estimation in statistics, control systems, and machine…
Non-homogeneous hidden Markov models (NHHMM) are a subclass of dependent mixture models used for semi-supervised learning, where both transition probabilities between the latent states and mean parameter of the probability distribution of the responses (for a given state) depend on the set of covariates. A priori w…
Simplified explanation of ML for mixtures and OT.
A new risk measure (FRM) for EM FI returns helps investors protect against volatility and policy instability.
Model-based clustering approaches concern the paradigm of exploratory data analysis relying on the finite mixture model to automatically find a latent structure governing observed data. They are one of the most popular and successful approaches in cluster analysis. The mixture density estimation is generally performed …
EM algorithm converges to global max in latent Gaussian tree models.
Regression mixture models are widely studied in statistics, machine learning and data analysis. Fitting regression mixtures is challenging and is usually performed by maximum likelihood by using the expectation-maximization (EM) algorithm. However, it is well-known that the initialization is crucial for EM. If the init…
Determining the 3D structures of biological molecules is a key problem for both biology and medicine. Electron Cryomicroscopy (Cryo-EM) is a promising technique for structure estimation which relies heavily on computational methods to reconstruct 3D structures from 2D images. This paper introduces the challenging Cryo-…
ELU algorithm improves on EM for over-specified Gaussian mixtures.
We review a resent {\em time-dependent} performance measure for economical time series -- the (optimal) investment horizon approach. For stock indices, the approach shows a pronounced gain-loss asymmetry that is {\em not} observed for the individual stocks that comprise the index. This difference may hint towards an sy…
Researchers use shape analysis to recover protein structures from Cryo-EM data.
Cryo-EM reconstruction is reformulated as a stochastic inverse problem to handle structural heterogeneity.
Expectation Maximization (EM) is among the most popular algorithms for maximum likelihood estimation, but it is generally only guaranteed to find its stationary points of the log-likelihood objective. The goal of this article is to present theoretical and empirical evidence that over-parameterization can help EM avoid …
Inference and learning for probabilistic generative networks is often very challenging and typically prevents scalability to as large networks as used for deep discriminative approaches. To obtain efficiently trainable, large-scale and well performing generative networks for semi-supervised learning, we here combine tw…
This paper presents an optimal allocation problem in a financial market with one risk-free and one risky asset, when the market is driven by a stochastic market price of risk. We solve the problem in continuous time, for an investor with a Constant Relative Risk Aversion (CRRA) utility, under two scenarios: when the ma…
Mixed linear regression involves the recovery of two (or more) unknown vectors from unlabeled linear measurements; that is, where each sample comes from exactly one of the vectors, but we do not know which one. It is a classic problem, and the natural and empirically most popular approach to its solution has been the E…
SOLVAR efficiently analyzes cryo-EM data's structural variability.