MM (majorization--minimization) algorithms are an increasingly popular tool for solving optimization problems in machine learning and statistical estimation. This article introduces the MM algorithm framework in general and via three popular example applications: Gaussian mixture regressions, multinomial logistic regre…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Deep-learning improves 6x6-mm OCTA angiograms by reducing noise and artifacts.
Non-convex optimization is ubiquitous in machine learning. Majorization-Minimization (MM) is a powerful iterative procedure for optimizing non-convex functions that works by optimizing a sequence of bounds on the function. In MM, the bound at each iteration is required to \emph{touch} the objective function at the opti…
Unified approach for federated learning using MM optimization.
In this paper, we study a popular method for inference of the Bradley-Terry model parameters, namely the MM algorithm, for maximum likelihood estimation and maximum a posteriori probability estimation. This class of models includes the Bradley-Terry model of paired comparisons, the Rao-Kupper model of paired comparison…
Support vector machines (SVMs) are an important tool in modern data analysis. Traditionally, support vector machines have been fitted via quadratic programming, either using purpose-built or off-the-shelf algorithms. We present an alternative approach to SVM fitting via the majorization--minimization (MM) paradigm. Alg…
Penalized estimation can conduct variable selection and parameter estimation simultaneously. The general framework is to minimize a loss function subject to a penalty designed to generate sparse variable selection. The majorization-minimization (MM) algorithm is a computational scheme for stability and simplicity, and …
Unified algorithm for tensor decomposition supports multiple loss functions and models.
Sufficient dimension reduction (SDR) using distance covariance (DCOV) was recently proposed as an approach to dimension-reduction problems. Compared with other SDR methods, it is model-free without estimating link function and does not require any particular distributions on predictors (see Sheng and Yin, 2013, 2016). …
Paper explores MM strategies that can refuse to quote or provide single-sided quotes.
NS-GAN mode collapse due to sample weighting inversion, solved with MM-nsat.
DGMM improves Gaussian mixture modeling efficiency and stability.
The purpose of this paper is the study of the roots in the mapping class groups. Let be a compact oriented surface, possibly with boundary, let $\PP$ be a finite set of punctures in the interior of , and let $\MM (Σ, \PP)$ denote the mapping class group of $(Σ, \PP)$. We prove that, if is of genus 0, then ea…
Let $(\MM ,{\tilde g})$ be an -dimensional smooth compact Riemannian manifold. We consider the singularly perturbed Allen-Cahn equation $$ ε^2Δ_{ {\tilde g}} {u}\,+\, (1 - {u}^2)u \,=\,0\quad \mbox{in } \MM, $$ where is a small parameter. Let $\KK\subset \MM$ be an -dimensional smooth minimal submanifold …
New algorithm speeds up NMF with -divergence.
New algorithm improves on EM for streaming data, outperforming existing methods.
The paper classifies Poincaré complexes as topological manifolds.
This paper proposes MM-DAGs for analyzing traffic congestion, learning multiple DAGs jointly.
Develops RL for optimal market-making in non-Markov processes.
The Bradley-Terry model is a popular approach to describe probabilities of the possible outcomes when elements of a set are repeatedly compared with one another in pairs. It has found many applications including animal behaviour, chess ranking and multiclass classification. Numerous extensions of the basic model have a…
We present a selective sampling method designed to accelerate the training of deep neural networks. To this end, we introduce a novel measurement, the minimal margin score (MMS), which measures the minimal amount of displacement an input should take until its predicted classification is switched. For multi-class linear…
Proposes a new metric for comparing shapes in different spaces.
A new method trains physics-constrained neural networks more efficiently.
Proposes MM-DUST for efficient generalized lasso solution paths.
Kernel k-Means algorithm improves clustering of non-linear data.
Malignant Pleural Mesothelioma (MPM) or malignant mesothelioma (MM) is an atypical, aggressive tumor that matures into cancer in the pleura, a stratum of tissue bordering the lungs. Diagnosis of MPM is difficult and it accounts for about seventy-five percent of all mesothelioma diagnosed yearly in the United States of …
Proposes MM-KTD for efficient RL learning with reduced sample size.
Mixture-of-experts (MoE) models are a powerful paradigm for modeling of data arising from complex data generating processes (DGPs). In this article, we demonstrate how different MoE models can be constructed to approximate the underlying DGPs of arbitrary types of data. Due to the probabilistic nature of MoE models, we…
The measure concentration property of an mm-space is roughly described as that any 1-Lipschitz map on to a metric space is almost close to a constant map. The target space is called the screen. The case of is widely studied in many literature (see \cite{gromov}, \cite{ledoux}, \cite{mil2}…
We consider Markov models of stochastic processes where the next-step conditional distribution is defined by a kernel density estimator (KDE), similar to Markov forecast densities and certain time-series bootstrap schemes. The KDE Markov models (KDE-MMs) we discuss are nonlinear, nonparametric, fully probabilistic repr…
We consider practical data characteristics underlying federated learning, where unbalanced and non-i.i.d. data from clients have a block-cyclic structure: each cycle contains several blocks, and each client's training data follow block-specific and non-i.i.d. distributions. Such a data structure would introduce client …
MM-DREX adapts LLM experts for financial trading via dynamic routing.
This paper addresses the problem of blind demixing of instantaneous mixtures in a multiple-input multiple-output communication system. The main objective is to present efficient blind source separation (BSS) algorithms dedicated to moderate or high-order QAM constellations. Four new iterative batch BSS algorithms are p…
Let be the energy of some knot for any from certain class of functions. The problem is to find knots with extremal values of energy. We discuss the notion of the locally perturbed knot. The knot circle minimizes some energies and maximizes some others. So, is there any energy such that the circle ne…
This paper considers the problem of robustly estimating a structured covariance matrix with an elliptical underlying distribution with known mean. In applications where the covariance matrix naturally possesses a certain structure, taking the prior structure information into account in the estimation procedure is benef…
We model the behavior of three agent classes acting dynamically in a limit order book of a financial asset. Namely, we consider market makers (MM), high-frequency trading (HFT) firms, and institutional brokers (IB). Given a prior dynamic of the order book, similar to the one considered in the Queue-Reactive models [14,…
New lower bounds improve logistic log-likelihood optimization and inference.
The Expectation-Maximization (EM) algorithm for mixture models often results in slow or invalid convergence. The popular convergence proof affirms that the likelihood increases with Q; Q is increasing in the M -step and non-decreasing in the E-step. The author found that (1) Q may and should decrease in some E-steps; (…
New approach for distributed learning of Gaussian mixtures.
Proposes MELODIC family for simultaneous binary logistic regression.
New algorithm solves fair PCA, robust PCA, and sparse PCA problems efficiently.
A new method estimates expectations from subtractive mixture models without sampling.
New methods for parameter estimation in mechanistic models using data-consistent inversion.
This paper uses deep RL to optimize market quotes from LOB data.
IMM uses imitation learning and predictive representation learning to improve market making strategies.
In this paper we introduce some new copulas emerging from shock models. It was shown earlier that reflected maxmin copulas (RMM for short) are not just some specific singular copulas; they contain many important absolutely continuous copulas including the negative quadrant dependent part of the Eyraud-Farlie-Gumbel-Mor…
Proposes a new method for estimating sparse precision matrices in GMRF-MM models.
Unified NMF models for various noise distributions, improving feature extraction.