Non-convex optimization is ubiquitous in machine learning. Majorization-Minimization (MM) is a powerful iterative procedure for optimizing non-convex functions that works by optimizing a sequence of bounds on the function. In MM, the bound at each iteration is required to \emph{touch} the objective function at the opti…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
MM (majorization--minimization) algorithms are an increasingly popular tool for solving optimization problems in machine learning and statistical estimation. This article introduces the MM algorithm framework in general and via three popular example applications: Gaussian mixture regressions, multinomial logistic regre…
Unified approach for federated learning using MM optimization.
Proposes MM-KTD for efficient RL learning with reduced sample size.
Deep-learning improves 6x6-mm OCTA angiograms by reducing noise and artifacts.
Penalized estimation can conduct variable selection and parameter estimation simultaneously. The general framework is to minimize a loss function subject to a penalty designed to generate sparse variable selection. The majorization-minimization (MM) algorithm is a computational scheme for stability and simplicity, and …
MM-DREX adapts LLM experts for financial trading via dynamic routing.
Paper explores MM strategies that can refuse to quote or provide single-sided quotes.
IMM uses imitation learning and predictive representation learning to improve market making strategies.
NS-GAN mode collapse due to sample weighting inversion, solved with MM-nsat.
New methods for parameter estimation in mechanistic models using data-consistent inversion.
The purpose of this paper is the study of the roots in the mapping class groups. Let be a compact oriented surface, possibly with boundary, let $\PP$ be a finite set of punctures in the interior of , and let $\MM (Σ, \PP)$ denote the mapping class group of $(Σ, \PP)$. We prove that, if is of genus 0, then ea…
Develops RL for optimal market-making in non-Markov processes.
Let $(\MM ,{\tilde g})$ be an -dimensional smooth compact Riemannian manifold. We consider the singularly perturbed Allen-Cahn equation $$ ε^2Δ_{ {\tilde g}} {u}\,+\, (1 - {u}^2)u \,=\,0\quad \mbox{in } \MM, $$ where is a small parameter. Let $\KK\subset \MM$ be an -dimensional smooth minimal submanifold …
The paper classifies Poincaré complexes as topological manifolds.
This paper proposes MM-DAGs for analyzing traffic congestion, learning multiple DAGs jointly.
In this paper, we study a popular method for inference of the Bradley-Terry model parameters, namely the MM algorithm, for maximum likelihood estimation and maximum a posteriori probability estimation. This class of models includes the Bradley-Terry model of paired comparisons, the Rao-Kupper model of paired comparison…
We present a selective sampling method designed to accelerate the training of deep neural networks. To this end, we introduce a novel measurement, the minimal margin score (MMS), which measures the minimal amount of displacement an input should take until its predicted classification is switched. For multi-class linear…
Proposes MM-DUST for efficient generalized lasso solution paths.
Malignant Pleural Mesothelioma (MPM) or malignant mesothelioma (MM) is an atypical, aggressive tumor that matures into cancer in the pleura, a stratum of tissue bordering the lungs. Diagnosis of MPM is difficult and it accounts for about seventy-five percent of all mesothelioma diagnosed yearly in the United States of …
Bayesian optimization for function-valued responses, addressing worst case deviations.
Support vector machines (SVMs) are an important tool in modern data analysis. Traditionally, support vector machines have been fitted via quadratic programming, either using purpose-built or off-the-shelf algorithms. We present an alternative approach to SVM fitting via the majorization--minimization (MM) paradigm. Alg…
The measure concentration property of an mm-space is roughly described as that any 1-Lipschitz map on to a metric space is almost close to a constant map. The target space is called the screen. The case of is widely studied in many literature (see \cite{gromov}, \cite{ledoux}, \cite{mil2}…
Unified algorithm for tensor decomposition supports multiple loss functions and models.
Let be the energy of some knot for any from certain class of functions. The problem is to find knots with extremal values of energy. We discuss the notion of the locally perturbed knot. The knot circle minimizes some energies and maximizes some others. So, is there any energy such that the circle ne…
Mixture-of-experts (MoE) models are a powerful paradigm for modeling of data arising from complex data generating processes (DGPs). In this article, we demonstrate how different MoE models can be constructed to approximate the underlying DGPs of arbitrary types of data. Due to the probabilistic nature of MoE models, we…
We model the behavior of three agent classes acting dynamically in a limit order book of a financial asset. Namely, we consider market makers (MM), high-frequency trading (HFT) firms, and institutional brokers (IB). Given a prior dynamic of the order book, similar to the one considered in the Queue-Reactive models [14,…
DGMM improves Gaussian mixture modeling efficiency and stability.
Sufficient dimension reduction (SDR) using distance covariance (DCOV) was recently proposed as an approach to dimension-reduction problems. Compared with other SDR methods, it is model-free without estimating link function and does not require any particular distributions on predictors (see Sheng and Yin, 2013, 2016). …
This paper uses deep RL to optimize market quotes from LOB data.
This paper considers the problem of robustly estimating a structured covariance matrix with an elliptical underlying distribution with known mean. In applications where the covariance matrix naturally possesses a certain structure, taking the prior structure information into account in the estimation procedure is benef…
In this paper we introduce some new copulas emerging from shock models. It was shown earlier that reflected maxmin copulas (RMM for short) are not just some specific singular copulas; they contain many important absolutely continuous copulas including the negative quadrant dependent part of the Eyraud-Farlie-Gumbel-Mor…
A new method trains physics-constrained neural networks more efficiently.
Proposes a new method for estimating sparse precision matrices in GMRF-MM models.
New algorithm speeds up NMF with -divergence.
We consider Markov models of stochastic processes where the next-step conditional distribution is defined by a kernel density estimator (KDE), similar to Markov forecast densities and certain time-series bootstrap schemes. The KDE Markov models (KDE-MMs) we discuss are nonlinear, nonparametric, fully probabilistic repr…
RLMM extends psychometric models to larger tasks.
Proposes a new metric for comparing shapes in different spaces.
Study models illiquid stock prices and finds low correlation due to constant prices.
Unified NMF models for various noise distributions, improving feature extraction.
Some aspects of the multidimensional soliton geometry are considered. The relation between soliton equations in 2+1 dimensions and the Self-Dual Yang-Mills and Bogomolny equations are discussed.
Deformation estimation of elastic object assuming an internal organ is important for the computer navigation of surgery. The aim of this study is to estimate the deformation of an entire three-dimensional elastic object using displacement information of very few observation points. A learning approach with a neural net…
Exponential Lasso improves Lasso's robustness to outliers and heavy-tailed noise.
We revisit and demonstrate the Epps effect using two well-known non-parametric covariance estimators; the Malliavin and Mancino (MM), and Hayashi and Yoshida (HY) estimators. We show the existence of the Epps effect in the top 10 stocks from the Johannesburg Stock Exchange (JSE) by various methods of aggregating Trade …
Leveraging on the convexity of the Lasso problem , screening rules help in accelerating solvers by discarding irrelevant variables, during the optimization process. However, because they provide better theoretical guarantees in identifying relevant variables, several non-convex regularizers for the Lasso have been prop…
A new framework for predictive clustering and optimization.
New method uses subtractive mixture models for approximate inference.
New algorithm improves on EM for streaming data, outperforming existing methods.