Photography method solves manifold invariants.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper refutes EM convergence theory and introduces a new EM algorithm.
Expectation Maximization (EM) is among the most popular algorithms for estimating parameters of statistical models. However, EM, which is an iterative algorithm based on the maximum likelihood principle, is generally only guaranteed to find stationary points of the likelihood objective, and these points may be far from…
We derive both {\em local} and {\em global} generalized {\em Bianchi identities} for classical Lagrangian field theories on gauge-natural bundles. We show that globally defined generalized Bianchi identities can be found without the {\em a priori} introduction of a connection. The proof is based on a {\em global} decom…
New uncertainty principle limits compression in distributed learning, suggesting optimal methods.
Improves EM algorithm for better local optima in mixture models.
As an automatic method of determining model complexity using the training data alone, Bayesian linear regression provides us a principled way to select hyperparameters. But one often needs approximation inference if distribution assumption is beyond Gaussian distribution. In this paper, we propose a Bayesian linear reg…
The study of solutions with fixed energy of certain classes of Lagrangian (or Hamiltonian) systems is reduced, via the classical Maupertuis--Jacobi variational principle, to the study of geodesics in Riemannian manifolds. We are interested in investigating the problem of existence of brake orbits and homoclinic orbits,…
When a gauge-natural invariant variational principle is assigned, to determine {\em canonical} covariant conservation laws, the vertical part of gauge-natural lifts of infinitesimal principal automorphisms -- defining infinitesimal variations of sections of gauge-natural bundles -- must satisfy generalized Jacobi equat…
Gradient descent on LSE objectives implicitly performs EM, leading to collapse without volume control.
We present a new statistical learning paradigm for Boltzmann machines based on a new inference principle we have proposed: the latent maximum entropy principle (LME). LME is different both from Jaynes maximum entropy principle and from standard maximum likelihood estimation.We demonstrate the LME principle BY deriving …
Learning with hidden variables is a central challenge in probabilistic graphical models that has important implications for many real-life problems. The classical approach is using the Expectation Maximization (EM) algorithm. This algorithm, however, can get trapped in local maxima. In this paper we explore a new appro…
In this paper, we provide an information-theoretic interpretation of the Vector Quantized-Variational Autoencoder (VQ-VAE). We show that the loss function of the original VQ-VAE can be derived from the variational deterministic information bottleneck (VDIB) principle. On the other hand, the VQ-VAE trained by the Expect…
Optimizing distributed learning systems is an art of balancing between computation and communication. There have been two lines of research that try to deal with slower networks: {\em communication compression} for low bandwidth networks, and {\em decentralization} for high latency networks. In this paper, We explore a…
Why does Deep Learning work? What representations does it capture? How do higher-order representations emerge? We study these questions from the perspective of group theory, thereby opening a new approach towards a theory of Deep learning. One factor behind the recent resurgence of the subject is a key algorithmic step…
This paper uses dynamical systems to analyze and ensure convergence of the Bayesian EM algorithm.
New method optimizes clustering with better log-likelihood landscape.
This paper establishes a statistical versus computational trade-off for solving a basic high-dimensional machine learning problem via a basic convex relaxation method. Specifically, we consider the {\em Sparse Principal Component Analysis} (Sparse PCA) problem, and the family of {\em Sum-of-Squares} (SoS, aka Lasserre/…
Paper derives constraints for Bayesian Knowledge Tracing parameters.
Enhances large language models' reasoning through simpler off-policy reinforcement learning.
SOLVAR efficiently analyzes cryo-EM data's structural variability.
In this paper we propose solving localized multiple kernel learning (LMKL) using LMKL-Net, a feedforward deep neural network. In contrast to previous works, as a learning principle we propose {\em parameterizing} both the gating function for learning kernel combination weights and the multiclass classifier in LMKL usin…
A {\em blink} is a plane graph with an arbitrary bipartition of its edges. As a consequence of a recent result of Martelli, I show that the homeomorphisms classes of closed oriented 3-manifolds are in 1-1 correspondence with specific classes of blinks. In these classes, two blinks are equivalent if they are linked by a…
A framework for efficient multi-objective optimization using entropy search.
Why does Deep Learning work? What representations does it capture? How do higher-order representations emerge? We study these questions from the perspective of group theory, thereby opening a new approach towards a theory of Deep learning. One factor behind the recent resurgence of the subject is a key algorithmic step…
We consider the problem of discriminative factor analysis for data that are in general non-Gaussian. A Bayesian model based on the ranks of the data is proposed. We first introduce a new {\em max-margin} version of the rank-likelihood. A discriminative factor model is then developed, integrating the max-margin rank-lik…
Mirror descent method improved RL algorithms.
Many real world problems can now be effectively solved using supervised machine learning. A major roadblock is often the lack of an adequate quantity of labeled data for training. A possible solution is to assign the task of labeling data to a crowd, and then infer the true label using aggregation methods. A well-known…
In Biology, all motor enzymes operate on the same principle: they trap favourable brownian fluctuations in order to generate directed forces and to move. Whether it is possible or not to copy one such strategy to play the market was the starting point of our investigations. We found the answer is yes. In this paper we …
Study optimal offline RL with uncertainty sets and distribution shifts.
The K-Mean and EM algorithms are popular in clustering and mixture modeling, due to their simplicity and ease of implementation. However, they have several significant limitations. Both coverage to a local optimum of their respective objective functions (ignoring the uncertainty in the model space), require the apriori…
New findings show sparse signals in MRA model require fewer measurements than previously thought.
Data-driven anomaly detection methods suffer from the drawback of detecting all instances that are statistically rare, irrespective of whether the detected instances have real-world significance or not. In this paper, we are interested in the problem of specifically detecting anomalous instances that are known to have …
Improved pricing of vanilla options using modified Adams method and sinh-acceleration.
Paper addresses global convergence of MLR estimation under weak data conditions.
Paper proposes a faster SPIDER-EM variant for large-scale nonconvex optimization.
Let be either a Bernoulli random walk or a Brownian motion with drift, and let , . This paper solves the general optimal prediction problem \sup_{0\leqτ\leq T}\sE[f(M_T-B_τ)], where the supremum is over all stopping times adapted to the natural…
We consider the geometric formulation of the Hamiltonian formalism for field theory in terms of {\em Hamiltonian connections} and {\em multisymplectic forms}. In this framework the covariant Hamilton equations for Mechanics and field theory are defined in terms of multisymplectic --forms, where is the dimens…
EGMM improves clustering by better handling uncertainty with evidential framework.
A new framework predicts hidden Markov model regimes online.
The EM algorithm is one of many important tools in the field of statistics. While often used for imputing missing data, its widespread applications include other common statistical tasks, such as clustering. In clustering, the EM algorithm assumes a parametric distribution for the clusters, whose parameters are estimat…
We develop a probabilistic framework for deep learning based on the Deep Rendering Mixture Model (DRMM), a new generative probabilistic model that explicitly capture variations in data due to latent task nuisance variables. We demonstrate that max-sum inference in the DRMM yields an algorithm that exactly reproduces th…
Bayesian networks (BN) are used in a big range of applications but they have one issue concerning parameter learning. In real application, training data are always incomplete or some nodes are hidden. To deal with this problem many learning parameter algorithms are suggested foreground EM, Gibbs sampling and RBE algori…
Graph convolutional networks (GCNs) suffer from the irregularity of graphs, while more widely-used convolutional neural networks (CNNs) benefit from regular grids. To bridge the gap between GCN and CNN, in contrast to previous works on generalizing the basic operations in CNNs to graph data, in this paper we address th…
Gradient EM converges globally for over-parameterized Gaussian mixtures.
EM algorithm converges in KL divergence for exponential families via mirror descent.
The study characterizes Hermitian manifolds with parallel Bismut-Strominger torsion.
sEM uses optimal transport to improve EM algorithm for better convergence and avoiding local optima.