JaxSGMC simplifies SG-MCMC for Bayesian deep learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study shows spherical embedding and immersion components are related to homotopy groups.
SG-PALM learns interpretable tensor models for high-dimensional data.
SG-NTF completes HDI tensors with spectral mapping and spatio-temporal gating.
Stochastic gradient Markov chain Monte Carlo (SG-MCMC) has been increasingly popular in Bayesian learning due to its ability to deal with large data. A standard SG-MCMC algorithm simulates samples from a discretized-time Markov chain to approximate a target distribution. However, the samples are typically highly correl…
Stochastic gradient MCMC (SG-MCMC) has played an important role in large-scale Bayesian learning, with well-developed theoretical convergence properties. In such applications of SG-MCMC, it is becoming increasingly popular to employ distributed systems, where stochastic gradients are computed based on some outdated par…
We define strongly Gauduchon spaces and the class SG which are generalization of strongly Gauduchon manifolds in complex spaces. Comparing with the case of Kahlerian, the strongly Gauduchon space and the class SG are similar to the Kahler space and the Fujiki class C respectively. Some properties about these complex sp…
There has been recent interest in developing scalable Bayesian sampling methods such as stochastic gradient MCMC (SG-MCMC) and Stein variational gradient descent (SVGD) for big-data analysis. A standard SG-MCMC algorithm simulates samples from a discrete-time Markov chain to approximate a target distribution, thus samp…
Unified framework for isotropic SG noise in posterior sampling.
Stochastic gradient Markov Chain Monte Carlo (SG-MCMC) has been developed as a flexible family of scalable Bayesian sampling algorithms. However, there has been little theoretical analysis of the impact of minibatch size to the algorithm's convergence rate. In this paper, we prove that under a limited computational bud…
Stochastic gradient Markov chain Monte Carlo (SG-MCMC) has become increasingly popular for simulating posterior samples in large-scale Bayesian modeling. However, existing SG-MCMC schemes are not tailored to any specific probabilistic model, even a simple modification of the underlying dynamical system requires signifi…
In an earlier paper (math.SG/0110169), we introduced absolute gradings on the three-manifold invariants developed in math.SG/0101206 and math.SG/0105202. Coupled with the surgery long exact sequences, we obtain a number of three- and four-dimensional applications of this absolute grading including strengthenings of the…
Streaming variational Bayes (SVB) is successful in learning LDA models in an online manner. However previous attempts toward developing online Monte-Carlo methods for LDA have little success, often by having much worse perplexity than their batch counterparts. We present a streaming Gibbs sampling (SGS) method, an onli…
This is the second in a series of papers on a new equivariant cohomology that takes values in a vertex algebra. In an earlier paper, the first two authors gave a construction of the cohomology functor on the category of O(sg) algebras. The new cohomology theory can be viewed as a kind of "chiralization'' of the classic…
This paper surveys ML applications in SG for cyberattacks.
The singular braids with strands, , were introduced independently by Baez and Birman. It is known that the monoid formed by the singular braids is embedded in a group that is known as singular braid group, denoted by . There has been another generalization of braid groups, denoted by , $n \ge…
A new method solves l1-regularized optimization problems efficiently and sparsely.
Recent advances in Bayesian learning with large-scale data have witnessed emergence of stochastic gradient MCMC algorithms (SG-MCMC), such as stochastic gradient Langevin dynamics (SGLD), stochastic gradient Hamiltonian MCMC (SGHMC), and the stochastic gradient thermostat. While finite-time convergence properties of th…
Study shows a modified cobordism category's first derivative is equivalent to a Thom spectrum.
Stochastic gradient MCMC (SG-MCMC) algorithms have proven useful in scaling Bayesian inference to large datasets under an assumption of i.i.d data. We instead develop an SG-MCMC algorithm to learn the parameters of hidden Markov models (HMMs) for time-dependent data. There are two challenges to applying SG-MCMC in this…
We review the relations between compact complex manifolds carrying various types of Hermitian metrics (Kähler, balanced or {\it strongly Gauduchon}) and those satisfying the -lemma or the degeneration at of the Frölicher spectral sequence, as well as the behaviour of these properties under h…
An ensemble of neural networks is known to be more robust and accurate than an individual network, however usually with linearly-increased cost in both training and testing. In this work, we propose a two-stage method to learn Sparse Structured Ensembles (SSEs) for neural networks. In the first stage, we run SG-MCMC wi…
We propose a unifying view of two different Bayesian inference algorithms, Stochastic Gradient Markov Chain Monte Carlo (SG-MCMC) and Stein Variational Gradient Descent (SVGD), leading to improved and efficient novel sampling schemes. We show that SVGD combined with a noise term can be framed as a multiple chain SG-MCM…
In this paper, we define the set of singular grid diagrams which provides a unified description for singular links, singular Legendrian links, singular transverse links, and singular braids. We also classify the complete set of all equivalence relations on which induce the bijection onto e…
New hyperbolicity concepts expand manifold study.
Continuing the program of math.SG/0012067 and math.SG/0310450, we introduce refinements of the Donaldson-Smith standard surface count which are designed to count nodal pseudoholomorphic curves and curves with a prescribed decomposition into reducible components. In cases where a corresponding analogue of the Gromov-Tau…
The aim of this article is to introduce invariants of oriented, smooth, closed four-manifolds, built using the Floer homology theories defined in two earlier papers (math.SG/0101206 and math.SG/0105202). This four-dimensional theory also endows the corresponding three-dimensional theories with additional structure: an …
In math.SG/0303255, we discussed the connected components of the space of surface group representations for any compact connected semisimple Lie group and any closed compact (orientable or nonorientable) surface. In this sequel, we generalize the results in math.SG/0303255 in two directions: we consider general compact…
Study assesses data-driven and physics-based SGS models for transcritical combustion.
Proposes SNML for selecting word2vec Skip-gram dimensionality.
In this paper, we study the efficiency of a {\bf R}estarted {\bf S}ub{\bf G}radient (RSG) method that periodically restarts the standard subgradient method (SG). We show that, when applied to a broad class of convex optimization problems, RSG method can find an -optimal solution with a lower complexity than the SG m…
Significant success has been realized recently on applying machine learning to real-world applications. There have also been corresponding concerns on the privacy of training data, which relates to data security and confidentiality issues. Differential privacy provides a principled and rigorous privacy guarantee on mac…
We propose a stochastic gradient Markov chain Monte Carlo (SG-MCMC) algorithm for scalable inference in mixed-membership stochastic blockmodels (MMSB). Our algorithm is based on the stochastic gradient Riemannian Langevin sampler and achieves both faster speed and higher accuracy at every iteration than the current sta…
The paper defines subgroups of camomile type and studies singular braids and links.
Recent growing adoption of experimentation in practice has led to a surge of attention to multiarmed bandits as a technique to reduce the opportunity cost of online experiments. In this setting, a decision-maker sequentially chooses among a set of given actions, observes their noisy rewards, and aims to maximize her cu…
We propose the stochastic average gradient (SAG) method for optimizing the sum of a finite number of smooth convex functions. Like stochastic gradient (SG) methods, the SAG method's iteration cost is independent of the number of terms in the sum. However, by incorporating a memory of previous gradient values the SAG me…
We consider Model-Agnostic Meta-Learning (MAML) methods for Reinforcement Learning (RL) problems, where the goal is to find a policy using data from several tasks represented by Markov Decision Processes (MDPs) that can be updated by one step of stochastic policy gradient for the realized MDP. In particular, using stoc…
The posteriors over neural network weights are high dimensional and multimodal. Each mode typically characterizes a meaningfully different representation of the data. We develop Cyclical Stochastic Gradient MCMC (SG-MCMC) to automatically explore such distributions. In particular, we propose a cyclical stepsize schedul…
We describe a method for learning word embeddings with data-dependent dimensionality. Our Stochastic Dimensionality Skip-Gram (SD-SG) and Stochastic Dimensionality Continuous Bag-of-Words (SD-CBOW) are nonparametric analogs of Mikolov et al.'s (2013) well-known 'word2vec' models. Vector dimensionality is made dynamic b…
New streaming methods improve convergence rates for optimization problems.
Recently, Stochastic Gradient Markov Chain Monte Carlo (SG-MCMC) methods have been proposed for scaling up Monte Carlo computations to large data problems. Whilst these approaches have proven useful in many applications, vanilla SG-MCMC might suffer from poor mixing rates when random variables exhibit strong couplings …
Artifical Neural Networks are a particular class of learning systems modeled after biological neural functions with an interesting penchant for Hebbian learning, that is "neurons that wire together, fire together". However, unlike their natural counterparts, artificial neural networks have a close and stringent couplin…
We study a new aggregation operator for gradients coming from a mini-batch for stochastic gradient (SG) methods that allows a significant speed-up in the case of sparse optimization problems. We call this method AdaBatch and it only requires a few lines of code change compared to regular mini-batch SGD algorithms. We p…
Corrected CBOW performs similarly to Skip-gram.
Based on our recent adaptation of the adiabatic limit construction to the case of complex structures, we prove the fact that the deformation limiting manifold of any holomorphic family of Moishezon manifolds is Moishezon. Two new ingredients, hopefully of independent interest, are introduced. The first one associates w…
A new approach RA improves stochastic optimization by executing multiple steps between subsample updates.
Stochastic gradient Markov chain Monte Carlo (SG-MCMC) methods are Bayesian analogs to popular stochastic optimization methods; however, this connection is not well studied. We explore this relationship by applying simulated annealing to an SGMCMC algorithm. Furthermore, we extend recent SG-MCMC methods with two key co…
Asynchronous parallel implementations of stochastic gradient (SG) have been broadly used in solving deep neural network and received many successes in practice recently. However, existing theories cannot explain their convergence and speedup properties, mainly due to the nonconvexity of most deep learning formulations …