End-to-end training of DBMs with improved gradient estimation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We provide initial seedings to the Quick Shift clustering algorithm, which approximate the locally high-density regions of the data. Such seedings act as more stable and expressive cluster-cores than the singleton modes found by Quick Shift. We establish statistical consistency guarantees for this modification. We then…
An image pattern can be represented by a probability distribution whose density is concentrated on different low-dimensional subspaces in the high-dimensional image space. Such probability densities have an astronomical number of local modes corresponding to typical pattern appearances. Related groups of modes can join…
We identify and study two common failure modes for early training in deep ReLU nets. For each we give a rigorous proof of when it occurs and how to avoid it, for fully connected and residual architectures. The first failure mode, exploding/vanishing mean activation length, can be avoided by initializing weights from a …
Sparse-mode DMD disambiguates local and global modes in spatiotemporal data.
Mode connectivity is a recently introduced frame- work that empirically establishes the connected- ness of minima by finding a high accuracy curve between two independently trained models. To investigate the limits of this setup, we examine the efficacy of this technique in extreme cases where the input models are trai…
EDLP samples flat modes in discrete spaces using entropy.
Mathematical analysis shows annealing prevents mode collapse in Gaussian mixtures.
Estimates modes and ridges in mixed Euclidean and directional spaces.
Mixed membership factorization is a popular approach for analyzing data sets that have within-sample heterogeneity. In recent years, several algorithms have been developed for mixed membership matrix factorization, but they only guarantee estimates from a local optimum. Here, we derive a global optimization (GOP) algor…
Optimizes neural network training by dynamically updating Tucker decomposition ranks.
Modified PCA algorithm with continual learning preserves features of previous modes for multimode process monitoring.
cKAM improves adaptive sampling by incorporating a cyclical stepsize scheme.
Develops an MS-inspired algorithm for regression mode finding and space partitioning.
We propose an estimation method for the conditional mode when the conditioning variable is high-dimensional. In the proposed method, we first estimate the conditional density by solving quantile regressions multiple times. We then estimate the conditional mode by finding the maximum of the estimated conditional density…
Deep ensembles have been empirically shown to be a promising approach for improving accuracy, uncertainty and out-of-distribution robustness of deep learning models. While deep ensembles were theoretically motivated by the bootstrap, non-bootstrap ensembles trained with just random initialization also perform well in p…
Simple mode exploration methods do not improve performance in neural networks.
When and why can a neural network be successfully trained? This article provides an overview of optimization algorithms and theory for training neural networks. First, we discuss the issue of gradient explosion/vanishing and the more general issue of undesirable spectrum, and then discuss practical solutions including …
Method for initializing Gaussian mixtures for variational inference with multi-modal distributions.
This paper presents a new way of selecting an initial solution for the k-modes algorithm that allows for a notion of mathematical fairness and a leverage of the data that the common initialisations from literature do not. The method, which utilises the Hospital-Resident Assignment Problem to find the set of initial clu…
BDMBC clusters data with varying densities using a new PLLS measure.
Hybrid model for multimodal distributions using diffusion and classification.
Analyzes learning dynamics of RNNs under locality constraints.
We discuss a natural form of Ricci--flow conjugation between two distinct general relativistic data sets given on a compact -dimensional manifold . We establish the existence of the relevant entropy functionals for the matter and geometrical variables, their monotonicity properties, and the associated conve…
Proves stability of gravitational instantons, proving operator positivity.
Neural networks can learn optimal auction mechanisms and satisfy mode connectivity.
Study on mode stability of gravitational instantons of type D.
We introduce the functional mean-shift algorithm, an iterative algorithm for estimating the local modes of a surrogate density from functional data. We show that the algorithm can be used for cluster analysis of functional data. We propose a test based on the bootstrap for the significance of the estimated local modes …
With the network methods and random matrix theory, we investigate the interaction structure of communities in financial markets. In particular, based on the random matrix decomposition, we clarify that the local interactions between the business sectors (subsectors) are mainly contained in the sector mode. In the secto…
We consider the Yang-Mills flow on hyperbolic 3-space. The gauge connection is constructed from the frame-field and (not necessarily compatible) spin connection components. The fixed points of this flow include zero Yang-Mills curvature configurations, for which the spin connection has zero torsion and the associated R…
Improves DPGMM sampler by better initializing subclusters for more effective clustering.
Koopman mode analysis applied to neural networks for training optimization.
In this paper, we show that Generative Adversarial Networks (GANs) suffer from catastrophic forgetting even when they are trained to approximate a single target distribution. We show that GAN training is a continual learning problem in which the sequence of changing model distributions is the sequence of tasks to the d…
Bayesian Non-negative Matrix Factorization (NMF) is a promising approach for understanding uncertainty and structure in matrix data. However, a large volume of applied work optimizes traditional non-Bayesian NMF objectives that fail to provide a principled understanding of the non-identifiability inherent in NMF-- an i…
New geometric approach realizes 5D bulk theories with 4D edge modes.
Human computer interaction facilitates intelligent communication between humans and computers, in which gesture recognition plays a prominent role. This paper proposes a machine learning system to identify dynamic gestures using tri-axial acceleration data acquired from two public datasets. These datasets, uWave and So…
HiSS sampling overcomes local mode traps in rugged discrete spaces.
Many generative models have to combat . The conventional wisdom to this end is by reducing through training a statistical distance (such as -divergence) between the generated distribution and provided data distribution. But this is more of a heuristic than a guarantee. The statistical distanc…
Bayesian taut splines estimate modes in probability densities.
New principle for supersymmetric localization on Lie groups.
Proposes neuron alignment to optimize mode connectivity in neural networks.
The paper studies kernel smoothing and mean shift for directional data, deriving convergence rates and mode estimation.
New method for Bayesian neural networks reduces inference difficulty.
The paper proposes a method to identify power system oscillation modes using blind source separation.
Jeffreys Flow improves robustness of Boltzmann generators for rare event sampling.
Study reveals issues with neural autoregressive models and proposes mode recovery cost.
Neurons predict future scalar inputs by learning top modes of lag vectors.
New approach to counterfactual reasoning in AI and psychology.