Advances scalable clustering and density mode finding.
problem Joint clustering and density mode finding for large datasets.
method Proposes a concave-convex relaxation with parallel algorithm optimization.
result Optimizes a tight bound for cluster assignment variables.
Study on learning dynamics in deep neural networks, proving properties and confirming empirical observations.
problem Understanding the learning dynamics of deep neural networks, especially in binary classification.
method Proved properties of learning in binary classification under strong assumptions, extended existing results from linear case.
result Classification error follows a sigmoidal shape in nonlinear architectures, parallel independent modes of learning, gradient starvation phenomenon.
Jeffreys Flow improves robustness of Boltzmann generators for rare event sampling.
problem Rare events and metastable trapping in sampling physical systems with rough energy landscapes.
method Introduces Jeffreys Flow, a robust generative framework using Parallel Tempering distillation and symmetric Jeffreys divergence to mitigate mode collapse and improve mode coverage.
result Minimizing Jeffreys divergence suppresses mode collapse and corrects inaccuracies in multi-modal distributions.
Proposes a method to improve SLMC for multimodal distributions.
problem Difficulty of applying SLMC to multimodal distributions.
method Parallel adaptive annealing with VAE-SLMC.
result Can proficiently obtain accurate samples from multimodal distributions.
The paper analyzes zero modes on product manifolds and provides estimates for their norms.
problem Analyzing zero modes on product Riemannian manifolds.
method Using the zero mode equation and non-increasing condition on |\varphi|, the paper derives estimates for the norms of the vector field A.
result The derived estimates are sharp in even dimensions and provide insights into the behavior of zero modes.
New sampler tackles noisy posterior sampling with multiple modes.
problem Efficient sampling from complex posterior distributions with multiple isolated modes.
method Integrates parallel tempering with Nosé-Hoover dynamics for stochastic gradient.
result Efficiently draws representative samples from noisy posterior distributions.
A new method combines MCMC results to avoid failures in parallel computing.
problem Parallel MCMC's sensitivity to subposterior sampling issues leads to failures.
method Parallel Active Inference (PAI) uses Gaussian Process (GP) surrogate modeling and active learning.
result PAI successfully combines MCMC results where previous methods fail.
Paper proves a spinor inequality for magnetic fields on spin manifolds.
problem Proving a spinor inequality for magnetic fields on spin manifolds.
method Analyzing the zero mode equation and using the Yamabe constant.
result The inequality ∥dA∥n/2>Y(Mn,[g])/(4vn1/2) holds for non-trivial solutions. Generative Adversarial Networks have become one of the most studied frameworks for unsupervised learning due to their intuitive formulation. They have also been shown to be capable of generating convincing examples in limited domains, such as low-resolution images. However, they still prove difficult to train in practi…
A new training method speeds up ResNet training by 3x with minimal accuracy loss.
problem Training ResNets is slow due to dependencies between modules.
method Serial-parallel hybrid training strategy with data augmentation and downsampling.
result Significant speedup over traditional methods with comparable accuracy.
Estimates eigenvalues of modified Dirac operators with multi-forms.
problem Estimating eigenvalues of modified Dirac operators with multi-forms.
method Analyzes eigenvalues of multi-form modified Dirac operators constructed from a standard Dirac operator.
result Provides estimates for eigenvalues of modified Dirac operators with multi-forms.
Characterizes neutral deformation modes of minimal surfaces.
problem Understanding the energy content of deformation modes of minimal surfaces.
method Analyzes the energy content of stretching, drilling, and bending modes of minimal surfaces.
result All isometries of a minimal surface are globally neutral and give rise to soft elasticity.
This work tackles GAN training instability through parallel tempering.
problem Training instability and mode collapse in GANs.
method Introduces a parallel tempering framework to stabilize GAN training.
result Significantly reduces gradient variance and improves training efficiency.
A fundamental aspect of biological information processing is the ubiquity of sequence-function relationships -- functions that map the sequence of DNA, RNA, or protein to a biochemically relevant activity. Most sequence-function relationships in biology are quantitative, but only recently have experimental techniques f…
A new method normalizes activations to match batch normalization without batch dependence.
problem Performance degradation with batch-independent normalization techniques.
method Proxy-Normalizing Activations
result Proxy-Normalization technique emulates batch normalization's behavior and performance.
Combines local MCMC chains to speed up sampling.
problem Designing efficient MCMC chains that mix well over the whole state space.
method Combining parallel chains prioritized by kernel Stein discrepancy, combining samples using novel probability estimation.
result Significant speedups in sampling from multimodal distributions.
New method trains neural samplers without simulation, but fails due to mode collapse.
problem Training neural samplers without simulation.
method Time-dependent normalizing flow with Langevin preconditioning.
result Langevin preconditioning is crucial for avoiding mode collapse.
Paper discovers simplicial complexes connecting trained models for improved ensembling.
problem Improving robustness and accuracy of deep learning ensembles.
method Identifies mode-connecting simplicial complexes on loss surfaces.
result Efficiently builds simplicial complexes for ensembling, outperforming independent ensembles.
New tool detects 'fleeting modes' causing excess risk in financial markets.
problem Detecting portfolios with statistically significant excess risk in financial markets.
method Random Matrix Theory to identify 'fleeting modes' independent of underlying correlation structure.
result Fleeting modes exist in both futures and equity markets, and momentum is a source of excess risk.
Gradient-guided nested sampling improves posterior inference efficiency.
problem Efficiently sampling from complex posterior distributions.
method Gradient-guided nested sampling combining differentiable programming, Hamiltonian slice sampling, clustering, mode separation, dynamic nested sampling, and parallelization.
result Significantly faster mode discovery and more accurate partition function estimates.
pLSTM tackles long-range language modeling and computer vision tasks with parallelizable linear source transition mark networks.
problem Challenges of existing recurrent architectures in handling sequences and multi-dimensional data.
method Introduces pLSTM, a parallelizable linear source transition mark network for linear graphs and DAGs, addressing vanishing/exploding activation/gradient issues.
result pLSTM outperforms Transformers in long-range tasks like arrow-pointing extrapolation and image size extrapolation.
DMD separates mixed time series with uncorrelated components.
problem Separating mixed time series with uncorrelated components.
method Dynamic Mode Decomposition (DMD) applied to a data matrix of mixed time series.
result DMD can approximate the mixing matrix of uncorrelated time series.
ICAL improves deep learning model accuracy and NLL with optimized batch labeling.
problem Deep Bayesian Active Learning for efficient model training.
method ICAL uses HSIC to measure dependency and optimizes batch size scaling.
result Significant improvements in model accuracy and NLL on image datasets.
The paper improves GP regression for sparse sensor data in structural mode shape reconstruction.
problem Reconstructing full-field structural mode shapes from sparse sensor data.
method Physics-Constrained Single-Output Gaussian Process (CONS-SOGP) framework.
result The proposed method provides more accurate and reliable mode shapes.
A parallel Fortran framework for neural networks and deep learning.
problem Developing efficient parallel Fortran for neural networks and deep learning.
method Simple interface, activation functions, stochastic gradient descent, Fortran 2018 collective subroutines, parallelism with derived types and collective operations.
result Ease of use and computational performance similar to existing machine learning frameworks, suitable for production.
Improved lower bound for parallel tempering's mixing time.
problem Slow convergence and mixing in multimodal target distributions.
method Presented a new lower bound for the spectral gap of parallel tempering.
result Improved the best existing bound on spectral gap with polynomial dependence on parameters.
New method samples from complex Bayesian posteriors efficiently.
problem Challenges in Bayesian inference for complex datasets.
method Bayesian nonparametric learning with posterior bootstrap.
result Scalable sampling from multimodal posteriors.
In this paper we investigate the relationship between the existence of parallel semi-Riemannian metrics of a connection and the reducibility of the associated holonomy group. The question as to whether the holonomy group necessarily reduces in the presence of a specified number of independent parallel semi-Riemannian m…
Stacking improves inference for multimodal Bayesian posterior distributions.
problem Difficulty of MCMC in moving between modes and underestimation of posterior uncertainty.
method Parallel runs of MCMC, variational, or mode-based inference, combined using Bayesian stacking.
result Stacking efficiently samples from multimodal posterior distributions and represents uncertainty better than variational inference.
Unified model predicts multi-mode failure with multi-sensor data.
problem Independent failure mode and RUL prediction ignores inherent relationship.
method Hierarchical Bayesian framework with Cox model, Gaussian process, and multinomial distributions.
result Robust uncertainty quantification and accurate prediction of multi-mode failure.
Parallel score matching accelerates DPM training and improves density estimation.
problem Extended training periods and limited modeling flexibility in DPMs.
method Partitioning the learning task into independent time sub-intervals and modeling the score at each time point separately.
result Significant acceleration of training process and improved density estimation performance.
Mode connectivity framework shows resilience to model differences.
problem Investigating the limits of mode connectivity in extreme cases.
method Empirically examines mode connectivity between differently trained models.
result The procedure is resilient to changes in model training or initialization.
Communication costs, resulting from synchronization requirements during learning, can greatly slow down many parallel machine learning algorithms. In this paper, we present a parallel Markov chain Monte Carlo (MCMC) algorithm in which subsets of data are processed independently, with very little communication. First, w…
We introduce a novel approach for parallelizing MCMC inference in models with spatially determined conditional independence relationships, for which existing techniques exploiting graphical model structure are not applicable. Our approach is motivated by a model of seismic events and signals, where events detected in d…
FastKCI speeds up KCI tests for causal inference on large datasets.
problem Cubic computational complexity of kernel-based conditional independence tests.
method Mixture-of-experts approach with parallel Gaussian process inference.
result Substantial computational speedups with maintained statistical power.
LoRA-Curve connects independent LoRA optima through continuous low-loss valleys, improving Bayesian model averaging.
problem Challenges in estimating epistemic uncertainty in LoRA-based Bayesian inference.
method Introduces LoRA-Curve, a segmented Bézier curve parameterization in the LoRA space, with free and anchored configurations.
result Empirically shows that connecting independent LoRA optima through continuous low-loss valleys improves mutual information of the predictive distribution.
In this addendum to our article "Superconnections and Parallel Transport" we give an alternate construction to the parallel transport of a superconnection contained in Corollary 4.4 of \cite{D1}, which has the advantage that is independent on the various ways a superconnection splits as a connection plus a bundle endom…
K-Metamodes clusters security data without converting categorical attributes.
problem Clustering heterogeneous security data sets with categorical and numerical attributes.
method Frequency-based distance function for ensemble-based k-modes clustering, adapted feature discretisation.
result Higher effectiveness compared to previous methods on public security data sets.
ICSGLD improves efficiency in posterior sampling for big data.
problem Efficient posterior sampling for large datasets.
method Embarrassingly parallel multiple-chain CSGLD with efficient interactions.
result ICSGLD is more efficient than a single-chain CSGLD.
Neurons predict future scalar inputs by learning top modes of lag vectors.
problem Predicting future scalar inputs with physiological delays.
method Normal Mode Decomposition to extract independently evolving modes.
result Temporal filters of neurons correspond to left eigenvectors of a generalized eigenvalue problem.
FPETS speeds up TTS by 600X and reduces errors.
problem High latency and errors in end-to-end TTS systems.
method Non-autoregressive, fully parallel approach with UFANS and trainable position encoding.
result Significant speed up and better quality audios with fewer errors.
Using the theory of extensors developed in a previous paper we present a theory of the parallelism structure on arbitrary smooth manifold. Two kinds of Cartan connection operators are introduced and both appear in intrinsic versions (i.e., frame independent) of the first and second Cartan structure equations. Also, the…
With the rapidly growing scales of statistical problems, subset based communication-free parallel MCMC methods are a promising future for large scale Bayesian analysis. In this article, we propose a new Weierstrass sampler for parallel MCMC based on independent subsets. The new sampler approximates the full data poster…
We apply belief propagation to a Bayesian bipartite graph composed of discrete independent hidden variables and discrete visible variables. The network is the Discrete counterpart of Independent Component Analysis (DICA) and it is manipulated in a factor graph form for inference and learning. A full set of simulations …
New method predicts RUL and failure modes from partial data.
problem Predicting RUL and failure modes from incomplete data.
method Formulated as vector General Value Function (GVF) prediction on an absorbing degradation process, using TD(n,λ) for estimation. result TD improves RUL and failure-mode prediction compared to Monte Carlo methods, especially under scarce complete labels.
Dropout neural networks can approximate any function with high probability.
problem Approximating functions with dropout neural networks.
method Two universal approximation theorems for dropout neural networks in random and deterministic modes.
result Dropout neural networks can approximate any function in probability and in Lq. A new dynamical formulation of log-PCA captures local principal modes of geodesic variations.
problem Learning principal variations of random probability measures under Wasserstein geometry.
method Introducing a new dynamical formulation of log-PCA as a variational approach.
result Deriving a general statistical convergence rate for empirical WT-PCA.
Proposes batch-weight method to correct mode imbalance in domain adaptation.
problem Mode imbalance between source and target distributions affects unsupervised domain transfer.
method Proposes batch-weight method to re-weight training samples.
result Effective in several image-to-image translation tasks.