SIMD operations boost Bayesian computations up to 6x faster.
problem Expensive Bayesian computations are computationally intensive and parallelizable.
method Demonstrated the utility of SIMD operations for Bayesian applications using standard libraries.
result Up to 6x improvement in floating point arithmetic performance.
This work discusses AAD for financial model calibration and its parallelization benefits.
problem Calibrating stochastic financial models using Automatic Adjoint Differentiation.
method Demonstrates the use of Automatic Adjoint Differentiation for functions in financial models and its parallelization potential.
result Theoretical and numeric results show that AAD allows perfect SIMD parallelization and is efficient.
Paper speeds up matrix multiplication on Intel PIII using SIMD.
problem Efficiently multiplying large matrices for faster algorithm performance.
method Implemented matrix-matrix multiply using Intel Pentium SIMD architecture.
result Average performance 2.09 times faster than public domain routines.
Modified BFGS and LBFGS++ libraries boost performance for non-parallelizable functions.
problem Improving performance of non-parallelizable functions using SIMD and AAD.
method Modifications to BFGS and LBFGS++ libraries, utilizing SIMD and Automatic Differentiation (AAD).
result Up to 3.8 times faster for European Swaption curve calibration and 1.4 times faster for LMM model calibration.
Improves matrix multiplication throughput for asymmetric bit-width operands.
problem Matrix multiplications between asymmetric bit-width operands, especially 8- and 4-bit, are not efficiently handled by existing SIMD instructions.
method Proposes a new SIMD matrix multiplication instruction that uses mixed precision on inputs (8- and 4-bit) and accumulates into 16-bit output, improving throughput.
result Offers 2x improvement in throughput compared to existing symmetric-operand-size instructions, with negligible overflow.
Cryptotree enables accurate predictions on encrypted data using Random Forests.
problem Applying machine learning to private data while preserving confidentiality.
method Adapting Neural RF to CKKS scheme for HE operations on encrypted data.
result Cryptotree achieves better prediction results on encrypted data than regular RF.
We introduce an online tensor decomposition based approach for two latent variable modeling problems namely, (1) community detection, in which we learn the latent communities that the social actors in social networks belong to, and (2) topic modeling, in which we infer hidden topics of text articles. We consider decomp…
SparseTrain uses dynamic sparsity in training deep neural networks on CPUs.
problem Training deep neural networks efficiently on general-purpose processors.
method Exploits dynamic zeros introduced by ReLU in feature maps and gradients.
result Significantly speeds up training on CPUs, up to 1.51x.
Gibbs sampling is a workhorse for Bayesian inference but has several limitations when used for parameter estimation, and is often much slower than non-sampling inference methods. SAME (State Augmentation for Marginal Estimation) \cite{Doucet99,Doucet02} is an approach to MAP parameter estimation which gives improved pa…
Q-NETs use neural networks to estimate integrals of low-dimensional functions efficiently.
problem Estimating integrals of multidimensional functions with costly evaluations.
method Fixed neural networks (Q-NETs) that operate on proxy function parameters to calculate exact integrals over subsets of dimensions.
result Q-NETs can calculate integrals over any subset of dimensions without resampling or retraining the proxy.
We introduce topological parallelisms of oriented lines (briefly called oriented parallelisms). Every topological parallelism (of lines) on PG(3,R) gives rise to a parallelism of oriented lines, but we show that even the most homogeneous parallelisms of oriented lines other than the Clifford parallelism do not necessar…
Parallelizes MCTS for continuous domains using leaf and root parallelization.
problem Solving challenging tasks in continuous domains using MCTS.
method Extends existing parallelization strategies to continuous domains, focusing on leaf and root parallelization.
result Proposes two final selection strategies for continuous states in root parallelization.
We present an overview of techniques for quantizing convolutional neural networks for inference with integer weights and activations. Per-channel quantization of weights and per-layer quantization of activations to 8-bits of precision post-training produces classification accuracies within 2% of floating point networks…
The paper parallelizes HMM inference for efficient long-term computations.
problem Efficiently computing inference in long-term hidden Markov models.
method Parallelization using associative elements and operators for sum-product and max-product algorithms.
result The proposed parallel algorithms are computationally efficient for long time horizons.
New rational parallelisms found on complex manifolds that are not flat.
problem Finding non-flat rational parallelisms on complex manifolds.
method Examined rational parallelisms on compact complex manifolds, discovering non-flat examples.
result Discovered rational parallelisms on compact complex manifolds that are not flat.
Introduces a natural parallel translation for navigation data.
problem Navigation data geometric representation and parallelism.
method Introduces a natural parallel translation using Riemannian parallelism.
result The natural parallel translation preserves the Randers norm and has a finite-dimensional holonomy group.
This study compares parallel SMC and MCMC for Bayesian deep learning, showing SMC parallel is faster.
problem Efficiently performing Bayesian deep learning with parallel computing.
method Compared sequential Monte Carlo (SMC) and Markov chain Monte Carlo (MCMC) in parallel settings.
result Parallel SMC achieves similar convergence as a single SMC but with reduced communication time.
We prove a conjecture formulated by Pablo M. Chacon and Guillermo A. Lobos in [Pseudo-parallel Lagrangian submanifolds in complex space forms, Differential Geom. Appl.] stating that every Lagrangian pseudo-parallel submanifold of a complex space form of dimension at least 3 is semi-parallel.
The paper explores parallel 1-forms on special Finsler manifolds and their properties.
problem Investigating parallel 1-forms on specific Finsler manifolds.
method Analyzing Landsberg manifolds, metrizability freedom, and specific Finsler metrics.
result Landsberg surfaces with parallel 1-forms are necessarily Berwaldian, and the metrizability freedom is at least 2.
This paper surveys parallel submanifolds in Riemannian and pseudo-Riemannian manifolds.
problem Understanding parallel submanifolds in Riemannian and pseudo-Riemannian manifolds.
method Comprehensive survey of parallel submanifolds.
result Extrinsic invariants of parallel submanifolds do not vary from point to point.
We propose a new integrated method of exploiting model, batch and domain parallelism for the training of deep neural networks (DNNs) on large distributed-memory computers using minibatch stochastic gradient descent (SGD). Our goal is to find an efficient parallelization strategy for a fixed batch size using P process…
Characterizes regular parallelisms in 3D space with 2-torus action.
problem Characterizing regular parallelisms in 3D space with 2-torus action.
method Characterization using compactness, equivalence relations, and properties of complex vector spaces.
result There is a 1-dimensional subtorus fixing every parallel class, leading to 2- or 3-dimensional regular parallelisms.
Paper studies second order symmetric parallel tensors in generalized f.pk-space forms.
problem Exploring properties of second order symmetric parallel tensors in generalized f.pk-space forms.
method Analyzes the properties of second order symmetric parallel tensors and deduces the existence or non-existence of certain tensors and hypersurfaces.
result There does not exist second order skew-symmetric parallel tensor in f.pk-space form. There is no parallel hypersurface in a generalized f.pk-space form but there is semi-parallel hypersurface.
Classifies simply-connected pluriclosed manifolds with parallel Bismut torsion.
problem Classifying specific types of manifolds with parallel Bismut torsion.
method Complete classification through mathematical analysis.
result Established a splitting theorem for certain manifolds.
Study on generalized ξ-parallel maps in Riemannian geometry.
problem Characterizing and understanding generalized ξ-parallel maps.
method Defined energy functional, derived first variation formula, and Euler-Lagrange equation.
result Established fundamental properties and relationships with harmonic and biharmonic maps.
This paper improves parallel belief propagation for scalable machine learning.
problem Efficient parallelization of belief propagation for large-scale machine learning tasks.
method Use of scalable relaxed schedulers to parallelize belief propagation.
result Our approach outperforms previous methods in scalability and convergence time.
Cyclic Data Parallelism reduces memory usage and balances gradient communications.
problem Training large deep learning models requires efficient parallelism to scale.
method Cyclic Data Parallelism shifts micro-batches from simultaneous to sequential execution, balancing memory and gradient communications.
result Cyclic Data Parallelism reduces total memory usage and balances gradient communications.
The paper examines parallel one forms on Riemannian and Finslerian manifolds.
problem Existence of parallel one forms on Riemannian and Finslerian manifolds.
method Using Finslerian settings, the paper investigates the existence of parallel one forms on Riemannian manifolds and Finslerian manifolds, proving conditions for their existence and non-existence.
result Conditions for the existence and non-existence of parallel one forms on Riemannian and Finslerian manifolds.
Betten and Riesinger constructed Parallelisms of PG(3,R) with automorphism group SO(3,R) by applying the reducible SO(3,R)-action to a rotational Betten spread. This was generalized by the present author so as to include oriented parallelisms (i.e., p…
The paper studies parallel spinor flows on 3D Cauchy hypersurfaces and provides initial data characterizations.
problem Characterizing parallel spinors on Ricci flat Lorentzian four-manifolds.
method Evolution flow defined by parallel spinors, proving preservation of constraints, solving left-invariant flows.
result Initial data characterization of parallel spinors on Ricci flat Lorentzian four-manifolds.
A submanifold of a pseudo-Riemannian manifold is said to have parallel mean curvature vector if the mean curvature vector field H is parallel as a section of the normal bundle. Submanifolds with parallel mean curvature vector are important since they are critical points of some natural functionals. In this paper, we su…
A new algorithm for parallel transport on shape spaces is presented and compared to existing methods.
problem Statistical analysis of shape data, especially in time series and optimization.
method Pole ladder algorithm for parallel transport on Kendall shape spaces, compared to integration methods.
result The pole ladder algorithm is a more efficient method for parallel transport.
Investigates parallel spinors on Lorentzian four-manifolds using differential geometry.
problem Characterizing and classifying Lorentzian four-manifolds with parallel spinors.
method Formulated parallel spinor flow equations and used parabolic pairs theory.
result Characterized all parallel Cauchy pairs on simply connected Cauchy surfaces and classified compact three-manifolds.
Compact 3D Cotton-parallel manifolds are always conformally flat.
problem Understanding the properties of compact 3D Cotton-parallel manifolds.
method Analyzing the Cotton tensor and its parallelism condition.
result Compact 3D Cotton-parallel manifolds are conformally flat.
We characterize compact locally conformal parallel G2 (respectively, Spin(7)) manifolds as fiber bundles over S1 with compact nearly Kähler (respectively, compact nearly parallel G2) fiber. A more specific characterization is provided when the local parallel structures are flat.
Linear algebra approach for parallel deep learning models.
problem Training large DNNs in distributed environments.
method Linear algebraic approach to model parallelism.
result Manual development of backward operators for gradient-based training.
Study on balanced Hermitian threefolds with parallel Bismut torsion.
problem Characterizing compact, balanced BTP threefolds.
method Detailed description of all compact, balanced BTP threefolds.
result Characterization of all compact, balanced BTP threefolds.
Study on pseudo-Riemannian metrics on Lie groups, finding new non-Einstein examples.
problem Characterizing and finding non-Einstein pseudo-Riemannian metrics on Lie groups.
method Analyzing left invariant metrics, using double extension process, and constructing examples.
result Construction of infinitely many new explicit examples of non-Einstein pseudo-Riemannian metrics on Lie groups.
This work optimizes deep learning training by combining data and model parallelism.
problem Training large models with multiple GPUs suffers from high communication overhead and statistical efficiency loss.
method Hybrid parallelization combining data and model parallelism.
result Hybrid training provides significant speedup compared to data parallelism alone.
In [dLMu05], DeLellis and Müller proved a quantitative version of Codazzi's theorem, namely for a smooth embedded surface Σ⊆R3 with area normalized to H2(Σ)=4π, it was shown that ∥AΣ−id∥L2(Σ)≤C∥AΣ0∥L2(Σ) , and building on…
A geometry with parallel skew-symmetric torsion is a Riemannian manifold carrying a metric connection with parallel skew-symmetric torsion. Besides the trivial case of the Levi-Civita connection, geometries with non-vanishing parallel skew-symmetric torsion arise naturally in several geometric contexts, e.g. on natural…
Method predicts how probability distributions evolve over time.
problem Predicting how systems described by probability distributions evolve under different conditions.
method Wasserstein Parallel Transport
result Wasserstein Parallel Transport provides counterfactual comparisons of distributional dynamics.
The study classifies parallel mean curvature spheres in a sphere-hyperbolic product space.
problem Understanding surfaces with parallel mean curvature in a specific Riemannian product space.
method Analyzing the holomorphic quadratic differential and topological constraints.
result Classification of all parallel mean curvature spheres with vanishing differential.
NeLLoC improves image compression with parallel decoding.
problem Image compression with OOD generalization.
method Local autoregressive model with parallel decoding.
result Significant gains in compression runtime.
Classifies Calabi hypersurfaces with parallel Fubini-Pick form.
problem Classifying Calabi hypersurfaces with specific geometric properties.
method Classification based on parallel Fubini-Pick form and Levi-Civita connection.
result Classification of 2 and 3-dimensional Calabi hypersurfaces.
New connections found with specific torsion properties.
problem Understanding metric connections with specific torsion properties.
method Described Lorentzian manifolds with metric connections having parallel, skew-symmetric torsion.
result Found new Lorentzian manifolds with metric connections having parallel, skew-symmetric torsion.
Study pp-waves with lightlike parallel spinors in vacuum spacetimes.
problem Characterize pp-waves with lightlike parallel spinors in vacuum spacetimes.
method Parametrize pp-wave spacetimes, show correspondence with Riemannian metrics, prove parallel spinor condition.
result A pp-wave spacetime with a lightlike parallel spinor corresponds to a Ricci-flat metric with a parallel spinor.
Paper proves non-existence of certain hypersurfaces in complex quadric.
problem Non-existence of Hopf real hypersurfaces with parallel normal Jacobi operator.
method Introducing C-parallel and Reeb parallel normal Jacobi operators, proving non-existence theorems. result Non-existence of Hopf real hypersurfaces with C-parallel normal Jacobi operator.