Paper proposes Sinkformers for Transformers with doubly stochastic attention.
problem Improving Transformer models' accuracy in vision and natural language processing.
method Using Sinkhorn's algorithm to make attention matrices doubly stochastic instead of SoftMax normalization.
result Sinkformers enhance model accuracy in vision and natural language processing tasks.
Graph alignment problem solved with convex relaxations for correlated matrices.
problem Recovering hidden vertex permutations from correlated Gaussian matrices.
method Convex relaxations of the quadratic assignment problem over doubly stochastic matrices.
result The solution of the convex relaxation concentrates around the ground-truth permutation matrix for certain correlation parameters.
Geometric approach for unsupervised word embedding alignment.
problem Learning alignment between word embeddings of source and target languages.
method Formulates alignment as domain adaptation on the manifold of doubly stochastic matrices, employing Riemannian conjugate gradient algorithm.
result Empirically outperforms state-of-the-art methods on bilingual lexicon induction tasks.
We introduce a new method to handle permutations efficiently using variational inference.
problem Efficient probabilistic reasoning about permutations in high-dimensional spaces.
method We reparameterize the Birkhoff polytope to enable variational inference over permutations.
result Our method enables efficient and accurate Bayesian inference over permutations.
π-GNN learns soft permutations for graph representations, improving graph classification and regression.
problem Limitations of MPNNs in graph neural networks.
method Proposes π-GNN, which learns a soft permutation matrix for each graph, projecting graphs into a common vector space.
result π-GNN achieves performance competitive with state-of-the-art models on graph classification and regression tasks.
A new algorithm speeds up optimal transport for machine learning.
problem Optimal transport for machine learning with additional terms.
method Forward-backward splitting algorithm based on Bregman distances.
result Significant improvement in speed and performance for domain adaptation.
New algorithms learn graph structures privately, matching best results.
problem Private learning of graph structures with multiple blocks.
method Sum-of-squares relaxation and exponential mechanism for score function.
result Matches statistical utility of previous best non-private methods.
Accelerates Birkhoff projection for manifold-constrained hyper-connections with high accuracy and speed.
problem Inaccurate and slow Birkhoff projection in mHC implementations.
method Dual formulation, Newton's method, implicit differentiation, warp-level CUDA kernel.
result Substantial speedups and accuracy improvements in doubly stochastic projections.
Doubly-stochastic normalization improves robustness to heteroskedastic noise.
problem Robustness to heteroskedastic noise in affinity matrix construction.
method Doubly-stochastic normalization of the Gaussian kernel.
result Doubly-stochastic normalization converges to clean matrix with rate m−1/2 under heteroskedastic noise. This technical report proves components consistency for the Doubly Stochastic Dirichlet Process with exponential convergence of posterior probability. We also present the fundamental properties for DSDP as well as inference algorithms. Simulation toy experiment and real-world experiment results for single and multi-clu…
We propose an efficient method for estimating covariate effects in doubly-stochastic spatial models.
problem Computational demands and restrictive assumptions in existing doubly-stochastic spatial models.
method Penalized regression method for estimating covariate effects in doubly-stochastic point processes.
result Consistency and asymptotic normality of the covariate effect estimates achieved despite model misspecification.
Deep kernel processes unify various models using Gram matrices and kernel functions.
problem Unified representation of various deep learning models.
method Defining deep kernel processes with progressively transformed Gram matrices and sampling from inverse Wishart distributions.
result Deep Gaussian processes, BNNs, infinite BNNs, and infinite BNNs with bottlenecks can all be written as deep kernel processes.
A new method for deep Wishart processes improves kernel-based models.
problem Inference in deep Wishart processes is challenging due to the need for flexible distributions over positive semi-definite matrices.
method Developed a novel approach to flexible distributions over positive semi-definite matrices using the Bartlett decomposition of the Wishart probability density. Used this to create an approximate posterior for the DWP.
result Improved performance of inference in the DWP compared to DGP with equivalent prior.
It is of increasing importance to develop learning methods for ranking. In contrast to many learning objectives, however, the ranking problem presents difficulties due to the fact that the space of permutations is not smooth. In this paper, we examine the class of rank-linear objective functions, which includes popular…
The paper analyzes generalization properties of scalable kernel methods.
problem Understanding the generalization of doubly stochastic learning algorithms.
method Theoretical analysis of different variants of doubly stochastic learning algorithms in nonparametric regression.
result Derivation of generalization error convergence results for the algorithms.
FDSKL algorithm trains vertically partitioned data with kernels securely and efficiently.
problem Training vertically partitioned data with kernels while maintaining privacy.
method FDSKL algorithm using random features and doubly stochastic gradients for federated learning.
result FDSKL achieves sublinear convergence and guarantees data security.
This paper discusses properties of a Doubly Stochastic Poisson Process (DSPP) where the intensity process belongs to a class of affine diffusions. For any intensity process from this class we derive an analytical expression for probability distribution functions of the corresponding DSPP. A specification of our results…
The paper improves boundary detection and density estimation on noisy data.
problem Detecting boundary points and estimating density on noisy data from compact manifolds.
method Doubly stochastic scaling of the Gaussian heat kernel via Sinkhorn iterations.
result The new estimates of boundary points and density outperform standard methods, especially under noise.
New method reduces variance in complex probabilistic model optimization.
problem High variance in stochastic optimisation of complex models.
method Use recognition network to approximate optimal control variate for each mini-batch.
result Sub-optimal variance reduction is improved with new approach.
A distributed Nesterov method for arbitrary graphs achieves faster convergence.
problem Optimizing distributed systems over arbitrary graphs.
method Distributed Nesterov method ABN and its variation FROZEN. result Achieves acceleration compared to state-of-the-art methods.
Robustly infers manifold density and geometry under high-dimensional noise.
problem Inaccurate kernel density estimation under high-dimensional noise.
method Doubly stochastic normalization of Gaussian kernel.
result Robust tools for density estimation, noise magnitude estimation, and distance approximation.
Doubly SGD improves convergence for intractable objective optimization.
problem Optimizing objectives in sum of intractable expectations.
method Doubly SGD with doubly stochastic gradients and independent minibatching.
result Established convergence of doubly SGD under general conditions, including dependent component gradient estimators.
Gaussian processes (GPs) are a good choice for function approximation as they are flexible, robust to over-fitting, and provide well-calibrated predictive uncertainty. Deep Gaussian processes (DGPs) are multi-layer generalisations of GPs, but inference in these models has proved challenging. Existing approaches to infe…
The general perception is that kernel methods are not scalable, and neural nets are the methods of choice for nonlinear learning problems. Or have we simply not tried hard enough for kernel methods? Here we propose an approach that scales up kernel methods using a novel concept called "doubly stochastic functional grad…
ADSGD method speeds up model identification in sparse optimization.
problem Implicit model identification in sparse optimization problems.
method Accelerated Doubly Stochastic Gradient Method (ADSGD) for faster explicit model identification.
result ADSGD achieves faster explicit model identification and improved algorithm efficiency.
We propose a doubly stochastic primal-dual coordinate optimization algorithm for empirical risk minimization, which can be formulated as a bilinear saddle-point problem. In each iteration, our method randomly samples a block of coordinates of the primal and dual solutions to update. The linear convergence of our method…
We introduce a doubly stochastic proximal gradient algorithm for optimizing a finite average of smooth convex functions, whose gradients depend on numerically expensive expectations. Our main motivation is the acceleration of the optimization of the regularized Cox partial-likelihood (the core model used in survival an…
DSVNP uses global and local latent variables for improved neural process predictions.
problem Limited expressiveness of vanilla neural processes in capturing target-specific local variation.
method Introduces DSVNP combining global and local latent variables for prediction.
result Competitive prediction performance in multi-output regression and uncertainty estimation.
Bayesian Neural Networks built block-by-block with uncertainty estimates.
problem Building interpretable and uncertainty-aware neural networks.
method Bayesian Neural Networks (BNNs) constructed using blocks, with doubly stochastic variational inference for posterior approximation.
result Uncertainty estimates provided for Bayesian Neural Networks.
Paper characterizes nc-rank using gradient flow on symmetric space.
problem Characterizing the noncommutative rank of matrices.
method Interprets residuals as gradients of a convex function on symmetric space and uses unbounded gradient flow.
result Noncommutative corank equals half the minimum gradient-norm of residuals.
The Sinkhorn-Knopp algorithm converges quickly but the number of iterations is poorly understood.
problem Understanding the number of iterations required for the Sinkhorn-Knopp algorithm to converge.
method Analyzing the Sinkhorn-Knopp algorithm for matrices with a specific density threshold.
result The Sinkhorn-Knopp algorithm requires Ω(n1/2/ε) iterations for matrices with density γ<1/2. Enhances DGPs with adaptive RKHS Fourier features for better non-stationary pattern modeling.
problem Capturing complex non-stationary patterns in non-linear dynamical systems.
method Integrates ODE-based RKHS Fourier features into DGPs using convolution operations for adaptive amplitude and phase modulation. Uses a doubly stochastic variational inference framework.
result Improved predictive performance across various regression tasks.
S2M optimizes mining for diverse data subpopulations.
problem Scalability and uniformity in training sets with many labels and diverse data.
method Doubly-stochastic mining (S2M) computes per-example and minibatch losses on hardest labels/examples.
result S2M ensures good performance across all data subpopulations.
We model messaging activities as a hierarchical doubly stochastic point process with three main levels, and develop an iterative algorithm for inferring actors' relative latent positions from a stream of messaging activity data. Each of the message-exchanging actors is modeled as a process in a latent space. The actors…
A scalable algorithm improves AUC optimization for semi-supervised ordinal regression.
problem Optimizing AUC for semi-supervised ordinal regression with limited labeled data.
method Proposes QS3ORAO using quadruply stochastic gradients for scalable kernelized learning. result Converges to optimal solution at O(1/t) rate, demonstrating efficiency and effectiveness. Many machine learning applications are based on data collected from people, such as their tastes and behaviour as well as biological traits and genetic data. Regardless of how important the application might be, one has to make sure individuals' identities or the privacy of the data are not compromised in the analysis.…
Proposes a new simulator for complex arrival processes.
problem Modeling and simulating complex arrival processes with non-stationary and multi-dimensional rates.
method Integrates Monte Carlo and GANs to model a broad class of arrival processes.
result Consistent and efficient estimation of the simulator using Wasserstein distance.
Under the Basel II standards, the Operational Risk (OpRisk) advanced measurement approach is not prescriptive regarding the class of statistical model utilised to undertake capital estimation. It has however become well accepted to utlise a Loss Distributional Approach (LDA) paradigm to model the individual OpRisk loss…
We consider the problem of identifying current coupons for Agency backed To-be-Announced (TBA) Mortgage Backed Securities. In a doubly stochastic factor based model which allows for prepayment intensities to depend upon current and origination mortgage rates, as well as underlying investment factors, we identify the cu…
We obtain an explicit formula for the bilateral counterparty valuation adjustment of a credit default swaps portfolio referencing an asymptotically large number of entities. We perform the analysis under a doubly stochastic intensity framework, allowing for default correlation through a common jump process. The key ins…
A recurring problem when building probabilistic latent variable models is regularization and model selection, for instance, the choice of the dimensionality of the latent space. In the context of belief networks with latent variables, this problem has been adressed with Automatic Relevance Determination (ARD) employing…
We propose a unified framework for equity and credit risk modeling, where the default time is a doubly stochastic random time with intensity driven by an underlying affine factor process. This approach allows for flexible interactions between the defaultable stock price, its stochastic volatility and the default intens…
Bayesian approach to data association using Gaussian processes.
problem Separating data from different generating processes.
method Fully Bayesian approach with Gaussian process priors for structure encoding and doubly stochastic variational inference.
result Efficient learning scheme for deep Gaussian process priors.
DGPs improve air quality inference from sparse data.
problem Accurate air quality monitoring in unmonitored areas.
method Deep Gaussian Processes with Doubly Stochastic Variational Inference.
result DGPs outperform state-of-the-art models in AQ inference.
We introduce local expectation gradients which is a general purpose stochastic variational inference algorithm for constructing stochastic gradients through sampling from the variational distribution. This algorithm divides the problem of estimating the stochastic gradients over multiple variational parameters into sma…
Proves Seidel's conjectures about ideal tetrahedra in hyperbolic 3-space.
problem Determining the volume of ideal hyperbolic tetrahedra using algebraic maps.
method Analyzes the doubly stochastic Gram matrix of tetrahedron vertices.
result Volume is a monotonic function of the permanent and determinant of the Gram matrix.
The softmax representation of probabilities for categorical variables plays a prominent role in modern machine learning with numerous applications in areas such as large scale classification, neural language modeling and recommendation systems. However, softmax estimation is very expensive for large scale inference bec…
Method uses deep learning to estimate traffic intensity.
problem Estimating stochastic intensity of traffic processes.
method Deep neural networks for nonlinear filtering.
result Deep learning method accurately estimates traffic intensity.