Doubly-stochastic normalization improves robustness to heteroskedastic noise.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper proposes Sinkformers for Transformers with doubly stochastic attention.
We propose an efficient method for estimating covariate effects in doubly-stochastic spatial models.
Robustly infers manifold density and geometry under high-dimensional noise.
This technical report proves components consistency for the Doubly Stochastic Dirichlet Process with exponential convergence of posterior probability. We also present the fundamental properties for DSDP as well as inference algorithms. Simulation toy experiment and real-world experiment results for single and multi-clu…
FDSKL algorithm trains vertically partitioned data with kernels securely and efficiently.
Geometric approach for unsupervised word embedding alignment.
This paper discusses properties of a Doubly Stochastic Poisson Process (DSPP) where the intensity process belongs to a class of affine diffusions. For any intensity process from this class we derive an analytical expression for probability distribution functions of the corresponding DSPP. A specification of our results…
The paper improves boundary detection and density estimation on noisy data.
New method reduces variance in complex probabilistic model optimization.
Doubly stochastic learning algorithms are scalable kernel methods that perform very well in practice. However, their generalization properties are not well understood and their analysis is challenging since the corresponding learning sequence may not be in the hypothesis space induced by the kernel. In this paper, we p…
The softmax representation of probabilities for categorical variables plays a prominent role in modern machine learning with numerous applications in areas such as large scale classification, neural language modeling and recommendation systems. However, softmax estimation is very expensive for large scale inference bec…
Doubly SGD improves convergence for intractable objective optimization.
Graph alignment problem solved with convex relaxations for correlated matrices.
Gaussian processes (GPs) are a good choice for function approximation as they are flexible, robust to over-fitting, and provide well-calibrated predictive uncertainty. Deep Gaussian processes (DGPs) are multi-layer generalisations of GPs, but inference in these models has proved challenging. Existing approaches to infe…
The general perception is that kernel methods are not scalable, and neural nets are the methods of choice for nonlinear learning problems. Or have we simply not tried hard enough for kernel methods? Here we propose an approach that scales up kernel methods using a novel concept called "doubly stochastic functional grad…
ADSGD method speeds up model identification in sparse optimization.
We propose a doubly stochastic primal-dual coordinate optimization algorithm for empirical risk minimization, which can be formulated as a bilinear saddle-point problem. In each iteration, our method randomly samples a block of coordinates of the primal and dual solutions to update. The linear convergence of our method…
We introduce a doubly stochastic proximal gradient algorithm for optimizing a finite average of smooth convex functions, whose gradients depend on numerically expensive expectations. Our main motivation is the acceleration of the optimization of the regularized Cox partial-likelihood (the core model used in survival an…
DSVNP uses global and local latent variables for improved neural process predictions.
π-GNN learns soft permutations for graph representations, improving graph classification and regression.
Accelerates Birkhoff projection for manifold-constrained hyper-connections with high accuracy and speed.
It is of increasing importance to develop learning methods for ranking. In contrast to many learning objectives, however, the ranking problem presents difficulties due to the fact that the space of permutations is not smooth. In this paper, we examine the class of rank-linear objective functions, which includes popular…
New algorithms learn graph structures privately, matching best results.
Enhances DGPs with adaptive RKHS Fourier features for better non-stationary pattern modeling.
S2M optimizes mining for diverse data subpopulations.
We model messaging activities as a hierarchical doubly stochastic point process with three main levels, and develop an iterative algorithm for inferring actors' relative latent positions from a stream of messaging activity data. Each of the message-exchanging actors is modeled as a process in a latent space. The actors…
Edge features contain important information about graphs. However, current state-of-the-art neural network models designed for graph learning, e.g. graph convolutional networks (GCN) and graph attention networks (GAT), adequately utilize edge features, especially multi-dimensional edge features. In this paper, we build…
Optimal transport aims to estimate a transportation plan that minimizes a displacement cost. This is realized by optimizing the scalar product between the sought plan and the given cost, over the space of doubly stochastic matrices. When the entropy regularization is added to the problem, the transportation plan can be…
Many machine learning applications are based on data collected from people, such as their tastes and behaviour as well as biological traits and genetic data. Regardless of how important the application might be, one has to make sure individuals' identities or the privacy of the data are not compromised in the analysis.…
Proposes a new simulator for complex arrival processes.
The Sinkhorn-Knopp algorithm converges quickly but the number of iterations is poorly understood.
The data association problem is concerned with separating data coming from different generating processes, for example when data come from different data sources, contain significant noise, or exhibit multimodality. We present a fully Bayesian approach to this problem. Our model is capable of simultaneously solving the…
Under the Basel II standards, the Operational Risk (OpRisk) advanced measurement approach is not prescriptive regarding the class of statistical model utilised to undertake capital estimation. It has however become well accepted to utlise a Loss Distributional Approach (LDA) paradigm to model the individual OpRisk loss…
We consider the problem of identifying current coupons for Agency backed To-be-Announced (TBA) Mortgage Backed Securities. In a doubly stochastic factor based model which allows for prepayment intensities to depend upon current and origination mortgage rates, as well as underlying investment factors, we identify the cu…
We obtain an explicit formula for the bilateral counterparty valuation adjustment of a credit default swaps portfolio referencing an asymptotically large number of entities. We perform the analysis under a doubly stochastic intensity framework, allowing for default correlation through a common jump process. The key ins…
A recurring problem when building probabilistic latent variable models is regularization and model selection, for instance, the choice of the dimensionality of the latent space. In the context of belief networks with latent variables, this problem has been adressed with Automatic Relevance Determination (ARD) employing…
We prove, in the case of hyperbolic 3-space, a couple of conjectures raised by J. J. Seidel in "On the volume of a hyperbolic simplex", Stud. Sci. Math. Hung. 21, 243-249, 1986. These conjectures concern expressing the volume of an ideal hyperbolic tetrahedron as a monotonic function of algebraic maps. More precisely, …
We propose a unified framework for equity and credit risk modeling, where the default time is a doubly stochastic random time with intensity driven by an underlying affine factor process. This approach allows for flexible interactions between the defaultable stock price, its stochastic volatility and the default intens…
DGPs improve air quality inference from sparse data.
Word embedding, which encodes words into vectors, is an important starting point in natural language processing and commonly used in many text-based machine learning tasks. However, in most current word embedding approaches, the similarity in embedding space is not optimized in the learning. In this paper we propose a …
In this letter, we introduce a distributed Nesterov method, termed as , that does not require doubly-stochastic weight matrices. Instead, the implementation is based on a simultaneous application of both row- and column-stochastic weights that makes this method applicable to arbitrary (strongly-connected…
In this paper we provide a valuation formula for different classes of actuarial and financial contracts which depend on a general loss process, by using the Malliavin calculus. In analogy with the celebrated Black-Scholes formula, we aim at expressing the expected cash flow in terms of a building block. The former is r…
We provide simple schemes to build Bayesian Neural Networks (BNNs), block by block, inspired by a recent idea of computation skeletons. We show how by adjusting the types of blocks that are used within the computation skeleton, we can identify interesting relationships with Deep Gaussian Processes (DGPs), deep kernel l…
We introduce local expectation gradients which is a general purpose stochastic variational inference algorithm for constructing stochastic gradients through sampling from the variational distribution. This algorithm divides the problem of estimating the stochastic gradients over multiple variational parameters into sma…
Method uses deep learning to estimate traffic intensity.
State-space models (SSMs) are a highly expressive model class for learning patterns in time series data and for system identification. Deterministic versions of SSMs (e.g. LSTMs) proved extremely successful in modeling complex time series data. Fully probabilistic SSMs, however, are often found hard to train, even for …
We introduce a novel stochastic version of the non-reversible, rejection-free Bouncy Particle Sampler (BPS), a Markov process whose sample trajectories are piecewise linear. The algorithm is based on simulating first arrival times in a doubly stochastic Poisson process using the thinning method, and allows efficient sa…