New method reduces variance in complex probabilistic model optimization.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Doubly SGD improves convergence for intractable objective optimization.
Doubly-stochastic normalization improves robustness to heteroskedastic noise.
We propose an efficient method for estimating covariate effects in doubly-stochastic spatial models.
This technical report proves components consistency for the Doubly Stochastic Dirichlet Process with exponential convergence of posterior probability. We also present the fundamental properties for DSDP as well as inference algorithms. Simulation toy experiment and real-world experiment results for single and multi-clu…
Paper proposes Sinkformers for Transformers with doubly stochastic attention.
We propose a doubly stochastic primal-dual coordinate optimization algorithm for empirical risk minimization, which can be formulated as a bilinear saddle-point problem. In each iteration, our method randomly samples a block of coordinates of the primal and dual solutions to update. The linear convergence of our method…
Proposes a new simulator for complex arrival processes.
In this letter, we introduce a distributed Nesterov method, termed as , that does not require doubly-stochastic weight matrices. Instead, the implementation is based on a simultaneous application of both row- and column-stochastic weights that makes this method applicable to arbitrary (strongly-connected…
An interesting theme in complex differential geometry is to find a correspondence between algebraic objects and differential geometric objects. One of the most attractive is the non-abelian Hodge theory of Simpson. In this paper, pursuing an analogue of the non-abelian Hodge theory in the context of -difference modu…
FDSKL algorithm trains vertically partitioned data with kernels securely and efficiently.
Action-BED: Task-Driven Bayesian Experimental Design
In this paper, we develop a new accelerated stochastic gradient method for efficiently solving the convex regularized empirical risk minimization problem in mini-batch settings. The use of mini-batches is becoming a golden standard in the machine learning community, because mini-batch settings stabilize the gradient es…
Geometric approach for unsupervised word embedding alignment.
We study the problem of off-policy critic evaluation in several variants of value-based off-policy actor-critic algorithms. Off-policy actor-critic algorithms require an off-policy critic evaluation step, to estimate the value of the new policy after every policy gradient update. Despite enormous success of off-policy …
This paper discusses properties of a Doubly Stochastic Poisson Process (DSPP) where the intensity process belongs to a class of affine diffusions. For any intensity process from this class we derive an analytical expression for probability distribution functions of the corresponding DSPP. A specification of our results…
The stochastic gradient descent has been widely used for solving composite optimization problems in big data analyses. Many algorithms and convergence properties have been developed. The composite functions were convex primarily and gradually nonconvex composite functions have been adopted to obtain more desirable prop…
The paper improves boundary detection and density estimation on noisy data.
When training a machine learning model with observational data, it is often encountered that some values are systemically missing. Learning from the incomplete data in which the missingness depends on some covariates may lead to biased estimation of parameters and even harm the fairness of decision outcome. This paper …
Doubly stochastic learning algorithms are scalable kernel methods that perform very well in practice. However, their generalization properties are not well understood and their analysis is challenging since the corresponding learning sequence may not be in the hypothesis space induced by the kernel. In this paper, we p…
We study a doubly reflected backward stochastic differential equation (BSDE) with integrable parameters and the related Dynkin game. When the lower obstacle and the upper obstacle of the equation are completely separated, we construct a unique solution of the doubly reflected BSDE by pasting local solutions and…
Robustly infers manifold density and geometry under high-dimensional noise.
Doubly robust self-training improves semi-supervised learning by balancing labeled and pseudo-labeled data.
DSVNP uses global and local latent variables for improved neural process predictions.
It is of increasing importance to develop learning methods for ranking. In contrast to many learning objectives, however, the ranking problem presents difficulties due to the fact that the space of permutations is not smooth. In this paper, we examine the class of rank-linear objective functions, which includes popular…
The general perception is that kernel methods are not scalable, and neural nets are the methods of choice for nonlinear learning problems. Or have we simply not tried hard enough for kernel methods? Here we propose an approach that scales up kernel methods using a novel concept called "doubly stochastic functional grad…
Semi-supervised ordinal regression (SOR) problems are ubiquitous in real-world applications, where only a few ordered instances are labeled and massive instances remain unlabeled. Recent researches have shown that directly optimizing concordance index or AUC can impose a better ranking on the data than optimizing t…
We introduce a doubly stochastic proximal gradient algorithm for optimizing a finite average of smooth convex functions, whose gradients depend on numerically expensive expectations. Our main motivation is the acceleration of the optimization of the regularized Cox partial-likelihood (the core model used in survival an…
Graph alignment problem solved with convex relaxations for correlated matrices.
We consider learning problems over training sets in which both, the number of training examples and the dimension of the feature vectors, are large. To solve these problems we propose the random parallel stochastic algorithm (RAPSA). We call the algorithm random parallel because it utilizes multiple parallel processors…
Gaussian processes (GPs) are a good choice for function approximation as they are flexible, robust to over-fitting, and provide well-calibrated predictive uncertainty. Deep Gaussian processes (DGPs) are multi-layer generalisations of GPs, but inference in these models has proved challenging. Existing approaches to infe…
ADSGD method speeds up model identification in sparse optimization.
π-GNN learns soft permutations for graph representations, improving graph classification and regression.
New algorithms learn graph structures privately, matching best results.
This work concerns the study of certain finite-energy solutions of the anti-self-dual Yang-Mills equations on Euclidean 4-dimensional space which are periodic in two directions, so-called doubly-periodic instantons. We establish a circle of ideas involving equivalent analytical and algebraic-geometric descriptions of t…
The paper tackles robust policy learning from multiple data sources.
S2M optimizes mining for diverse data subpopulations.
DoubleGen addresses bias in generative modeling of counterfactuals.
Estimates outcomes under hypothetical scenarios using a flexible framework.
Accelerates Birkhoff projection for manifold-constrained hyper-connections with high accuracy and speed.
Contextual multi-armed bandit algorithms are widely used in sequential decision tasks such as news article recommendation systems, web page ad placement algorithms, and mobile health. Most of the existing algorithms have regret proportional to a polynomial function of the context dimension, . In many applications ho…
The paper examines Einstein doubly warped product manifolds with a semi-symmetric metric connection.
Characterizes spacetimes using doubly torqued vectors.
New invariant measures doubly slice links, disproving previous bounds.
Enhances DGPs with adaptive RKHS Fourier features for better non-stationary pattern modeling.
Characterizes a specific type of spacetime using vector fields.
Characterizes and examines gradient solitons on doubly warped product manifolds.
Paper extends causal inference to non-Euclidean data like images and distributions.