Doubly-stochastic normalization improves robustness to heteroskedastic noise.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper proposes Sinkformers for Transformers with doubly stochastic attention.
We propose an efficient method for estimating covariate effects in doubly-stochastic spatial models.
Robustly infers manifold density and geometry under high-dimensional noise.
Estimates outcomes under hypothetical scenarios using a flexible framework.
Paper defines new risk measures for elliptical distributions.
Corrects mismatch in consistency of nuisance estimators for doubly robust methods.
This technical report proves components consistency for the Doubly Stochastic Dirichlet Process with exponential convergence of posterior probability. We also present the fundamental properties for DSDP as well as inference algorithms. Simulation toy experiment and real-world experiment results for single and multi-clu…
Doubly SGD improves convergence for intractable objective optimization.
New method reduces variance in complex probabilistic model optimization.
The paper calculates moments and conditional risks for skewed elliptical distributions.
CPME embeds counterfactual outcomes in RKHS for flexible policy evaluation.
New bounds on self-normalized martingales improve online linear regression performance.
FDSKL algorithm trains vertically partitioned data with kernels securely and efficiently.
Geometric approach for unsupervised word embedding alignment.
We study the problem of off-policy critic evaluation in several variants of value-based off-policy actor-critic algorithms. Off-policy actor-critic algorithms require an off-policy critic evaluation step, to estimate the value of the new policy after every policy gradient update. Despite enormous success of off-policy …
This paper discusses properties of a Doubly Stochastic Poisson Process (DSPP) where the intensity process belongs to a class of affine diffusions. For any intensity process from this class we derive an analytical expression for probability distribution functions of the corresponding DSPP. A specification of our results…
The stochastic gradient descent has been widely used for solving composite optimization problems in big data analyses. Many algorithms and convergence properties have been developed. The composite functions were convex primarily and gradually nonconvex composite functions have been adopted to obtain more desirable prop…
The paper improves boundary detection and density estimation on noisy data.
The softmax representation of probabilities for categorical variables plays a prominent role in modern machine learning with numerous applications in areas such as large scale classification, neural language modeling and recommendation systems. However, softmax estimation is very expensive for large scale inference bec…
When training a machine learning model with observational data, it is often encountered that some values are systemically missing. Learning from the incomplete data in which the missingness depends on some covariates may lead to biased estimation of parameters and even harm the fairness of decision outcome. This paper …
Doubly stochastic learning algorithms are scalable kernel methods that perform very well in practice. However, their generalization properties are not well understood and their analysis is challenging since the corresponding learning sequence may not be in the hypothesis space induced by the kernel. In this paper, we p…
We study a doubly reflected backward stochastic differential equation (BSDE) with integrable parameters and the related Dynkin game. When the lower obstacle and the upper obstacle of the equation are completely separated, we construct a unique solution of the doubly reflected BSDE by pasting local solutions and…
A (1,1) knot K in a 3-manifold M is a knot that intersects each solid torus of a genus 1 Heegaard splitting of M in a single trivial arc. Choi and Ko developed a parameterization of this family of knots by a four-tuple of integers, which they call Schubert's normal form. This article presents an algorithm for construct…
Computing partition functions, the normalizing constants of probability distributions, is often hard. Variants of importance sampling give unbiased estimates of a normalizer Z, however, unbiased estimates of the reciprocal 1/Z are harder to obtain. Unbiased estimates of 1/Z allow Markov chain Monte Carlo sampling of "d…
Improved Transformer performance by addressing 'explaining away' effect.
This paper studies semiparametric contextual bandits, a generalization of the linear stochastic bandit problem where the reward for an action is modeled as a linear function of known action features confounded by an non-linear action-independent term. We design new algorithms that achieve regret …
DSVNP uses global and local latent variables for improved neural process predictions.
Paper extends LME models to allow sign constraints on coefficients with SDTN random effects.
Let be the space of properly embedded minimal tori in quotients of by two independent translations, with any fixed (even) number of parallel ends. After an appropriate normalization, we prove that is a 3-dimensional real analytic manifold that reduces to the finite coverings of the ex…
New method for estimating parameters in inverse problems using double robustness.
The general perception is that kernel methods are not scalable, and neural nets are the methods of choice for nonlinear learning problems. Or have we simply not tried hard enough for kernel methods? Here we propose an approach that scales up kernel methods using a novel concept called "doubly stochastic functional grad…
Paper tackles causal inference with partially labeled data, introducing robust methods.
Improved estimators for causal inference using cross-fitting and undersmoothing.
We introduce a doubly stochastic proximal gradient algorithm for optimizing a finite average of smooth convex functions, whose gradients depend on numerically expensive expectations. Our main motivation is the acceleration of the optimization of the regularized Cox partial-likelihood (the core model used in survival an…
Graph alignment problem solved with convex relaxations for correlated matrices.
Gaussian processes (GPs) are a good choice for function approximation as they are flexible, robust to over-fitting, and provide well-calibrated predictive uncertainty. Deep Gaussian processes (DGPs) are multi-layer generalisations of GPs, but inference in these models has proved challenging. Existing approaches to infe…
Paper proposes a new DR estimator for adaptive experiments with improved performance.
ADSGD method speeds up model identification in sparse optimization.
π-GNN learns soft permutations for graph representations, improving graph classification and regression.
New algorithms learn graph structures privately, matching best results.
We propose a doubly stochastic primal-dual coordinate optimization algorithm for empirical risk minimization, which can be formulated as a bilinear saddle-point problem. In each iteration, our method randomly samples a block of coordinates of the primal and dual solutions to update. The linear convergence of our method…
S2M optimizes mining for diverse data subpopulations.
Using an obstruction based on Donaldson's theorem, we derive strong restrictions on when a Seifert fibered space over an orientable base surface can smoothly embed in . This allows us to classify precisely when smoothly embeds provided , where $…
Accelerates Birkhoff projection for manifold-constrained hyper-connections with high accuracy and speed.
Proposes a new simulator for complex arrival processes.
New method improves robustness of double robust estimators under complete misspecification.
Contextual multi-armed bandit algorithms are widely used in sequential decision tasks such as news article recommendation systems, web page ad placement algorithms, and mobile health. Most of the existing algorithms have regret proportional to a polynomial function of the context dimension, . In many applications ho…