Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

64128192256 · May 202619922001200920182026
48 results for total variance divergence

The paper analyzes the bias-variance tradeoff for Bregman divergences.

problem Understanding the bias-variance tradeoff for Bregman divergences.
method Analyzes the bias-variance tradeoff through operations in dual space.
result Derives several results including a generalized law of total variance and ensembling operations.

Improved KL divergence estimators for normalizing flows lead to faster convergence and better approximations.

problem Estimating KL divergences for normalizing flows efficiently and accurately.
method Path-gradient estimators for reverse and forward KL divergences.
result Path-gradient estimators lead to faster convergence and better approximation results.

POP3D is a new reinforcement learning algorithm that improves upon PPO.

problem The shortcomings of existing reinforcement learning algorithms.
method Policy Optimization with Penalized Point Probability Distance (POP3D) as a lower bound to the square of total variance divergence.
result POP3D is highly competitive compared to PPO in various benchmarks.

Deep learning models can have low bias and variance, contrary to classical theory.

problem Understanding the performance of deep learning models at high complexity.
method Developed a fine-grained bias-variance decomposition for random feature kernel regression, analyzing the effects of sampling, initialization, and labels.
result The variance terms exhibit non-monotonic behavior and can diverge at the interpolation boundary, even in the absence of label noise.

Generalizes bias-variance decomposition for Bregman divergences.

problem No specific problem stated; generalization of bias-variance for Bregman divergences.
method Provided a generalization of the bias-variance decomposition for Bregman divergences.
result A clear, standalone derivation of the bias-variance decomposition for Bregman divergences.

This paper generalizes beta divergence beyond its classical form associated with power variance functions of Tweedie models. Generalized form is represented by a compact definite integral as a function of variance function of the exponential dispersion model. This compact integral form simplifies derivations of many pr…

2013-06-14abs ↗pdf ↗

LMC algorithm converges to target in Chi-squared and Renyi divergence.

problem Sampling from target distribution using LMC with strong dissipativity and smoothness conditions.
method LMC algorithm with strong dissipativity and first-order smoothness, initialized with Gaussian.
result LMC reaches ε-neighborhood of target in Chi-squared and Renyi divergence in O(λ²dε⁻¹) steps.

Black box variational inference (BBVI) with reparameterization gradients triggered the exploration of divergence measures other than the Kullback-Leibler (KL) divergence, such as alpha divergences. In this paper, we view BBVI with generalized divergences as a form of estimating the marginal likelihood via biased import…

2017-09-21abs ↗pdf ↗

New method for fast, accurate Gaussian process regression with data guarantees.

problem Slow and unreliable GP inference methods for nonparametric regression.
method Developed a novel objective function (preconditioned Fisher divergence) for scalable approximate GP regression with finite-data guarantees.
result Minimizing the pF divergence provides pointwise mean and variance estimates with tight 2-Wasserstein distance bounds and comparable empirical performance to variational sparse GPs.

The paper decomposes unsupervised learning's generalization error into model, data, and variance components.

problem Understanding the components of unsupervised learning's generalization error.
method Information-geometric decomposition of the Kullback-Leibler generalization error.
result The optimal rank in εε-PCA is the noise floor, balancing model-error gain and data-bias cost.

Sample variance decay is shown in deep ReLU networks, impacting training dynamics.

problem Sample variance decay in deep ReLU networks during training.
method Decomposed total variance into sample variance and network-averaged sum of sample mean and variance.
result Sample variance decays in later layers of deep ReLU networks, impacting training dynamics.

We describe the underlying probabilistic interpretation of alpha and beta divergences. We first show that beta divergences are inherently tied to Tweedie distributions, a particular type of exponential family, known as exponential dispersion models. Starting from the variance function of a Tweedie model, we outline how…

2012-09-19abs ↗pdf ↗

Large batch sizes reduce gradient variance in DP-SGD, improving privacy.

problem Understanding why large batch sizes work in DP-SGD.
method Decomposed total gradient variance into subsampling and noise-induced variances, proving batch size independence in the limit.
result Large batch sizes reduce effective total gradient variance, improving privacy in DP-SGD.

On a compact nn-dimensional manifold MM, it is well known that a critical metric of the total scalar curvature, restricted to the space of metrics with unit volume, is Einstein. It has been conjectured that a critical metric of the total scalar curvature, restricted to the space of metrics with constant scalar curvat…

2017-10-20abs ↗pdf ↗

Optimized α\alpha-posteriors reduce KL divergence from true posterior in parametric misspecification.

problem Reduction of KL divergence from true posterior in parametric model misspecification.
method Derivation of Bernstein-von Mises theorem and optimization of α\alpha-posteriors.
result Optimized α\alpha-posteriors minimize KL divergence from true posterior, especially in severe misspecification.

Paper introduces Wasserstein total correlation for disentangled representation learning.

problem Learning disentangled representations from data.
method Adversarial training of a critic to estimate Wasserstein total correlation in variational and Wasserstein autoencoders.
result Proposed method achieves comparable disentanglement performance with less reconstruction loss.

New framework estimates staged tree models using hierarchical clustering on the probability simplex.

problem Estimating staged tree models with context-specific dependencies.
method Hierarchical clustering on the probability simplex, using simplex-based divergences and linkage methods.
result Total Variation divergence with Ward.D2 linkage produces staged trees with better model fit, structure recovery, and computational efficiency.

The paper proves gap properties for critical metrics under specific conditions.

problem Proving gap properties for critical metrics under divergence-free Bach tensor condition.
method Analyzing critical point equation of total scalar curvature with divergence-free Bach tensor.
result Proves gap properties for n5n \geq 5 and a similar condition for n=4n=4.

New dispersion indices based on inaccuracy and divergence introduced for information measures.

problem Measuring variability in uncertainty measures.
method Introducing new dispersion indices based on Kerridge inaccuracy and Kullback-Leibler divergence.
result Properties, bounds, and examples of new dispersion indices presented.

Paper proposes a method to stabilize estimation of KL divergence using a discriminator in RKHS.

problem High variance and instability in estimating KL divergence using neural network discriminators.
method Developed a novel construction of the discriminator in RKHS, controlled its complexity, and proved the consistency of the estimator.
result Reduced variance and stabilized training of KL divergence estimates.

ff-divergences are a general class of divergences between probability measures which include as special cases many commonly used divergences in probability, mathematical statistics and information theory such as Kullback-Leibler divergence, chi-squared divergence, squared Hellinger distance, total variation distance e…

2013-02-02abs ↗pdf ↗

A new model corrects inhomogeneity in Optimal Transport with Boundary.

problem Inhomogeneity in UROT models for Optimal Transport with Boundary.
method Proposed a modified entropic regularization term to make UROT models homogeneous.
result Homogeneous UROT model preserves properties of standard UROT while correcting inhomogeneity.

The paper develops estimators for variance in graph structures using fused lasso.

problem Variance estimation in graph-structured problems.
method Developed linear time estimator for homoscedastic case and total variation regularization estimator for heteroscedastic case.
result Minimax rates and consistency for variance estimation in various graph structures.

Privacy amplification improved through contraction coefficients and EγE_γ-divergence.

problem Improving privacy guarantees in iterative algorithms.
method Using contraction coefficients derived from EγE_γ-divergence to determine differential privacy parameters.
result Tighter bounds on differential privacy parameters of iterative algorithms.

Proves Sard conjecture for specific distributions, controlling divergence of vector fields.

problem Proving the Sard conjecture for certain types of distributions.
method Constructs a singular distribution capturing essential abnormal lifts, proving the conjecture for rank 3 distributions in dimension 4 and generic corank 1 distributions.
result Proves the Sard conjecture for generic co-rank one distributions.

This work analyzes the statistical properties of adaptive gradient methods.

problem Lack of understanding of the statistical properties of adaptive gradient methods.
method Theoretical analyses and experiments on the variance of update magnitudes.
result The variance of update magnitudes is an increasing and bounded function of time, not diverging.

Two approaches integrate qualitative views into portfolio optimization, showing aggregation methods outperform robust optimization.

problem Incorporating qualitative views into portfolio optimization models.
method Robust optimization and order aggregation methods.
result Aggregation methods outperform robust optimization in portfolio performance analysis.

Optimizes deep neural network initialization variance for better performance.

problem Improving deep neural network performance through optimal initialization variance.
method Using SGD dynamics and Fokker-Planck equations, we study the relationship between initialization and expected loss function.
result An optimal condition for initialization variance that leads to lower training loss and higher test accuracy.

We quantify predictive uncertainty using the posterior predictive variance.

problem Quantifying uncertainty in predictive models.
method Using the law of total variance, we generate expansions for the posterior predictive variance.
result Identify the main contributors to prediction intervals and quantify term-wise uncertainty.

A new model relaxes constraints on exponential dispersion models.

problem Tight conditions on cumulant function limit the class of exponential dispersion models.
method Introduces K-LED model with Legendre cumulant function and Bregman divergence guidance.
result The model allows for easier computation of mean parameter and includes various distributions.

New method for optimizing complex composite functions with reduced variance.

problem Optimizing multi-level composite functions with nested random and smooth mappings.
method Normalized proximal approximate gradient (NPAG) method with nested stochastic variance reduction.
result Total sample complexity of O(ε3)O(ε^{-3}) in expectation and O(N+Nε2)O(N+\sqrt{N}ε^{-2}) in finite-sum cases.

This work generalizes calibeating for a broader range of proper losses using Bregman divergence.

problem Calibration for a wide range of proper losses beyond Brier and log loss.
method Regret minimization based on Bregman divergence for a family of proper losses.
result U-calibration results for a family of Tsallis losses with logarithmic regret and dimension independence.

QP improves Gaussian process inference by minimizing Wasserstein distance.

problem Approximate inference in Gaussian processes using KL divergence is inadequate.
method Quantile Propagation (QP) minimizes Wasserstein distance instead of KL divergence.
result QP outperforms EP and variational Bayes in classification and Poisson regression.

The paper improves generalization bounds using interpolation between various divergences.

problem Improving generalization bounds in machine learning.
method Derives new PAC-Bayes generalization bounds based on (f,Γ)(f, Γ)-divergence and interpolates between various divergences.
result Connects derived bounds to earlier statistical learning results and provides practical training objectives.