Replacing MSE with f-divergence in diffusion models improves robustness under data contamination.
problem Improving robustness of diffusion models under data contamination.
method Replacing MSE with f-divergence in diffusion models.
result Empirical improvement in performance under data contamination.
AlphaNet improves supernets training with alpha-divergence.
problem Improving the uncertainty distillation in weight-sharing NAS.
method Proposes alpha-divergence for better uncertainty distillation in supernets.
result Significant improvements in model performance across various FLOPs regimes.
The article recovers tensor fields from partial data using weighted divergent ray transforms.
problem Recovering tensor fields from partial data.
method Weighted divergent ray transforms, unique continuation property of fractional Laplacian, explicit reconstruction formulas.
result Recovery of symmetric m-tensor fields and unique continuation for vector fields and symmetric 2-tensor fields. Proposes tail-adaptive f-divergences for better inference.
problem Inference difficulties with heavy-tailed importance weights.
method Tail-adaptive f-divergences that change convex function with importance weights tail.
result Significant advantages over classical KL and α-divergences.
Estimates causal effects using neural networks for balancing covariates.
problem Estimating causal effects from observational data.
method Neural Balancing Weights (NBW) using α-divergence for density ratio estimation. result Generalized approach for balancing multidimensional data.
A method for combining classifiers from multiple views using Bregman divergences.
problem Combining classifiers from multiple views with limited labeled data.
method Jointly learns view-specific and overall weighted majority vote classifiers using Bregman divergences.
result Empirical results show improved classifier performance with limited labeled data.
A new method for averaging model predictions using minimum divergence.
problem Improving model averaging methods, especially in small samples.
method Minimum divergence framework for model weight calculation.
result Empirically outperforms standard model averaging methods.
New aggregation strategy handles unbounded losses with regret bounds.
problem Online optimization with unbounded loss functions.
method Follow The Regularized Leader (FTRL) with φ-divergence.
result Worst regret bound for unbounded losses with alternative divergences.
We speed up marginal inference by ignoring factors that do not significantly contribute to overall accuracy. In order to pick a suitable subset of factors to ignore, we propose three schemes: minimizing the number of model factors under a bound on the KL divergence between pruned and full models; minimizing the KL dive…
A new algorithm, Weighted Contrastive Divergence (WCD), improves on Contrastive Divergence (CD) for learning Boltzmann architectures.
problem Computational infeasibility of exact gradient computation in Boltzmann architectures.
method Proposes Weighted Contrastive Divergence (WCD) as a modification of Contrastive Divergence (CD) with small modifications to the negative phase.
result Experimental results show significant improvement of WCD over standard CD and persistent CD with minimal additional computational cost.
A new, computationally friendly formula for a class of risk-averse preferences.
problem Characterizing a class of risk-averse preferences called uniformly weighted divergence preferences.
method Introducing a new formula that characterizes UWDP as the translation-invariant hull of state-independent expected utility.
result UWDP are the translation-invariant hull of state-independent expected utility over L0. Paper proposes a new strategy to improve initial performance of federated models.
problem Weight divergence in Federated Averaging (FedAvg) leads to poor initial performance in federated models.
method Local continual training with importance weights evaluated on a proxy dataset.
result The method significantly improves the initial performance of federated models with minimal extra communication costs.
Upper bounds for eigenvalues on submanifolds in weighted manifolds.
problem Eigenvalue bounds for submanifolds in weighted Riemannian manifolds.
method Proving upper bounds for divergence-type operators and Steklov problems on submanifolds.
result Reilly-type upper bounds for eigenvalues.
Paper proposes a method to train ML on sPlot background data without negative weights.
problem Training machine learning on data with sPlot background subtraction leads to negative weights and algorithm divergence.
method Proposes a rigorous mathematical approach to handle negative weights in sPlot background data.
result Allows the use of any machine learning method on sPlot background data samples without encountering negative weights.
Optimal weights improve particle-based approximations of discrete distributions.
problem Improving particle-based approximations of discrete distributions.
method Proving optimality of weights and showing how to compute them efficiently.
result Optimal weights can be computed from existing particle-based methods without extra costs.
Study on Lp affine surface areas and their inequalities for convex bodies.
problem Understanding weighted Lp affine surface areas in convex bodies. method Investigating valuations, isoperimetric inequalities, and connections to f divergences. result Established isoperimetric inequalities for weighted Lp affine surface areas. PSDR improves robustness against noisy labels by penalizing KL divergence between similar inputs.
problem Robust training of DNNs in datasets with noisy labels.
method Introduces PSDR, a manifold regularizer that penalizes KL divergence between similar inputs.
result Significantly improves robustness against noisy labels on benchmark datasets.
New algorithm improves accuracy of importance weights for diverse applications.
problem Improving accuracy of importance weights for various applications.
method Formulated multicalibrated partitions and developed an efficient algorithm.
result Algorithm significantly improves accuracy of importance weights.
New proof of Willmore inequality using geometric divergence inequality.
problem Proving the Willmore inequality for bounded domains.
method Using a parametric geometric inequality derived from a divergence form geometric differential inequality.
result New proofs of quantitative Willmore-type and weighted Minkowski inequalities.
Paper formalizes and analyzes a new bound for variational inference.
problem Lack of theoretical guarantees in variational algorithms.
method Introduces VR-IWAE bound, a generalization of IWAE.
result VR-IWAE bound leads to unbiased gradient estimators.
Study calibrates high-dimensional binary classifiers using angle between estimator and true weights.
problem Calibrating high-dimensional binary classifiers with provable properties.
method Interpolates with a chance classifier to construct well-calibrated predictor based on angle between estimator and true weights.
result Angular calibration approach is provably well-calibrated in high dimensions, minimizing Bregman divergence.
We compute all 2-covariant tensors naturally constructed from a semiriemannian metric which are divergence-free and have weight greater than -2. As a consequence, it follows a characterization of the Einstein tensor as the only, up to a constant factor, 2-covariant tensor naturally constructed from a semiriemannian met…
This paper presents a novel theoretical study of the general problem of multiple source adaptation using the notion of Renyi divergence. Our results build on our previous work [12], but significantly broaden the scope of that work in several directions. We extend previous multiple source loss guarantees based on distri…
Paper introduces f-divergence variational inference for broader application.
problem Variational inference limited to specific divergences.
method Generalizes variational inference to all f-divergences using f-divergence minimization.
result Unified framework for variational inference with arbitrary f-divergences.
This research explores using Alpha-Divergences in variational dropout for better inference.
problem Improving variational inference methods using alternative divergences.
method Extending the Stochastic Gradient Variational Bayes (SGVB) framework with Alpha-Divergences.
result The α-divergence with αightarrow1 yields the lowest training error and optimizes the ELBO. Estimates KL divergence with fairness considerations for sub-populations.
problem Fairly estimate KL divergence between distributions considering sub-populations.
method Proposes multi-group attribution for KL divergence estimation, derived from multi-calibration.
result Shows multi-group attribution provides better KL divergence estimates conditioned on sub-populations.
ETM identifies field-specific keywords in text classification.
problem Unsupervised text classification with field-specific keywords.
method Weighted Lasso penalty and pairwise Kullback-Leibler divergence penalty for topic separation.
result ETM improves topic coherence by 22% and 10% compared to LDA.
BHLR predicts hyperlink weights from data vectors using symmetric similarity functions and Bregman divergence.
problem Predicting hyperlink weights from data vectors in a general framework.
method BHLR learns a symmetric similarity function to minimize Bregman-divergence between hyperlink weights and estimated similarities.
result BHLR is statistically consistent and computationally tractable, providing theoretical guarantees for various methods.
DAIS minimizes symmetrized KL divergence between initial and target distributions.
problem Optimizing over initial distributions in importance sampling.
method Differentiable annealed importance sampling (DAIS) minimizing symmetrized KL divergence.
result DAIS minimizes symmetrized KL divergence between initial and target distributions.
The paper analyzes SBL pruning criteria under weakened assumptions.
problem Sparse Bayesian learning hyperparameter divergence and pruning.
method Analyzing marginal likelihood function under weakened Gaussian assumptions.
result Conditions for finite vs infinite hyperparameters lead to F-SBL pruning.
CO2 algorithm creates coresets for generic smooth divergences efficiently.
problem Efficiently creating coresets for generic smooth divergences.
method CO2 algorithm using functional Taylor expansion and maximum mean discrepancy minimization.
result Poly-logarithmically many data points suffice for Sinkhorn divergence approximation.
fBNNs use stochastic processes for variational inference in neural networks.
problem Difficulties in specifying priors and posteriors in high-dimensional weight spaces.
method Maximize Evidence Lower Bound (ELBO) on stochastic processes, using spectral Stein gradient estimator.
result fBNNs provide reliable uncertainty estimates and extrapolate well with structured priors.
In high-dimensional data, many sparse regression methods have been proposed. However, they may not be robust against outliers. Recently, the use of density power weight has been studied for robust parameter estimation and the corresponding divergences have been discussed. One of such divergences is the γ-divergence a…
Improved hypothesis testing and change-point detection using diffusion-based methods.
problem Limited power of score-based hypothesis tests and change-point detection.
method Extending score-based Fisher divergence to diffusion-divergence by multiplying score functions with a matrix-valued function or weight matrix.
result Theoretical quantification and demonstration of optimal performance of diffusion-based algorithms.
New Stein operator improves robustness in model inference.
problem Improving robustness in inference for unnormalized models.
method Density-power weighted Stein operator (γ-Stein operator). result Robust methods for goodness-of-fit testing and posterior approximation.
Contrastive Divergence (CD) and Persistent Contrastive Divergence (PCD) are popular methods for training the weights of Restricted Boltzmann Machines. However, both methods use an approximate method for sampling from the model distribution. As a side effect, these approximations yield significantly different biases and…
This paper introduces a new method to train normalizing flows using precision-recall divergences.
problem Training generative models with mode dropping and low-quality samples.
method Introduces PR-divergences and proposes a novel generative model to minimize precision-recall trade-offs.
result Normalizing flows can be trained to achieve specific precision-recall trade-offs using PR-divergences.
SLERP interpolation optimizes dynamic weight rebalancing in AMMs.
problem Optimizing dynamic weight rebalancing in automated market makers (AMMs).
method Riemannian geometry and SLERP interpolation.
result SLERP interpolation minimizes the KL divergence loss in dynamic weight rebalancing.
Generative neural samplers are probabilistic models that implement sampling using feedforward neural networks: they take a random input vector and produce a sample from a probability distribution defined by the network weights. These models are expressive and allow efficient computation of samples and derivatives, but …
New method compresses neural networks using random code, improving efficiency.
problem Large memory footprint of deep neural networks.
method Training a variational distribution over weights, encoding using Kullback-Leibler divergence.
result Achieves state-of-the-art compression rates and test performance.
Proposes a new method for kernel density estimation using stagewise minimization and a simple dictionary.
problem Kernel density estimation with data-adaptive weighting parameters and sparse representation.
method Stagewise minimization algorithm based on U-divergence and a simple dictionary. result Develops non-asymptotic error bound for the proposed estimator.
cGAN learns a distribution for causal inference without specifying P.
problem Enforcing strong ignorability in causal analyses of observational data.
method Generative adversarial network (GAN)-based model called the Counterfactual χ-GAN (cGAN). result Minimizes Pearson χ2 divergence, maximizing coverage and minimizing variance of ATE estimates. Paper addresses variational inference issues in Bayesian neural networks.
problem Negative infinite ELBO for function-space priors in BNNs.
method Regularized KL divergence for well-defined function-space variational inference.
result Method provides competitive uncertainty estimates for BNNs.
A new associative memory uses Sinkhorn divergence for efficient pattern retrieval.
problem Efficiently retrieving patterns from large datasets of weighted point clouds.
method Derived retrieval dynamics as a SHK gradient flow, discretized for a deterministic algorithm.
result Proved basin invariance, geometric convergence, and robust recovery from perturbations.
Modeling and learning turn-taking behaviors in multi-agent systems.
problem Modeling and predicting turn-taking behaviors in dynamic multi-agent systems.
method Individual behavior models (WFSTs) and multi-agent fusion model (logistic regression classifier).
result Accurately models and predicts turn-taking behaviors with high precision.
Study compares chi-squared divergence and KL-divergence posteriors for PAC-Bayesian bounds.
problem Investigates optimal posteriors for PAC-Bayesian bounds using chi-squared divergence.
method Analyzes bounds for three distance functions, derives FP equations for computation.
result Chi-squared divergence based posteriors have weaker bounds and worse test errors.
Adaptive sampling for multimodal distributions converges faster than classical methods.
problem Sampling from multimodal distributions efficiently.
method Adaptive linear dynamics with adaptive diffusion coefficients and vector fields, interpreted as weighted Wasserstein gradient flows.
result Derivative-free dynamics can achieve significantly faster convergence for nonconvex potentials.
A new method improves adversarial robustness by optimizing importance weights.
problem Adversarial training's non-uniform robustness across different data points.
method Doubly-robust instance reweighted adversarial training using distributionally robust optimization.
result Improves robustness against attacks on the weakest data points.