Paper studies statistical manifolds with logarithmic divergences.
problem Understanding statistical manifolds induced by logarithmic divergences.
method Constructs dual foliation of the statistical manifold.
result Extends dual foliation of a dually flat manifold.
Generalized dual discriminator GANs improve upon traditional GANs by using two discriminators and a flexible loss function.
problem Mode collapse in GANs.
method Introducing dual discriminator α-GANs and extending the approach to arbitrary functions. result The approach reduces the optimization problem to a linear combination of an f-divergence and a reverse f-divergence. A new method for non-negative matrix factorization using generalized dual divergence.
problem Non-negative matrix factorization for various noise structures.
method Theoretical framework based on generalized dual Kullback-Leibler divergence, with algorithms developed and proven convergence using Expectation-Maximization.
result Generalizes existing methods and provides an alternative for non-negative matrix factorizations.
New dual formulation reduces generalization error for ERM-fDR.
problem Generalization error in constrained optimization problems.
method Introduces a dual formulation of ERM-fDR using Legendre-Fenchel transform and implicit function theorem.
result Explicit characterizations of generalization error for algorithms under mild conditions.
In Riemannian geometry geodesics are integral curves of the Riemannian distance gradient. We extend this classical result to the framework of Information Geometry. In particular, we prove that the rays of level-sets defined by a pseudo-distance are generated by the sum of two tangent vectors. By relying on these vector…
Dual optimization connects ERM-fDR to normalization function.
problem Empirical risk minimization with f-divergence regularization.
method Dual formulation, Legendre-Fenchel transform, implicit function theorem, nonlinear ODE.
result Computational method to calculate normalization function efficiently.
The paper analyzes the bias-variance tradeoff for Bregman divergences.
problem Understanding the bias-variance tradeoff for Bregman divergences.
method Analyzes the bias-variance tradeoff through operations in dual space.
result Derives several results including a generalized law of total variance and ensembling operations.
A recently introduced canonical divergence D for a dual structure (g,∇,∇∗) is discussed in connection to other divergence functions. Finally, open problems concerning symmetry properties are outlined.
DualVDT improves time-series forecasting with a novel dual reparametrized structure.
problem Time-series forecasting with improved performance and analytical rigor.
method Dual reparametrized variational mechanisms on VAE, latent score based generative model, reverse time stochastic differential equation, variational ancestral sampling, KL divergence reduction.
result Advanced performance in time-series forecasting with reduced KL divergence.
Study of generalized Csiszár divergences and their application to Cramér-Rao bounds.
problem Deriving lower bounds for estimator variance using generalized divergences.
method Applied Eguchi's theory to derive Fisher information metric and dual affine connections.
result More widely applicable Cramér-Rao inequality for escort distributions.
This paper improves active learning by using robust divergences for committee disagreement.
problem Active learning with high measurement costs.
method Query by committee with Bregman divergence (including Kullback-Leibler divergence as a special case).
result The proposed method is more robust and performs as well as or better than conventional methods.
Dual-ISL improves implicit generative model training with convex optimization and explicit density approximation.
problem Training implicit generative models with robust and practical likelihood-free objectives.
method Introduces dual-ISL, a novel likelihood-free objective using a convex divergence derived from the invariant statistical loss (ISL) framework.
result Dual-ISL yields a convex optimization problem in the space of model densities, providing explicit density approximation and improved training stability.
Study finds no evidence dual-class stocks are effective predictors.
problem Investment efficiency of dual-class stocks.
method In-depth analysis of stock price divergence, innovative LSTM model training set selection.
result No compelling evidence dual-class stocks are effective predictors.
New tensors reveal full curvature structure from Riemann tensor.
problem Limited information from Ricci contraction of Riemann tensor.
method Contracting double dual of Riemann tensor to reveal full curvature.
result New tensors provide canonical parents of Einstein tensor.
We show that the variational representations for f-divergences currently used in the literature can be tightened. This has implications to a number of methods recently proposed based on this representation. As an example application we use our tighter representation to derive a general f-divergence estimator based on t…
We know SGAN may have a risk of gradient vanishing. A significant improvement is WGAN, with the help of 1-Lipschitz constraint on discriminator to prevent from gradient vanishing. Is there any GAN having no gradient vanishing and no 1-Lipschitz constraint on discriminator? We do find one, called GAN-QP. To construct a …
Defines T-duality and generalised Ricci flow relations using Courant algebroid relations.
problem Establishing compatibility between T-duality and generalised Ricci flow.
method Introducing Courant algebroid relations, invariant divergence operators, and generalised isometries.
result T-duality is compatible with generalised Ricci flow, and T-dual solutions are also solutions of generalised Ricci flow.
Dual-objective GANs reduce training instabilities with tunable α-loss parameters.
problem Training instabilities in Generative Adversarial Networks (GANs).
method Introduce (αD,αG)-GANs with dual objectives modeled using α-loss. result Upper bounds on estimation error show improved performance under certain conditions.
Optimal transport with f-divergence regularization using generalized Sinkhorn algorithm.
problem Optimal transport with f-divergence regularization. method Generalized Sinkhorn algorithm for solving optimal transport problems with various f-divergences. result Strong duality holds, optimums are attained, and convergence to an optimal solution is guaranteed under certain conditions.
Introduces a new divergence measure for optimal transport.
problem Optimal transport distances and information divergences.
method Infimal convolution formulation of proximal optimal transport divergence.
result Establishes connections to dynamic formulations and partial differential equations.
Risk measures for multivariate financial positions are studied in a utility-based framework. Under a certain incomplete preference relation, shortfall and divergence risk measures are defined as the optimal values of specific set minimization problems. The dual relationship between these two classes of multivariate ris…
Logarithmic divergences linked to curvature in statistical manifolds.
problem Understanding the geometric interpretation of curvature in statistical manifolds.
method Analyzing logarithmic L(α)-divergence and its equivalence to conformal transformations and Kurose's geometric divergence. result Logarithmic divergence is a canonical divergence of a statistical manifold with constant sectional curvature −α. New algorithm solves saddle point problems in Banach spaces.
problem Solving saddle point problems in real reflexive Banach spaces.
method Stochastic Bregman Primal-Dual Splitting Algorithm with relative smoothness and strong convexity assumptions.
result Almost sure convergence to saddle points under various conditions.
In the field of statistics, many kind of divergence functions have been studied as an amount which measures the discrepancy between two probability distributions. In the differential geometrical approach in statistics (information geometry), dually flat spaces play a key role. In a dually flat space, there exist dual a…
The paper studies a generalized Pythagorean theorem on dually flat spaces via toric geometry.
problem Understanding the geometry of dually flat spaces and their toric Kähler manifolds.
method Introducing a dually flat structure and Bregman divergence on the boundary of toric Kähler manifolds.
result A continuity and generalized Pythagorean theorem for the divergence on the boundary.
EGMU optimizes portfolios using KL divergence, ensuring positive solutions.
problem Constructing multi-factor target-exposure portfolios efficiently and accurately.
method Convex optimization framework minimizing KL divergence, with explicit solvers.
result Established feasibility and uniqueness of strictly positive solutions under convex-hull conditions.
Warped product affects divergences in information geometry.
problem Warped product's impact on divergences in information geometry.
method Study of warped product on information geometry.
result Warped product does not preserve canonical divergences.
Improved GAN training stability through tunable classification losses.
problem Training instabilities in GANs.
method Reformulated GAN value function using class probability estimation (CPE) losses, defined (αD,αG)-GANs. result Tuning (αD,αG) can alleviate training instabilities. Study on 3D Lie groups finds all generalized Einstein metrics.
problem Classifying generalized Einstein metrics on 3D Lie groups.
method Developed theory of left-invariant generalized pseudo-Riemannian metrics, computed Ricci tensor, determined all metrics.
result Determined all generalized Einstein metrics on three-dimensional Lie groups.
Method prevents model divergence in rapidly changing ad markets.
problem Model divergence due to rapid ad turnover and discontinuity.
method Dual ascent optimization with latent vector constraints.
result Significant reduction in diverging instances and improved user experience/revenue.
Graph-based methods provide a powerful tool set for many non-parametric frameworks in Machine Learning. In general, the memory and computational complexity of these methods is quadratic in the number of examples in the data which makes them quickly infeasible for moderate to large scale datasets. A significant effort t…
We propose in this paper a novel approach to tackle the problem of mode collapse encountered in generative adversarial network (GAN). Our idea is intuitive but proven to be very effective, especially in addressing some key limitations of GAN. In essence, it combines the Kullback-Leibler (KL) and reverse KL divergences …
Unified framework for debiased machine learning using Riesz representer and Bregman divergence.
problem Estimating causal and structural parameters in machine learning.
method Generalized Riesz regression for fitting Riesz representer via Bregman divergence minimization.
result Automatic covariate balancing and Neyman orthogonality properties for debiased estimation.
Bregman divergences play a central role in the design and analysis of a range of machine learning algorithms. This paper explores the use of Bregman divergences to establish reductions between such algorithms and their analyses. We present a new scaled isodistortion theorem involving Bregman divergences (scaled Bregman…
Develops torsion dual connections for statistical manifolds.
problem Defining statistical manifolds using dual connections.
method Introduces torsion dual connections and proves their properties.
result Curvature tensor of torsion dual connections has specific divergence.
This work improves understanding of symmetrizing Bregman divergences on positive definite matrices.
problem Understanding which mean to use for symmetrizing Bregman divergences on positive definite matrices.
method Axiomatic definition of mean functionals and variational principles over the cone of positive definite matrices.
result The arithmetic mean is canonical for forward symmetrization, and the arithmetic, log-Euclidean, and harmonic means for reverse symmetrization.
Dual Space Preconditioning speeds up gradient descent in overparameterized models.
problem Improving convergence of gradient descent in overparameterized linear models.
method Introducing a novel preconditioner of the form ablaK for convex K and applying it to overparameterized linear models. result The iterates of the preconditioned gradient descent converge to a solution W∞ satisfying XW∞=Y. New divergences improve estimation and GAN training performance.
problem Improving estimation and training in machine learning models.
method Function-space regularized Rényi divergences.
result New divergences reduce variance and improve training performance.
A recent proposal by Ryu and Takayanagi for a holographic interpretation of entanglement entropy in conformal field theories dual to supergravity on anti-de Sitter (adS) is generalized to include entanglement entropy of black holes living on the boundary of adS. The generalized proposal is verified in boundary dimensio…
Unified framework for reinforcement learning using entropic regularization.
problem Learning optimal controllers for unknown MDPs without diverging to dangerous regions.
method Entropic proximal policy optimization with α-divergences. result Unified perspective on actor-critic architectures and asymptotic analysis of solutions.
This paper connects optimal transport and information geometry using pseudo-Riemannian geometry.
problem Understanding the geometric structures of probability distributions.
method Introducing a new differential geometric connection between optimal transport and information geometry.
result A new information-geometric interpretation of the MTW tensor.
A framework for modular training of robust generative models.
problem Training large generative models is resource-intensive and requires heuristic tuning.
method Modular training using a gating mechanism and a minimax game to find a robust gate.
result The modular approach can theoretically outperform monolithic baselines and is scalable.
Positive definite matrices abound in a dazzling variety of applications. This ubiquity can be in part attributed to their rich geometric structure: positive definite matrices form a self-dual convex cone whose strict interior is a Riemannian manifold. The manifold view is endowed with a "natural" distance function whil…
Unified framework for unlearning in diffusion models using KL divergence and likelihood constraints.
problem Removing undesirable data or concepts while preserving utility of pretrained models.
method Constrained optimization framework based on reverse and forward KL divergences, and likelihood constraints.
result Our KL-constrained approach achieves superior retention-unlearning tradeoffs compared to weight-based baselines.
A new algorithm screens negligible components to efficiently approximate optimal transport distances.
problem Efficiently approximating the Sinkhorn distance between discrete measures.
method Screening of negligible components in the dual solution of the regularized Sinkhorn problem.
result Screenkhorn algorithm provides provable guarantees with smaller computational complexity.
New samplers minimize KL divergence for constrained and non-Euclidean geometries.
problem Efficient sampling from constrained and non-Euclidean distributions.
method Stein Variational Mirror Descent and Mirrored Stein Variational Gradient Descent.
result New samplers converge more rapidly and accurately than prior methods.
A new approach to risk-sensitive reinforcement learning tackles computational challenges.
problem Computational challenges in estimating risk-sensitive policies for MDPs with finite state and action spaces.
method Proposes a new risk measure called 'caution' and uses a stochastic primal-dual method with KL divergence.
result Demonstrates improved reliability in reward accumulation without additional computational costs.
Adversarial dynamics embedding improves MLE of exponential family models.
problem Maximum likelihood estimation of exponential family models with neural network parametrization.
method Adversarial dynamics embedding to estimate the dual sampler and primal model simultaneously.
result Adversarial dynamics embedding leads to more effective learning and improved estimators compared to existing methods.