Divergence functions play a key role as to measure the discrepancy between two points in the field of machine learning, statistics and signal processing. Well-known divergences are the Bregman divergences, the Jensen divergences and the f-divergences. In this paper, we show that the symmetric Bregman divergence can be …
Generalizes bias-variance decomposition for Bregman divergences.
problem No specific problem stated; generalization of bias-variance for Bregman divergences.
method Provided a generalization of the bias-variance decomposition for Bregman divergences.
result A clear, standalone derivation of the bias-variance decomposition for Bregman divergences.
A method to compute divergences between decomposable models, useful in supervised learning.
problem Computing exact divergences between high-dimensional distributions is intractable.
method Proposes an approach to compute exact alpha-beta divergences between marginal and conditional distributions of decomposable models.
result Tractable computation of marginal and conditional alpha-beta divergences.
Unified framework for high-dimensional online learning with non-divergent error bounds and adaptive gains.
problem Divergence of error bounds in high-dimensional online learning as data batches increase.
method Asynchronous decomposition framework with summary statistics and dynamic regularization.
result Non-divergent error bounds and adaptive gains in sparse online optimization.
We present a novel nonnegative tensor decomposition method, called Legendre decomposition, which factorizes an input tensor into a multiplicative combination of parameters. Thanks to the well-developed theory of information geometry, the reconstructed tensor is unique and always minimizes the KL divergence from an inpu…
Paper studies statistical manifolds with logarithmic divergences.
problem Understanding statistical manifolds induced by logarithmic divergences.
method Constructs dual foliation of the statistical manifold.
result Extends dual foliation of a dually flat manifold.
Gaussian processes improved for ocean current reconstruction and divergence identification.
problem Reconstructing ocean currents from sparse buoy data.
method Proposed a Helmholtz decomposition-based approach to Gaussian processes for better physical modeling.
result Improved inference on ocean currents and divergence identification with minimal computational cost.
New bounds found for optimizing non-convex functions with noisy data.
problem Limits of first-order stochastic optimization in non-convex settings.
method Divergence decomposition to construct challenging subclasses.
result Sharp lower bounds on noisy gradient queries for various non-convex classes.
New risk decompositions clarify domain adaptation issues.
problem Domain adaptation challenges with different training and test distributions.
method Representation Bayesian Risk Decompositions, hybrid argument.
result Clarifies factors (2) and (3) as reasons for generalization failure.
A graph theory approach defines curl and decomposes vector fields.
problem Defining curl for vector fields on graphs and decomposing them.
method Definition of curl as orthogonal complement of circulation-free fields, proving analogues of vector field theorems.
result Helmholtz-Hodge decomposition on graphs: gradient, curl, and harmonic fields.
New method for factor analysis using nuclear and ℓ0 norms.
problem Finding a low-rank plus sparse decomposition from noisy covariance matrix.
method Formulated an optimization problem with nuclear norm, ℓ0 norm, and KL divergence. Used alternating minimization algorithm. result Algorithm effectively decomposes covariance matrices in synthetic and real datasets.
Unified algorithm for tensor decomposition supports multiple loss functions and models.
problem Efficient tensor decomposition for various models and loss functions.
method Hierarchical combination of ADMM and MM for optimization.
result Wide-range applications can be solved by the proposed algorithm.
This is the second in a series of papers where we prove a conjecture of Deser and Schwimmer regarding the algebraic structure of ``global conformal invariants''; these are defined to be conformally invariant integrals of geometric scalars. The conjecture asserts that the integrand of any such integral can be expressed …
New GPs model edge functions on complex networks, capturing divergence and curl.
problem Modeling flow data on networks with independent learning of Hodge components.
method Developed Hodge-compositional edge GPs using Hodge decomposition.
result Hodge-compositional edge GPs can represent any edge function and capture flow relevance.
A new method for decomposing non-negative tensors using energy-based modeling.
problem Challenges in traditional tensor decomposition methods, especially global optimization and rank selection.
method Energy-based modeling of tensors, considering interactions between modes for global optimization.
result Demonstrates effectiveness in tensor completion and approximation, revealing a relationship between many-body and low-rank approximations.
Deep learning models can have low bias and variance, contrary to classical theory.
problem Understanding the performance of deep learning models at high complexity.
method Developed a fine-grained bias-variance decomposition for random feature kernel regression, analyzing the effects of sampling, initialization, and labels.
result The variance terms exhibit non-monotonic behavior and can diverge at the interpolation boundary, even in the absence of label noise.
On a compact Kahler manifold, one can define global invariants by integrating local invariants of the metric. Assume that a global invariant thus obtained depends only on the Kahler class. Then we show that the integrand can be decomposed into a Chern polynomial (the integrand of a Chern number) and divergences of one …
This paper addresses the estimation of the latent dimensionality in nonnegative matrix factorization (NMF) with the β-divergence. The β-divergence is a family of cost functions that includes the squared Euclidean distance, Kullback-Leibler and Itakura-Saito divergences as special cases. Learning the model order is impo…
CDFD analyzes circularity and directionality in weighted directed networks.
problem Analyzing circularity and directionality in weighted directed networks.
method CDFD framework separates flow into circular and acyclic components.
result CDFD yields a normalized circularity index capturing flow in cycles and directionality.
Unified framework for data-free sampling using Wasserstein gradient flows.
problem Efficient sampling from unnormalized distributions without data.
method Unified theoretical framework based on Wasserstein gradient flows.
result Unified form of velocity field for various f-divergences.
A new method splits surface flow discretizations into streamfunctions and harmonic fields.
problem Discretizing incompressible flows on surfaces with pressure and saddle-point structure.
method Discrete Helmholtz-Hodge decomposition for BDM elements on surfaces.
result Eliminates pressure and saddle-point structure, ensuring exact tangentiality and divergence-freeness.
Proposes a new neural head for asymmetric representation learning.
problem Asymmetric representation learning in directed relations.
method Role-aware neural convex divergence head.
result Role-aware projections improve directional accuracy over plain ICNN-Bregman heads.
The paper decomposes unsupervised learning's generalization error into model, data, and variance components.
problem Understanding the components of unsupervised learning's generalization error.
method Information-geometric decomposition of the Kullback-Leibler generalization error.
result The optimal rank in ε-PCA is the noise floor, balancing model-error gain and data-bias cost. We consider the problem of maximum a posteriori (MAP) inference in discrete graphical models. We present a parallel MAP inference algorithm called Bethe-ADMM based on two ideas: tree-decomposition of the graph and the alternating direction method of multipliers (ADMM). However, unlike the standard ADMM, we use an inexa…
MicroRNAs (miRNAs) play crucial roles in multifarious biological processes associated with human diseases. Identifying potential miRNA-disease associations contributes to understanding the molecular mechanisms of miRNA-related diseases. Most of the existing computational methods mainly focus on predicting whether a miR…
This paper forms part of a larger work where we prove a conjecture of Deser and Schwimmer regarding the algebraic structure of "global conformal invariants"; these are defined to be conformally invariant integrals of geometric scalars. The conjecture asserts that the integrand of any such integral can be expressed as a…
Framework fine-tunes foundation models with semi-supervised learning for downstream tasks and latent spaces.
problem Training foundation models with limited labelled data.
method Mutual information decomposition for downstream and latent spaces, semi-supervised fine-tuning.
result Significant improvements in classification tasks under low-labelled conditions.
Formula derived for renormalized area of minimal submanifolds in Poincaré-Einstein manifolds.
problem Calculating the renormalized area of minimal submanifolds in Poincaré-Einstein manifolds.
method Decomposition of extrinsic Q-curvature and application to renormalized area. result Renormalized area formula expressed as a linear combination of Euler characteristic and scalar conformal submanifold invariant.
Conditional diffusion models can approximate target distributions well with Gaussian-mixture reverse kernels.
problem Approximating target distributions in conditional diffusion models.
method Using finite Gaussian mixtures with ReLU-network logits as reverse kernels, reducing the problem to static conditional density approximation.
result The resulting neural reverse-kernel class is dense in conditional KL divergence under exact terminal matching.
TRM improves long-horizon LLM RL by masking divergent sequences.
problem Long-horizon reinforcement learning with LLMs suffers from off-policy mismatch and approximation errors.
method Derives and applies trust region bounds to control divergence, proposing Trust Region Masking.
result First non-vacuous monotonic improvement guarantees for long-horizon LLM-RL.
TRM improves long-horizon reinforcement learning for LLMs by masking divergent sequences.
problem Long-horizon reinforcement learning for LLMs suffers from off-policy mismatch and approximation errors.
method Derives and applies trust region bounds to control divergence, proposing Trust Region Masking.
result First non-vacuous monotonic improvement guarantees for long-horizon LLM-RL.
ADMM algorithm solves nonlinear matrix decompositions efficiently.
problem Nonlinear matrix decompositions for various applications.
method Alternating Direction Method of Multipliers (ADMM) for nonlinear matrix factorization.
result The method efficiently solves diverse nonlinear matrix decompositions.
In this paper, we present a general, multistage framework for graphical model approximation using a cascade of models such as trees. In particular, we look at the problem of covariance matrix approximation for Gaussian distributions as linear transformations of tree models. This is a new way to decompose the covariance…
First introduced by Fernholz in stochastic portfolio theory, functionally generated portfolio allows its investment performance to be attributed to directly observable and easily interpretable market quantities. In previous works we showed that Fernholz's multiplicatively generated portfolio has deep connections with o…
Study examines stability of MOTS in symmetric initial data sets.
problem Stability of marginally outer trapped surfaces in symmetric initial data sets.
method Characterizes instability through vector decomposition and properties of normal and tangent components.
result Characterizes instability of MOTS by the nature of zero sets and divergences.
The Hodge decomposition provides a very powerful mathematical method for the analysis of 2D and 3D vector fields. It states roughly that any vector field can be L2-orthogonally decomposed into a curl-free, divergence-free, and a harmonic field. The harmonic field itself can be further decomposed into three component…
We present transductive Boltzmann machines (TBMs), which firstly achieve transductive learning of the Gibbs distribution. While exact learning of the Gibbs distribution is impossible by the family of existing Boltzmann machines due to combinatorial explosion of the sample space, TBMs overcome the problem by adaptively …
This is the fifth in a series of papers where we prove a conjecture of Deser and Schwimmer regarding the algebraic structure of ``global conformal invariants''; these are defined to be conformally invariant integrals of geometric scalars. The conjecture asserts that the integrand of any such integral can be expressed a…
Investing is a compression problem, maximizing growth by minimizing divergence.
problem Maximizing long-term wealth and minimizing risk of ruin in investing.
method Decomposes investing into three terms: money, entropy, and divergence. Uses Kelly Criterion and universal portfolio theory.
result Investing can be seen as a compression problem, with optimal strategies minimizing divergence.
Reassesses calibration metrics in machine learning models.
problem Inconsistent reporting of calibration metrics in recent literature.
method Calibration-based decomposition of Bregman divergences, visualization of calibration and generalization error.
result New visualization technique for detecting trade-offs between calibration and generalization.
New analysis shows entropy term cancels out in likelihood-based OOD detection.
problem Curious likelihood values for out-of-distribution data.
method Decomposed average likelihood into KL divergence and entropy terms.
result Entropy term explains OOD behaviour and cancels out in expectation.
The paper examines VI for overparameterized BNNs, revealing a trade-off between likelihood and KL terms.
problem Critical issue in mean-field VI training for overparameterized BNNs.
method Theoretical and empirical study of overparameterized two-layer BNNs using VI.
result A trade-off between likelihood and KL terms in overparameterized regime, with KL scaling crucial.
RFM simplifies generative modeling on complex geometries without simulation.
problem Training generative models on non-Euclidean geometries is challenging.
method Riemannian Flow Matching (RFM) constructs a premetric for efficient vector field computation.
result RFM achieves state-of-the-art performance on various non-Euclidean datasets.
Method estimates shared and study-specific factors for multi-study data.
problem Covariance estimation for multi-study data with shared and study-specific components.
method Spectral decomposition for latent factors, surrogate Bayesian regressions for loadings and variances.
result Strong frequentist guarantees and superior performance in simulations and real data.
MAP inference for general energy functions remains a challenging problem. While most efforts are channeled towards improving the linear programming (LP) based relaxation, this work is motivated by the quadratic programming (QP) relaxation. We propose a novel MAP relaxation that penalizes the Kullback-Leibler divergence…
The paper analyzes MACD using operator theory.
problem Understanding the mathematical foundation of MACD.
method Developed a functional-analytic framework interpreting MACD as a phase-corrected, smoothed derivative operator.
result MACD is structurally equivalent to a band-pass filter and can be expressed as a finite difference of delayed and doubly averaged signals.
This is the first in a series of papers where we prove a conjecture of Deser and Schwimmer regarding the algebraic structure of ``global confor- mal invariants"; these are defined to be conformally invariant integrals of geometric scalars. The conjecture asserts that the integrand of any such integral can be expressed …
Score matching errors are not sufficient for measuring diffusion model quality.
problem The L2 score matching error is not a reliable measure of diffusion model performance. method Decomposed score errors into gradient and solenoidal components and analyzed their geometric properties.
result Only the gradient component of the score error affects the marginal distributional quality.