This work generalizes Log-Determinant divergences to infinite-dimensional settings.
problem Generalizing Log-Determinant divergences to infinite-dimensional spaces.
method Introducing a parametrized family of divergences, Alpha-Beta Log-Determinant divergences, for positive definite unitized trace class operators.
result The Alpha-Beta Log-Det divergences encompass various divergences and metrics, including the affine-invariant Riemannian distance and symmetric Stein divergence.
Study on geometric Jensen-Shannon divergence for Gaussian measures in Hilbert space.
problem Computing divergence between Gaussian measures in infinite-dimensional Hilbert space.
method Closed form expression and regularization for divergence calculation.
result Closed form expression and regularization for Geometric Jensen-Shannon divergence.
Optimizes machine learning by approximating log determinants efficiently.
problem Computational expense in calculating log determinants for large data sets.
method Demonstrates the optimality of Maximum Entropy methods in approximating log determinants.
result Reduction of KL divergence between proposal and true eigenvalue distribution by adding more moments.
An algorithm learns a kernel matrix from relative-distance constraints for semi-supervised clustering.
problem Learning metrics from relative-distance constraints to capture finer structures.
method Log determinant divergence for kernel matrix learning with relative-distance constraints.
result Kernels learned from relative-distance constraints yield better clusterings than existing methods.
Paper shows how sparse inversion speeds up log determinant derivatives.
problem Deriving log determinant derivatives for sparse matrices.
method Sparse inversion, selected inversion, accelerates computation.
result Derivative of log determinant can be computed faster with sparse inversion.
Bayesian approach estimates log-determinant with uncertainty quantification.
problem Intractable computation of log-determinant in large kernel matrices.
method Reinterpreting as Bayesian inference with prior bounds and evidence.
result Probabilistic estimates of log-determinant and uncertainty.
The paper develops divergences for Gaussian processes and RKHS settings.
problem Estimating divergences in infinite-dimensional spaces.
method Formulations of Alpha Log-Det divergences, continuity in norm, laws of large numbers, consistent estimation from finite samples.
result Infinite-dimensional divergences can be estimated from finite-dimensional versions with dimension-independent sample complexities.
Estimates log determinants using entropy for scalable machine learning.
problem Scalable calculation of matrix determinants is a bottleneck in machine learning.
method Maximum entropy framework with moment constraints for stochastic trace estimation.
result Significant improvement over state-of-the-art methods on various UFL sparse matrices.
New method estimates log-determinant using trace powers, avoiding classical limitations.
problem Estimating log-determinant of large matrices efficiently and accurately.
method Interpolating moment-generating function and its derivative at zero using trace powers.
result No continuous estimator using finite moments can be uniformly accurate over unbounded conditioning.
A new method reduces log determinant evaluation cost from cubic to quadratic.
problem Efficiently evaluating log determinants in machine learning.
method Variational Bayesian approximation with complexity O(n^2).
result State-of-the-art performance on synthetic and real-world datasets.
Novel algorithm speeds up log-determinant estimation for large matrices.
problem Efficiently estimating log-determinants of large positive definite matrices under memory constraints.
method Hierarchical algorithm based on block-wise computation of LDL decomposition.
result Accurate estimation of NTK log-determinants from a tiny fraction of the full dataset.
New scalable methods for log determinant computations speed up Gaussian process kernel learning.
problem Prohibitive computational cost of log determinant calculations for Gaussian process kernel learning.
method Stochastic approximations based on Chebyshev, Lanczos, and surrogate models.
result Lanczos method is superior for kernel learning, and surrogate models are highly efficient and accurate.
The paper improves support recovery in high-dimensional precision matrix estimation using meta learning.
problem Support recovery in high-dimensional precision matrix estimation with reduced sample complexity.
method Pooling samples from different tasks and using an improper ℓ1-regularized log-determinant Bregman divergence to estimate a single precision matrix. result The support of the improperly estimated single precision matrix is equal to the true support union with high probability.
Paper proves Shapley value convergence in Bayesian learning games.
problem Measuring contributions in cooperative games using Bayesian inference.
method Established convergence of Shapley value in parametric Bayesian learning games.
result Shapley value differences converge in probability to a limiting game.
Given i.i.d. observations of a random vector X∈Rp, we study the problem of estimating both its covariance matrix Σ∗, and its inverse covariance or concentration matrix {Θ∗=(Σ∗)−1.} We estimate Θ∗ by minimizing an ℓ1-penalized log-determinant Bregman divergence; in the multivariate G…
New method computes affine normal directions efficiently for sparse polynomials.
problem Computing affine normal directions is computationally expensive in high dimensions.
method Reduces third-order tensor contraction to matrix-free formulation using log-determinant gradient.
result Scalable implementations with near-linear scaling in dimension and sparsity.
A new algorithm improves sampling for graph learning models.
problem Euclidean proposals struggle near the boundary of PSD matrices.
method ConeMALA, a geometry-aware Langevin algorithm.
result ConeMALA achieves higher ESS/sec and stable diagnostics.
Consider a random vector with finite second moments. If its precision matrix is an M-matrix, then all partial correlations are non-negative. If that random vector is additionally Gaussian, the corresponding Markov random field (GMRF) is called attractive. We study estimation of M-matrices taking the role of inverse sec…
Proposes a new method to improve regression models with reweighted samples.
problem Improves regression models' performance under low sample sizes and covariate perturbations.
method Reparametrizes sample weights using a doubly non-negative matrix and solves the reweighted estimate efficiently.
result Adversarial reweighting strategy delivers promising results on various datasets.
MEMe efficiently approximates large-scale ML problems with hundreds of moments.
problem Efficient approximation in large-scale machine learning.
method Maximum entropy algorithm with hundreds of moments for computationally efficient approximations.
result Superior to existing approaches in fast log determinant estimation and Bayesian optimisation.
Paper improves efficiency in matrix computations for Gaussian processes.
problem Efficiency in matrix computations for Gaussian processes.
method Variance reduction via matrix factorization.
result Factorized estimator can be up to 1,000 times more efficient.
A new method speeds up training of deep models by avoiding Jacobian determinant computation.
problem Efficiently training deep neural networks with complex log-determinant terms.
method Relative gradients to compute Jacobian updates efficiently.
result Training neural networks with Jacobian log-determinant objectives becomes feasible.
DAGMA learns DAGs faster and more accurately using log-determinant acyclicity.
problem Learning directed acyclic graphs from data efficiently and accurately.
method DAGMA uses M-matrices and log-determinant acyclicity to optimize DAG learning.
result DAGMA achieves faster and more accurate DAG learning compared to existing methods.
The paper connects complex normalizing flows to Kähler-Ricci flows using geometric and statistical perspectives.
problem Understanding the relationship between complex normalizing flows and Kähler-Ricci flows.
method Develops connections between complex normalizing flows and Kähler-Ricci flows by relating the log determinant to Ricci curvature and using a Bayesian perspective.
result Reconciles the complex normalizing flow and Kähler-Ricci flow, showing they are related under certain conditions.
Develops interpolation methods for matrix functions in statistics and machine learning.
problem Estimating matrix functions in statistics and machine learning.
method Interpolates log-determinant and trace of matrix powers using modified sharp bounds.
result Accuracy and performance demonstrated in numerical examples.
This paper proposes an online MTL framework that improves scalability and robustness.
problem Efficient online multi-task learning with correlated and personalized task structures.
method Decomposes weight matrix into low-rank common structure and personalized patterns using nuclear norm and group lasso.
result Achieves sub-linear regret and improved performance with log-determinant function.
Matrix formulas for knot invariants derived from Tait graphs.
problem Computing knot invariants for alternating links.
method Squarefree matrix extraction from Tait graph vertices.
result Explicit formulas for CWRk for k≥4. New geometric structures defined on SPD matrices for better understanding.
problem Understanding SPD matrices and their geometric properties.
method Introducing Finslerian and dual information-geometric structures on James' bicone domain.
result Geodesics correspond to straight lines in coordinate systems, and new dissimilarities generalize existing ones.
We establish sharp Sobolev inequalities of order four on Euclidean d-balls for d greater than or equal to four. When d=4, our inequality generalizes the classical second order Lebedev-Milin inequality on Euclidean 2-balls. Our method relies on the use of scattering theory on hyperbolic d-balls. As an application, we ch…
Polylab is a MATLAB toolbox for multivariate polynomial modeling.
problem Efficiently modeling and manipulating multivariate polynomials across CPU and GPU.
method Unified symbolic-numeric interface, three aligned classes (MPOLY, MPOLY_GPU, MPOLY_HP), polynomial operations, differentiation, matrix computations.
result Advantages of MPOLY-HP for reduction-heavy simplification and large-scale computations, and the stochastic log-determinant variant for sparse regimes.
Paper introduces VDE, a variance-reduced determinant estimator.
problem Estimating determinants with low variance and efficiency.
method Combines variational inference and spherical normalizing flows.
result VDE achieves zero variance in ideal cases, requiring only one sample.
HCLM framework uses entropy regularization for open learning systems.
problem Real-world AI challenges and limitations of deep learning.
method Dynamical and information-theoretic framework with entropy regularization.
result Geometric entropy surrogates, especially log-determinant covariance entropy, induce stronger and more stable information forces.
New algorithm for online portfolio selection with reduced runtime.
problem Maximizing total return in online portfolio selection.
method Minimizes current logarithmic loss regularized by log-determinant of Hessian.
result Achieves regret guarantee similar to Universal Portfolios with reduced runtime.
The curvature of the noncommutative torus Tθ2 (θ irrational) endowed with a noncommutative conformal metric has been the focus of attention of several recent works. Continuing the approach taken in the paper [A. Connes and H. Moscovici, http://arxiv.org/abs/1110.3500] we extend the study of the curvature to twist…
Invertible ResNets enable classification, density estimation, and generation.
problem Enforcing invertibility in ResNets without architectural changes.
method Simple normalization during training to make ResNets invertible.
result Invertible ResNets achieve competitive performance with single architecture.
New divergences extend Bregman and skew Jensen, including f-divergences.
problem Developing new divergences to include f-divergences.
method Introducing g-Bregman and skew g-Jensen divergences, showing they include f-divergences.
result g-divergences generalize existing divergences and inequalities.
The L1-regularized Gaussian maximum likelihood estimator (MLE) has been shown to have strong statistical guarantees in recovering a sparse inverse covariance matrix, or alternatively the underlying graph structure of a Gaussian Markov Random Field, from very limited samples. We propose a novel algorithm for solving the…
Paper shows how to break down a specific type of divergence into simpler parts.
problem Understanding and simplifying divergence functions.
method Decomposes the symmetric Bregman divergence into two types of Jensen divergences and a Bregman divergence, and extends this to include f-divergences.
result Sum decomposition of divergence into simpler parts is possible.
A new coordinate system for SPD matrices simplifies computations and generative modeling.
problem Computing and modeling SPD matrices
method Reverse telescoping coordinate system
result Significantly reduces computational complexity and facilitates generative modeling.
New iterative solvers speed up Gaussian process regression with derivatives.
problem Scaling Gaussian process regression with derivatives for high-dimensional problems and large budgets.
method Iterative solvers using fast matrix-vector multiplications and pivoted Cholesky preconditioning.
result Bayesian optimization with derivatives can now scale to high-dimensional problems and large evaluation budgets.
New algorithm speeds up cluster-based compressive sensing tasks.
problem Efficiently solving multiple compressive sensing tasks with shared information.
method Combines Monte Carlo sampling with iterative linear solvers to avoid explicit covariance matrix computation.
result Up to thousands of times faster and orders of magnitude more memory-efficient compared to existing methods.
Study explores relationship between Hölder and FDPD divergences.
problem Understanding the relationship between Hölder and FDPD divergences.
method Intersection and generalization of divergence families, proving nonnegativity, deriving inequalities.
result Established ξ-Hölder divergence and derived inequalities. Gradient descent on LSE objectives implicitly performs EM, leading to collapse without volume control.
problem Gradient collapse in autoencoders without volume control.
method Introduced a single-layer encoder with an LSE objective and InfoMax regularization for volume control.
result Gradient--responsibility identity holds exactly; LSE alone collapses; variance prevents dead components; decorrelation prevents redundancy.
Unified representation of density-power-based divergences simplifies estimation to M-estimation.
problem Outliers in density estimation.
method Define a norm-based Bregman density power divergence (NB-DPD) that reduces to M-estimation.
result NB-DPD connects and generalizes existing divergences, highlighting robustness properties.
This paper improves active learning by using robust divergences for committee disagreement.
problem Active learning with high measurement costs.
method Query by committee with Bregman divergence (including Kullback-Leibler divergence as a special case).
result The proposed method is more robust and performs as well as or better than conventional methods.
Low-rank matrix is desired in many machine learning and computer vision problems. Most of the recent studies use the nuclear norm as a convex surrogate of the rank operator. However, all singular values are simply added together by the nuclear norm, and thus the rank may not be well approximated in practical problems. …
New divergence measures improve KL approximation.
problem Improving KL divergence approximation without AC condition.
method Introduced α-geodesical skew divergence. result Properties of α-geodesical skew divergence studied. New method to study group invariants using divergence spectra.
problem Understanding group invariants through divergence.
method Introducing divergence spectrum to compare classical notions and study relatively hyperbolic groups.
result Existence of groups with exponential divergence but different divergence spectra.