This work presents a parametrized family of divergences, namely Alpha-Beta Log- Determinant (Log-Det) divergences, between positive definite unitized trace class operators on a Hilbert space. This is a generalization of the Alpha-Beta Log-Determinant divergences between symmetric, positive definite matrices to the infi…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study on geometric Jensen-Shannon divergence for Gaussian measures in Hilbert space.
Algorithms for Gaussian process, marginal likelihood methods or restricted maximum likelihood methods often require derivatives of log determinant terms. These log determinants are usually parametric with variance parameters of the underlying statistical models. This paper demonstrates that, when the underlying matrix …
The paper develops divergences for Gaussian processes and RKHS settings.
The log-determinant of a kernel matrix appears in a variety of machine learning problems, ranging from determinantal point processes and generalized Markov random fields, through to the training of Gaussian processes. Exact calculation of this term is often intractable when the size of the kernel matrix exceeds a few t…
New method estimates log-determinant using trace powers, avoiding classical limitations.
Novel algorithm speeds up log-determinant estimation for large matrices.
Paper proves Shapley value convergence in Bayesian learning games.
The paper improves support recovery in high-dimensional precision matrix estimation using meta learning.
Given i.i.d. observations of a random vector , we study the problem of estimating both its covariance matrix , and its inverse covariance or concentration matrix {.} We estimate by minimizing an -penalized log-determinant Bregman divergence; in the multivariate G…
The scalable calculation of matrix determinants has been a bottleneck to the widespread application of many machine learning methods such as determinantal point processes, Gaussian processes, generalised Markov random fields, graph models and many others. In this work, we estimate log determinants under the framework o…
The ability of many powerful machine learning algorithms to deal with large data sets without compromise is often hampered by computationally expensive linear algebra tasks, of which calculating the log determinant is a canonical example. In this paper we demonstrate the optimality of Maximum Entropy methods in approxi…
New method computes affine normal directions efficiently for sparse polynomials.
A new algorithm improves sampling for graph learning models.
We consider the problem of metric learning subject to a set of constraints on relative-distance comparisons between the data items. Such constraints are meant to reflect side-information that is not expressed directly in the feature vectors of the data items. The relative-distance constraints used in this work are part…
Consider a random vector with finite second moments. If its precision matrix is an M-matrix, then all partial correlations are non-negative. If that random vector is additionally Gaussian, the corresponding Markov random field (GMRF) is called attractive. We study estimation of M-matrices taking the role of inverse sec…
Evaluating the log determinant of a positive definite matrix is ubiquitous in machine learning. Applications thereof range from Gaussian processes, minimum-volume ellipsoids, metric learning, kernel learning, Bayesian neural networks, Determinental Point Processes, Markov random fields to partition functions of discret…
Proposes a new method to improve regression models with reweighted samples.
Paper improves efficiency in matrix computations for Gaussian processes.
A new method speeds up training of deep models by avoiding Jacobian determinant computation.
For applications as varied as Bayesian neural networks, determinantal point processes, elliptical graphical models, and kernel learning for Gaussian processes (GPs), one must compute a log determinant of an positive definite matrix, and its derivatives - leading to prohibitive computatio…
DAGMA learns DAGs faster and more accurately using log-determinant acyclicity.
The paper connects complex normalizing flows to Kähler-Ricci flows using geometric and statistical perspectives.
Develops interpolation methods for matrix functions in statistics and machine learning.
New geometric structures defined on SPD matrices for better understanding.
Efficient approximation lies at the heart of large-scale machine learning problems. In this paper, we propose a novel, robust maximum entropy algorithm, which is capable of dealing with hundreds of moments and allows for computationally efficient approximations. We showcase the usefulness of the proposed method, its eq…
We establish sharp Sobolev inequalities of order four on Euclidean d-balls for d greater than or equal to four. When d=4, our inequality generalizes the classical second order Lebedev-Milin inequality on Euclidean 2-balls. Our method relies on the use of scattering theory on hyperbolic d-balls. As an application, we ch…
Multi-Task Learning (MTL) can enhance a classifier's generalization performance by learning multiple related tasks simultaneously. Conventional MTL works under the offline or batch setting, and suffers from expensive training cost and poor scalability. To address such inefficiency issues, online learning techniques hav…
Polylab is a MATLAB toolbox for multivariate polynomial modeling.
Paper introduces VDE, a variance-reduced determinant estimator.
HCLM framework uses entropy regularization for open learning systems.
New algorithm for online portfolio selection with reduced runtime.
The curvature of the noncommutative torus ( irrational) endowed with a noncommutative conformal metric has been the focus of attention of several recent works. Continuing the approach taken in the paper [A. Connes and H. Moscovici, http://arxiv.org/abs/1110.3500] we extend the study of the curvature to twist…
We show that standard ResNet architectures can be made invertible, allowing the same model to be used for classification, density estimation, and generation. Typically, enforcing invertibility requires partitioning dimensions or restricting network architectures. In contrast, our approach only requires adding a simple …
Gaussian processes (GPs) with derivatives are useful in many applications, including Bayesian optimization, implicit surface reconstruction, and terrain reconstruction. Fitting a GP to function values and derivatives at points in dimensions requires linear solves and log determinants with an ${n(d+1) \times n(d…
In this paper, we introduce new classes of divergences by extending the definitions of the Bregman divergence and the skew Jensen divergence. These new divergence classes (g-Bregman divergence and skew g-Jensen divergence) satisfy some properties similar to the Bregman or skew Jensen divergence. We show these g-diverge…
The L1-regularized Gaussian maximum likelihood estimator (MLE) has been shown to have strong statistical guarantees in recovering a sparse inverse covariance matrix, or alternatively the underlying graph structure of a Gaussian Markov Random Field, from very limited samples. We propose a novel algorithm for solving the…
Divergence functions play a key role as to measure the discrepancy between two points in the field of machine learning, statistics and signal processing. Well-known divergences are the Bregman divergences, the Jensen divergences and the f-divergences. In this paper, we show that the symmetric Bregman divergence can be …
A new coordinate system for SPD matrices simplifies computations and generative modeling.
Study explores relationship between Hölder and FDPD divergences.
New algorithm speeds up cluster-based compressive sensing tasks.
While normalizing flows have led to significant advances in modeling high-dimensional continuous distributions, their applicability to discrete distributions remains unknown. In this paper, we show that flows can in fact be extended to discrete events---and under a simple change-of-variables formula not requiring log-d…
Unified representation of density-power-based divergences simplifies estimation to M-estimation.
This paper improves active learning by using robust divergences for committee disagreement.
Gradient descent on LSE objectives implicitly performs EM, leading to collapse without volume control.
New divergence measures improve KL approximation.
The paper improves semi-supervised learning using -divergences and -Rényi divergences.
Low-rank matrix is desired in many machine learning and computer vision problems. Most of the recent studies use the nuclear norm as a convex surrogate of the rank operator. However, all singular values are simply added together by the nuclear norm, and thus the rank may not be well approximated in practical problems. …