Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

91183274365 · Jun 202019922001200920172026
48 results for matrix β-mean

A new method for distributed PCA using matrix β-mean.

problem Efficiently aggregating PCA results across multiple machines with reduced computational overhead.
method Proposes a novel DPCA method that incorporates eigenvalue information using the matrix β-mean.
result The matrix β-mean method improves robustness and stability of eigenvector ordering.

We show that the objective function of conventional k-means clustering can be expressed as the Frobenius norm of the difference of a data matrix and a low rank approximation of that data matrix. In short, we show that k-means clustering is a matrix factorization problem. These notes are meant as a reference and intende…

2015-12-23abs ↗pdf ↗

Conventional machine learning algorithms cannot be applied until a data matrix is available to process. When the data matrix needs to be obtained from a relational database via a feature extraction query, the computation cost can be prohibitive, as the data matrix may be (much) larger than the total input relation size…

2019-10-11abs ↗pdf ↗

Novel mean estimation method under user-level differential privacy reduces noise in continual mean estimates.

problem Maintaining accurate running mean estimates under user-level differential privacy.
method Developed a novel mean estimation specific factorization under approximate differential privacy.
result Achieved asymptotically lower mean-squared error bounds in continual mean estimation.

A clustering algorithm uses the left Gram matrix for high dimensional data.

problem Clustering high dimensional data with many features and few objects.
method The algorithm uses the normalized left Gram matrix G = XX'/P to cluster objects based on row means.
result The algorithm provides the most accurate cluster configuration more than twice as often as competitors.

Kernel clustering algorithm improved for large datasets using incomplete Cholesky factorization.

problem Large memory usage in kernel-based clustering for large-scale datasets.
method Approximate the kernel matrix using incomplete Cholesky factorization and apply linear kk-means clustering.
result The proposed method achieves similar performance to kernel kk-means clustering but handles large-scale datasets efficiently.

Study of metrics on positive-definite matrices from power potential, linking to power means.

problem Understanding metrics on positive-definite matrices derived from power potential.
method Explicit expressions for geodesics and distance function derived from Hessian of power potential.
result Geodesics and distance function converge to weighted matrix geometric mean as β tends to zero.

Paper proposes a method to break symmetries in Bayesian matrix factorization.

problem Symmetries in posterior distribution reduce MCMC sampling efficiency.
method Modification to Gaussian prior mean and covariance to break symmetries.
result Breaking symmetries leads to lower autocorrelation and reconstruction errors.

Efficiently reduces rank of non-negative matrices with quadratic time complexity.

problem Efficiently reducing the rank of non-negative matrices.
method Formulated rank reduction as a mean-field approximation using a log-linear model.
result Optimal solution for minimizing KL divergence can be computed in closed form.

The original k-means clustering method works only if the exact vectors representing the data points are known. Therefore calculating the distances from the centroids needs vector operations, since the average of abstract data points is undefined. Existing algorithms can be extended for those cases when the sole input i…

2013-03-24abs ↗pdf ↗

In this work, the possibility of clustering correlated random variables was examined, both because of their mutual similarity and because of their similarity to the principal components. The k-means algorithm and spectral algorithms were used for clustering. For spectral methods, the similarity matrix was both the matr…

2019-09-07abs ↗pdf ↗

Binary data matrices can represent many types of data such as social networks, votes, or gene expression. In some cases, the analysis of binary matrices can be tackled with nonnegative matrix factorization (NMF), where the observed data matrix is approximated by the product of two smaller nonnegative matrices. In this …

2018-12-17abs ↗pdf ↗

Paper analyzes VI for location-scale families, proving robustness guarantees for mean and correlation recovery.

problem Misspecification in VI for intractable target densities.
method Variational inference on location-scale families with symmetries.
result VI recovers mean and correlation matrix under specific symmetries.

Method selects number of communities in weighted networks.

problem Selecting the number of communities in weighted networks.
method Proposes a novel weighted DCSBM and uses a sequential testing framework with spectral clustering and matrix scaling.
result Method is consistent in estimating the true number of communities under mild conditions.

The paper computes an approximation to the sample Frechet mean of graph sets using spectral information.

problem Characterizing the location of a set of graphs in a metric space.
method The Frechet mean is computed for sets of large graphs using the pseudometric defined by the norm between eigenvalues of adjacency matrices.
result An algorithm to approximate the sample Frechet mean of undirected unweighted graphs is described.

The asymptotic concentration of the Fr{é}chet mean of IID random variables on a Rieman-nian manifold was established with a central limit theorem by Bhattacharya \& Patrangenaru (BP-CLT) [6]. This asymptotic result shows that the Fr{é}chet mean behaves almost as the usual Euclidean case for sufficiently concentrated di…

2019-06-18abs ↗pdf ↗

Signed graphs encode positive (attractive) and negative (repulsive) relations between nodes. We extend spectral clustering to signed graphs via the one-parameter family of Signed Power Mean Laplacians, defined as the matrix power mean of normalized standard and signless Laplacians of positive and negative edges. We pro…

2019-05-15abs ↗pdf ↗

Multilayer graphs encode different kind of interactions between the same set of entities. When one wants to cluster such a multilayer graph, the natural question arises how one should merge the information different layers. We introduce in this paper a one-parameter family of matrix power means for merging the Laplacia…

2018-03-01abs ↗pdf ↗

In this paper, the Riemannian gradient algorithm and the natural gradient algorithm are applied to solve descent direction problems on the manifold of positive definite Hermitian matrices, where the geodesic distance is considered as the cost function. The first proposed problem is control for positive definite Hermiti…

2019-04-05abs ↗pdf ↗

Most recent results in matrix completion assume that the matrix under consideration is low-rank or that the columns are in a union of low-rank subspaces. In real-world settings, however, the linear structure underlying these models is distorted by a (typically unknown) nonlinear transformation. This paper addresses the…

2015-12-29abs ↗pdf ↗

The paper addresses portfolio allocation with uncertain covariance matrices, finding a logarithmic risk dependence.

problem Portfolio allocation with uncertain covariance matrices.
method Calculates the expected value of CARA utility function over a distribution of covariance matrices, considering uncertainty in future returns and covariances.
result Marginalization introduces a logarithmic dependence on risk, leading to lower allocation levels for higher uncertainties.

In this paper, we consider the streaming memory-limited matrix completion problem when the observed entries are noisy versions of a small random fraction of the original entries. We are interested in scenarios where the matrix size is very large so the matrix is very hard to store and manipulate. Here, columns of the o…

2015-04-13abs ↗pdf ↗

We study the stability vis a vis adversarial noise of matrix factorization algorithm for matrix completion. In particular, our results include: (I) we bound the gap between the solution matrix of the factorization method and the ground truth in terms of root mean square error; (II) we treat the matrix factorization as …

2012-06-18abs ↗pdf ↗

Proposes a new K-means method for efficient clustering of nonlinear data.

problem Challenges of kernel K-means, including high memory usage and computational inefficiency.
method Combines linear and nonlinear approaches using explicit feature maps based on spectral analysis.
result Demonstrates Explicit Kernel Minkowski Weighted K-means (Explicit KMWK-means) reduces memory usage and improves efficiency.

We generalize the recently discovered relationship between JT gravity and double-scaled random matrix theory to the case that the boundary theory may have time-reversal symmetry and may have fermions with or without supersymmetry. The matching between variants of JT gravity and matrix ensembles depends on the assumed s…

2019-07-07abs ↗pdf ↗

Enhances ROM simulation for multivariate systems with exact Kollo skewness.

problem Modeling multivariate systems with high dimensions and specific higher moments.
method Extends Random Orthogonal Matrix simulation to match target Kollo skewness.
result Established conditions and developed a general approach for constructing admissible values.

K-means -- and the celebrated Lloyd algorithm -- is more than the clustering method it was originally designed to be. It has indeed proven pivotal to help increase the speed of many machine learning and data analysis techniques such as indexing, nearest-neighbor search and prediction, data compression; its beneficial u…

2019-08-23abs ↗pdf ↗