Deviance-style normalization for sparse, jointly overdispersed count matrices
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This work compresses heavy-tailed weight matrices for tighter generalization bounds.
A gamma process dynamic Poisson factor analysis model is proposed to factorize a dynamic count matrix, whose columns are sequentially observed count vectors. The model builds a novel Markov chain that sends the latent gamma random variables at time as the shape parameters of those at time , which are linked …
CMF is a technique for simultaneously learning low-rank representations based on a collection of matrices with shared entities. A typical example is the joint modeling of user-item, item-property, and user-feature matrices in a recommender system. The key idea in CMF is that the embeddings are shared across the matrice…
Efficiently infers sparse networks from count data with reduced memory usage.
In banking practice, rating transition matrices have become the standard approach of deriving multi-year probabilities of default (PDs) from one-year PDs, the latter normally being available from Basel ratings. Rating transition matrices have gained in importance with the newly adopted IFRS 9 accounting standard. Here,…
Based on a new atomic norm, we propose a new convex formulation for sparse matrix factorization problems in which the number of nonzero elements of the factors is assumed fixed and known. The formulation counts sparse PCA with multiple factors, subspace clustering and low-rank sparse bilinear regression as potential ap…
We define a family of probability distributions for random count matrices with a potentially unbounded number of rows and columns. The three distributions we consider are derived from the gamma-Poisson, gamma-negative binomial, and beta-negative binomial processes. Because the models lead to closed-form Gibbs sampling …
We describe a way of representing finite biquandles with n elements as 2n x 2n block matrices. Any finite biquandle defines an invariant of virtual knots through counting homomorphisms. The counting invariants of non-quandle biquandles can reveal information not present in the knot quandle, such as the non-triviality o…
Proposes PSCCA for estimating correlations and canonical correlations in sparse count data.
Spectral clustering with edge counting detects communities in sparse models.
p-SNE embeds Poisson count data into low dimensions preserving structure.
BeBold improves exploration in sparse-reward tasks by regulating visitation counts.
Method estimates number of clusters in Block Markov Chain trajectories.
Ranky solves SVD for large sparse matrices in distributed systems.
We tackle anomaly detection in sparse time series data.
Sparse matrices are favorable objects in machine learning and optimization. When such matrices are used, in place of dense ones, the overall complexity requirements in optimization can be significantly reduced in practice, both in terms of space and run-time. Prompted by this observation, we study a convex optimization…
PHIBP predicts infectious disease outbreaks in sparse data regions.
We present a Bayesian tensor factorization model for inferring latent group structures from dynamic pairwise interaction patterns. For decades, political scientists have collected and analyzed records of the form "country took action toward country at time "---known as dyadic events---in order to form an…
Novel Bayesian method for high-dimensional count data prediction.
Paper proposes VAE-BPTF for better tensor factorization of sparse, imbalanced count data.
Develops a model for RNA-seq data clustering.
Random projections help in representing sparse graphs efficiently.
Pruning at initialization fails to find sparse subnetworks, revealing information-theoretic barriers.
New bootstraps improve speed and accuracy for graph count functionals.
This paper proposes learning to jump for generative modeling of sparse, skewed, heavy-tailed data.
Algorithm counts intersections of normal curves efficiently.
A framework estimates multiple precision matrices with shared structures.
Given two data matrices and , sparse canonical correlation analysis (SCCA) is to seek two sparse canonical vectors and to maximize the correlation between and . However, classical and sparse CCA models consider the contribution of all the samples of data matrices and thus cannot identify an unde…
Proposes a method to handle sparse multiway count data with false zeros using zero-truncated Poisson regression.
Origin-destination (OD) matrices are often used in urban planning, where a city is partitioned into regions and an element (i, j) in an OD matrix records the cost (e.g., travel time, fuel consumption, or travel speed) from region i to region j. In this paper, we partition a day into multiple intervals, e.g., 96 15-min …
RGAM builds more accurate models by preferring linear features over non-linear ones.
ENTED efficiently decomposes binary and count tensors using nonparametric Gaussian processes.
Improved COD algorithm reduces streaming AMM errors and uses less space.
Improves classification of microbiome data using mixture distributions.
Many applications of machine learning involve the analysis of large data frames-matrices collecting heterogeneous measurements (binary, numerical, counts, etc.) across samples-with missing values. Low-rank models, as studied by Udell et al. [30], are popular in this framework for tasks such as visualization, clustering…
We present a theory for Euclidean dimensionality reduction with subgaussian matrices which unifies several restricted isometry property and Johnson-Lindenstrauss type results obtained earlier for specific data sets. In particular, we recover and, in several cases, improve results for sets of sparse and structured spars…
Multiresolution Matrix Factorization (MMF) was recently introduced as a method for finding multiscale structure and defining wavelets on graphs/matrices. In this paper we derive pMMF, a parallel algorithm for computing the MMF factorization. Empirically, the running time of pMMF scales linearly in the dimension for spa…
Method estimates sparse inverse covariance and partial correlation matrices efficiently.
Geodesic distance matrices can reveal shape properties that are largely invariant to non-rigid deformations, and thus are often used to analyze and represent 3-D shapes. However, these matrices grow quadratically with the number of points. Thus for large point sets it is common to use a low-rank approximation to the di…
Regular integer lattices are characterized by k unit vectors that build up their generator matrices. These have rank k for D-lattices, and are rank-deficient for A-lattices, for E_6 and E_7. We count lattice points inside hypercubes centered at the origin for all three types, as if classified by maximum infinity norm i…
Algorithm matches vertices of correlated Erdős-Rényi graphs efficiently.
New method estimates sparse covariance matrices in logit mixtures.
Paper proposes Monarch matrices for scalable probabilistic circuits.
Given the superposition of a low-rank matrix plus the product of a known fat compression matrix times a sparse matrix, the goal of this paper is to establish deterministic conditions under which exact recovery of the low-rank and sparse components becomes possible. This fundamental identifiability issue arises with tra…
New matrix reveals cluster info in sparse directed graphs.
In this letter, we propose an algorithm for recovery of sparse and low rank components of matrices using an iterative method with adaptive thresholding. In each iteration, the low rank and sparse components are obtained using a thresholding operator. This algorithm is fast and can be implemented easily. We compare it w…
New algorithms reduce communication in GNN training.