The study uses a ReLU network to discern geometric structure in data via the Data Information Matrix.
problem Understanding the geometric structure of real data in high-dimensional spaces.
method Employing a ReLU neural network trained as a classifier and the Data Information Matrix (DIM) to discern a singular foliation structure.
result The singular points of the foliation are measure zero, and a local regular foliation exists almost everywhere.
MFAI uses gradient boosted trees to leverage auxiliary info for scalable Bayesian matrix factorization.
problem Matrix factorization struggles with poor data quality, especially high sparsity and low SNR.
method Integrates gradient boosted trees into probabilistic matrix factorization framework.
result MFAI effectively leverages auxiliary information, improving model performance.
New matrix reveals cluster info in sparse directed graphs.
problem Analyzing cluster information in directed graphs.
method Proposed complex non-backtracking matrix integrating Hermitian adjacency matrix and non-backtracking matrix properties.
result The complex non-backtracking matrix holds cluster information, especially for sparse directed graphs.
Transposable data represents interactions among two sets of entities, and are typically represented as a matrix containing the known interaction values. Additional side information may consist of feature vectors specific to entities corresponding to the rows and/or columns of such a matrix. Further information may also…
Framework for robust matrix estimation with side information.
problem High-dimensional matrix estimation with restrictive structure.
method Flexible framework decomposing matrix into four components: interaction, row, column, and residual.
result Improved imputation accuracy and treatment-effect estimation with side information.
We extend kernelized matrix factorization with a fully Bayesian treatment and with an ability to work with multiple side information sources expressed as different kernels. Kernel functions have been introduced to matrix factorization to integrate side information about the rows and columns (e.g., objects and users in …
The matrix-based Renyi's α-order entropy functional was recently introduced using the normalized eigenspectrum of a Hermitian matrix of the projected data in a reproducing kernel Hilbert space (RKHS). However, the current theory in the matrix-based Renyi's α-order entropy functional only defines the entropy of a single…
Predicting unobserved entries of a partially observed matrix has found wide applicability in several areas, such as recommender systems, computational biology, and computer vision. Many scalable methods with rigorous theoretical guarantees have been developed for algorithms where the matrix is factored into low-rank co…
Given a sparse rating matrix and an auxiliary matrix of users or items, how can we accurately predict missing ratings considering different data contexts of entities? Many previous studies proved that utilizing the additional information with rating data is helpful to improve the performance. However, existing methods …
The nonnegative matrix factorization is a widely used, flexible matrix decomposition, finding applications in biology, image and signal processing and information retrieval, among other areas. Here we present a related matrix factorization. A multi-objective optimization problem finds conical combinations of templates …
This work formulates a novel song recommender system as a matrix completion problem that benefits from collaborative filtering through Non-negative Matrix Factorization (NMF) and content-based filtering via total variation (TV) on graphs. The graphs encode both playlist proximity information and song similarity, using …
Proposes a transductive matrix completion method with calibration for multi-task learning.
problem Improving multi-task learning with multiple related data sources.
method Transductive matrix completion with calibration constraint.
result The proposed algorithm recovers incomplete feature and target matrices with improved results.
In this short note we extend some of the recent results on matrix completion under the assumption that the columns of the matrix can be grouped (clustered) into subspaces (not necessarily disjoint or independent). This model deviates from the typical assumption prevalent in the literature dealing with compression and r…
We introduce RSE to measure robustness in estimation problems.
problem Estimating statistical models from observed data.
method Developed theory for spectral functions of measures to compute RSE.
result RSE reveals a reciprocal relationship with problem complexity.
Real-world data such as digital images, MRI scans and electroencephalography signals are naturally represented as matrices with structural information. Most existing classifiers aim to capture these structures by regularizing the regression matrix to be low-rank or sparse. Some other methodologies introduce factorizati…
Paper improves ML estimation from incomplete data with robust M-estimator.
problem Estimating parameters from incomplete data with improved accuracy.
method Developed a robust M-estimator and a sandwich estimator for standard errors.
result Improved estimation accuracy with smaller standard errors than ML estimates.
Paper proposes a transfer learning method for improving matrix completion.
problem Improving estimation of a low-rank target matrix using auxiliary data.
method Transfer learning procedure leveraging prior information on favorable source datasets.
result Method outperforms traditional methods when source datasets are close to the target matrix.
A new method for state estimation in state-space models using incomplete data.
problem State estimation in nonlinear state-space models with incomplete observations.
method Statistical analysis of incomplete observations, score function, observed information matrices, EM-gradient-particle filtering.
result Maximum likelihood estimation of state-vector with explicit form of observed information matrix.
Proposes TS-NMF for 2D clustering, preserving spatial info.
problem Loss of spatial information in 2D data.
method Semi-Nonnegative Matrix Factorization with manifold learning.
result Improves clustering performance compared to state-of-the-art.
Paper finds formulas for mutual information and MMSE in matrix tensor product problems.
problem High-dimensional inference problems involving matrix tensor products.
method Single-letter formulas for mutual information and MMSE, using new techniques.
result Analytical formulas describe leading order terms in mutual information and MMSE.
A new multi-view clustering method using deep matrix decomposition and partition alignment.
problem Improving multi-view clustering methods to better utilize data representations and view-specific structures.
method Deep matrix decomposition for partition representations, joint use of partition representations, and alternating optimization.
result Demonstrated effectiveness on six benchmark datasets compared to state-of-the-art methods.
In matrix factorization, available graph side-information may not be well suited for the matrix completion problem, having edges that disagree with the latent-feature relations learnt from the incomplete data matrix. We show that removing these contested edges improves prediction accuracy and scalability. We…
Paper improves compressed sensing with prior probability information.
problem Enhancing compressed sensing accuracy with prior information.
method Designing a sensing matrix and sparse recovery algorithm using probability-based prior information.
result Proposed methods outperform existing CS systems in simulations.
This paper improves sample efficiency in noisy inductive matrix completion with side-information.
problem Improving sample efficiency in noisy inductive matrix completion with side-information.
method Nonconvex projected gradient descent algorithm with spectral initialization.
result Achieves linear convergence and stable recovery at a sample complexity governed by the effective side-information dimension.
Proposes a privacy-preserving recommendation system using matrix factorization and differential privacy.
problem Privacy leakage in recommendation systems when anonymizing user data is not sufficient.
method Uses matrix factorization and differential privacy via the Gaussian mechanism.
result Demonstrates excellent utility for privacy-preserving recommendation systems.
Transfer knowledge from multiple sources to improve matrix completion.
problem Matrix completion with noisy data.
method Aggregating singular subspaces information from multiple sources to solve a two-way PCA problem and transform into a low-dimensional linear regression.
result Guaranteed statistical efficiency in transforming the high-dimensional target matrix completion problem.
Optimal transfer learning for missing not-at-random matrix completion using source data.
problem Matrix completion in a Missing Not-at-Random setting with incomplete and noisy source data.
method Active sampling of rows and columns, feature shift in latent space, minimax lower bounds, computationally efficient estimation framework.
result Achieves minimax lower bound for active sampling setting, avoiding incoherence assumptions.
This paper analyzes Barlow Twins' representation efficiency using information-geometric methods.
problem Understanding and comparing the efficiency of self-supervised learning methods.
method Introduces an information-geometric framework to quantify representation efficiency and applies it to Barlow Twins.
result Proves that Barlow Twins achieves optimal representation efficiency (η=1).
New method clusters matrix-valued data by latent variables.
problem Clustering matrix-valued data with hidden structure.
method Latent variable model with hierarchical clustering.
result Algorithm attains clustering consistency in high dimensions.
A fast optimization method for matrix completion with side information.
problem Matrix completion with and without side information.
method fastImpute based on non-convex gradient descent.
result fastImpute converges to a global minimum and recovers the matrix accurately.
Paper improves matrix-valued data classification using nonparametric LDA.
problem Classification of matrix-valued data in neuroimaging and signal processing.
method Nonparametric LDA based on NPMLE for vectorized and scaled matrices.
result Improves classification performance across various data structures.
Study confirms sparse coding in whole brain using MRI data.
problem Sparse coding in the whole brain's neural activities.
method Applied various matrix factorization methods to fMRI data.
result Sparse coding hypothesis in information representation in the whole human brain is confirmed.
A deep neural network is a hierarchical nonlinear model transforming input signals to output signals. Its input-output relation is considered to be stochastic, being described for a given input by a parameterized conditional probability distribution of outputs. The space of parameters consisting of weights and biases i…
New algorithm predicts missing matrix entries using side information, outperforming existing methods.
problem Learning a partially observed matrix with side information.
method Mixed-projection ADMM algorithm for optimization.
result Our algorithm achieves 2.3% lower objective value and 41% lower reconstruction error than benchmarks.
Proposes a robust factor analysis for matrix data.
problem Robust factor analysis for matrix data with heavy-tailed or contaminated data.
method Bilinear factor analysis based on the matrix-variate t distribution. result Significantly higher breakdown point than traditional methods.
A new method preserves useful information in data rows with outlying cells.
problem Preserving useful information in data rows with outlying cells.
method Cellwise robust Minimum Covariance Determinant (cellMCD) method using observed likelihood and a penalty term on cellwise outliers.
result The cellMCD method performs well in simulations and on real data.
Nonnegative Matrix Factorization (NMF) has been continuously evolving in several areas like pattern recognition and information retrieval methods. It factorizes a matrix into a product of 2 low-rank non-negative matrices that will define parts-based, and linear representation of nonnegative data. Recently, Graph regula…
Algorithm refines matrix ratings using hierarchical graph clustering.
problem Matrix completion with side information from social graphs.
method Hierarchical graph clustering followed by iterative refinement of matrix ratings.
result Achieves optimal sample complexity for matrix completion.
Relational learning can be used to augment one data source with other correlated sources of information, to improve predictive accuracy. We frame a large class of relational learning problems as matrix factorization problems, and propose a hierarchical Bayesian model. Training our Bayesian model using random-walk Metro…
Introduces a new model for mapping matrices to matrices, subsuming linear regression.
problem Learning matrix-to-matrix mappings from data.
method Partial trace regression model, leveraging quantum information theory.
result Relevance demonstrated in matrix-to-matrix regression and positive semidefinite matrix completion.
DPERC efficiently estimates covariance matrices for mixed data with missing values.
problem Estimating covariance matrices for datasets with missing values and mixed features.
method Direct Parameter Estimation for Randomly Missing Data with Categorical Features (DPERC).
result DPERC outperforms other methods in estimating covariance matrices for mixed data with missing values.
This work improves knowledge distillation by transferring full kernel matrices efficiently.
problem Efficiently transferring full pairwise similarity matrices for model compression in deep learning.
method The authors propose a method to transfer the full similarity matrix effectively using the Nyström method, decomposing it into partial matrices.
result The difference between the full kernel matrices of teacher and student can be well bounded by partial matrices, improving optimization efficiency.
The matrix-based Renyi's α-entropy functional and its multivariate extension were recently developed in terms of the normalized eigenspectrum of a Hermitian matrix of the projected data in a reproducing kernel Hilbert space (RKHS). However, the utility and possible applications of these new estimators are rather new an…
Boolean matrix has been used to represent digital information in many fields, including bank transaction, crime records, natural language processing, protein-protein interaction, etc. Boolean matrix factorization (BMF) aims to find an approximation of a binary matrix as the Boolean product of two low rank Boolean matri…
We show that the Kullback-Leibler distance is a good measure of the statistical uncertainty of correlation matrices estimated by using a finite set of data. For correlation matrices of multivariate Gaussian variables we analytically determine the expected values of the Kullback-Leibler distance of a sample correlation …
SMM preserves matrix data structure for SVM classification.
problem Preserving spatial correlations in matrix data for SVM.
method SMM uses spectral elastic net combining nuclear and Frobenius norms.
result SMM improves SVM performance on matrix data.
Deep neural networks reveal a low-dimensional manifold structure in data.
problem Understanding the structure of data for better model performance.
method Model-centric analysis of the data manifold using the local data matrix and Fisher information matrix.
result The dataset lies on a data leaf with a dimension bounded by the number of labels.
Efficiently factorizes coupled matrix tensor data for better accuracy and speed.
problem Poor computation efficiency in existing N-CMTF algorithms.
method Column-wise element selection to prevent frequent gradient updates.
result More accurate and computationally efficient factorization.