We propose a new method to model multi-way similarities into hypergraphs for clustering.
problem Clustering real-valued data using hypergraphs with multi-way similarities.
method Formulate multi-way similarities using kernel functions, establish connections to hypergraph cut, and develop a fast spectral clustering algorithm.
result Our method outperforms existing graph and heuristic modeling methods in clustering performance.
Develops algorithms for multi-way similarity clustering in hypergraphs.
problem Challenges of spectral clustering in multi-way similarity settings.
method Hypergraph Spectral Clustering (HSC) and Hypergraph Spectral Clustering with Local Refinement (HSCLR).
result Achieves optimal performance under the weighted stochastic block model.
We investigate the distribution of eigenvalues of the weighted Laplacian on closed weighted Riemannian manifolds of nonnegative Bakry-Émery Ricci curvature. We derive some universal inequalities among eigenvalues of the weighted Laplacian on such manifolds. These inequalities are quantitative versions of the previous t…
Paper reviews multi-way graph signal processing for tensor data.
problem Maximizing use of multi-way structure in irregular tensor data.
method Generalizes GSP to multi-way data, focusing on graph signals across tensor modes.
result Synthesizes common themes in combining GSP with tensor analysis.
A new method for analyzing multi-source, multi-way data reduces dimensionality and reveals shared and individual structures.
problem Analyzing multi-source, multi-way data from different high-throughput technologies.
method Multiple Linked Tensor Factorization (MULTIFAC) extending CP decomposition with L2 penalties and EM algorithm for incomplete data.
result MULTIFAC approximates underlying signal, identifies shared and unshared structures, and imputes missing data.
High-dimensional linear classifiers, such as the support vector machine (SVM) and distance weighted discrimination (DWD), are commonly used in biomedical research to distinguish groups of subjects based on a large number of features. However, their use is limited to applications where a single vector of features is mea…
We extend multi-way, multivariate ANOVA-type analysis to cases where one covariate is the view, with features of each view coming from different, high-dimensional domains. The different views are assumed to be connected by having paired samples; this is a common setup in recent bioinformatics experiments, of which we a…
Bayesian method identifies multi-way interactions among predictors.
problem Identifying meaningful interactions among multiple variables.
method Factorization mechanism and Gibbs sampling for posterior inference.
result Posterior consistency of the regression model.
Bayesian model predicts iron deficiency from multi-source multi-way molecular data.
problem Predicting iron deficiency in rhesus monkeys from multi-source multi-way molecular data.
method Developed a Bayesian approach with a linear model incorporating multi-way dependence and varying signal sizes across sources.
result Model accurately classifies iron deficiency in monkeys and outperforms simpler models.
We propose multi-way, multilingual neural machine translation. The proposed approach enables a single neural translation model to translate between multiple languages, with a number of parameters that grows only linearly with the number of languages. This is made possible by having a single attention mechanism that is …
SiBBlInGS discovers interpretable building blocks across states in multi-way data.
problem Identifying interpretable units (Building Blocks) in multi-state, multi-way data.
method Graph-based dictionary learning approach for sparse BBs and temporal traces.
result Captures per-trial variability and state-specific vs. state-invariant components.
Tensor factorization uncovers hidden patterns in student behavior data.
problem Discovering low-dimensional structure in high-dimensional behavioral data.
method Non-negative tensor factorization applied to wearable sensor data.
result Tensor factorization reveals clusters of students with different behaviors.
Translating potential disease biomarkers between multi-species 'omics' experiments is a new direction in biomedical research. The existing methods are limited to simple experimental setups such as basic healthy-diseased comparisons. Most of these methods also require an a priori matching of the variables (e.g., genes o…
Paper connects tensor regression and Gaussian processes for multi-way data analysis.
problem Learning high-order correlations from multi-way data.
method Demonstrates connections between low-rank tensor regression and Gaussian processes, proving oracle inequality and learning curve.
result Low-rank tensor regression is equivalent to constrained Bayesian inference in Gaussian processes, with learning dependent on eigenvalues and variable correlations.
New method reduces complexity of fuzzy decision trees in Big Data.
problem Reducing complexity in multi-way fuzzy decision trees for Big Data classification.
method Two-step process: 1) Probability integral transform, 2) Ruspini strong fuzzy partition.
result Up to 6 million fewer leaves with similar classification accuracy.
NeSGD efficiently updates tensor-based features for online model learning in multi-way data.
problem Incremental updates of tensor-based features and model coefficients for evolving data distributions.
method Online NeSGD for CP decomposition of tensor data.
result Proposed NeSGD method significantly improves classification accuracy and adapts to changing data distributions.
Proposes a multi-objective variational autoencoder for smart infrastructure damage detection.
problem Detecting and diagnosing damage in smart infrastructure using multi-way data.
method Multi-objective variational autoencoder (MVA) method for smart infrastructure damage detection and diagnosis.
result The method accurately detects structural damage, estimates severity, and captures damage locations.
Develops a tensor network framework to reduce RNN complexity for high-dimensional sequence modeling.
problem Exponential parameter growth in RNNs for large multidimensional data.
method Embeds a multi-linear graph filter in a tensor network architecture to approximate RNN hidden states.
result Demonstrates superior performance and reduced complexity compared to traditional RNNs.
Develops a new tensor classification method for high-dimensional data.
problem Efficient learning algorithms exploiting tensorial structure in high-dimensional multi-way arrays.
method Tensor Train Multi-way Multi-level Kernel (TT-MMK) combining Canonical Polyadic decomposition, Dual Structure-preserving Support Vector Machine, and Tensor Train approximation.
result The TT-MMK method provides higher prediction accuracy and is more reliable computationally compared to other techniques.
KReTTaH uses tensor trains and Hadamard overparameterization for fast, interpretable multi-way data imputation.
problem Multi-way data imputation for high-dimensional functional MRI and dynamic graph recovery.
method Reformulates imputation as RKHS regression with TT-constrained coefficients and Hadamard overparameterization. Optimizes TT coefficients and kernel matrices on Riemannian manifolds.
result Consistently outperforms state-of-the-art methods in modeling accuracy.
KReTTaH uses tensor trains and Hadamard overparameterization for fast, interpretable multi-way data imputation.
problem Multi-way data imputation in high-dimensional spaces.
method Reformulates imputation as RKHS regression with TT-constrained coefficients, optimized on manifold frameworks.
result Consistently outperforms state-of-the-art methods in accuracy.
Deep learning method clusters multi-view data matrices.
problem Clustering heterogeneous relational data matrices.
method Deep collective matrix tri-factorization (DCMTF).
result Discover latent clusters across input matrices and their associations.
We present an approach for penalized tensor decomposition (PTD) that estimates smoothly varying latent factors in multi-way data. This generalizes existing work on sparse tensor decomposition and penalized matrix decompositions, in a manner parallel to the generalized lasso for regression and smoothing problems. Our ap…
Estimates higher-order Poincaré constants for weighted manifolds.
problem Estimating constants for weighted manifolds and their applications.
method Introducing higher-order Poincaré constants and estimating them from above.
result Upper bounds for eigenvalues and isoperimetric constants.
We propose a framework for the linear prediction of a multi-way array (i.e., a tensor) from another multi-way array of arbitrary dimension, using the contracted tensor product. This framework generalizes several existing approaches, including methods to predict a scalar outcome from a tensor, a matrix from a matrix, or…
A multi-way factor analysis model is introduced for tensor-variate data of any order. Each data item is represented as a (sparse) sum of Kruskal decompositions, a Kruskal-factor analysis (KFA). KFA is nonparametric and can infer both the tensor-rank of each dictionary atom and the number of dictionary atoms. The model …
Advances neural tri-factorization for clustering and discordance analysis of multi-typed data.
problem Challenges in analyzing heterogeneous, multimodal relational data.
method Deep collective matrix tri-factorization for spectral clustering and cluster association learning.
result Demonstrates efficacy over previous non-neural approaches in clustering and discordance analysis.
New method detects communities in hypergraphs by embedding them into a vector space.
problem Detecting communities in hypergraphs with multi-way interactions.
method Augmenting non-uniform hypergraphs, embedding into a vector space, using an alternative updating scheme.
result Asymptotic consistencies in community detection and hypergraph estimation established.
High-dimensional tensors or multi-way data are becoming prevalent in areas such as biomedical imaging, chemometrics, networking and bibliometrics. Traditional approaches to finding lower dimensional representations of tensor data include flattening the data and applying matrix factorizations such as principal component…
The problem of Hybrid Linear Modeling (HLM) is to model and segment data using a mixture of affine subspaces. Different strategies have been proposed to solve this problem, however, rigorous analysis justifying their performance is missing. This paper suggests the Theoretical Spectral Curvature Clustering (TSCC) algori…
A new method identifies class-specific covariates in multi-class prediction tasks.
problem Identifying covariates specifically associated with one or more outcome classes in multi-class prediction tasks.
method Introducing multi forests (MuFs) with multi-way and binary splits to measure class-associated discriminatory ability.
result The multi-class VIM specifically ranks class-associated covariates highly, unlike conventional VIMs.
Topological data analysis (TDA) has emerged as one of the most promising techniques to reconstruct the unknown shapes of high-dimensional spaces from observed data samples. TDA, thus, yields key shape descriptors in the form of persistent topological features that can be used for any supervised or unsupervised learning…
Probabilistic PARAFAC2 improves robustness to noise.
problem Improving robustness of PARAFAC2 to noise and determining the number of factors.
method Developed two probabilistic formulations of PARAFAC2 with variational procedures for inference.
result Probabilistic PARAFAC2 is more robust to noise and model order misspecification.
Extracting latent low-dimensional structure from high-dimensional data is of paramount importance in timely inference tasks encountered with `Big Data' analytics. However, increasingly noisy, heterogeneous, and incomplete datasets as well as the need for {\em real-time} processing of streaming data pose major challenge…
GRTR framework uses graph regularization to improve financial forecasting.
problem High computational costs and economic domain knowledge loss in tensor models.
method Graph-Regularized Tensor Regression (GRTR) framework incorporating economic domain knowledge.
result Improved performance in multi-way financial forecasting with reduced computational costs.
Tensors or {\em multi-way arrays} are functions of three or more indices (i,j,k,⋯) -- similar to matrices (two-way arrays), which are functions of two indices (r,c) for (row,column). Tensors have a rich history, stretching over almost a century, and touching upon numerous disciplines; but they have only recent…
A new method constrains PARAFAC2 for better pattern recovery.
problem Challenges in analyzing multi-way measurements with variations across one mode.
method AO-ADMM approach to fit PARAFAC2 model with flexible constraints.
result The proposed method allows for flexible constraints, recovers patterns accurately, and is computationally efficient.
NeCPD improves online tensor decomposition using SGD with Hessian analysis and NAG.
problem Efficiently decompose multi-way tensors in online data processing.
method NeCPD solver based on SGD with Hessian analysis and NAG.
result NeCPD provides more accurate results than existing methods.
Improved source code summarization using extended Tree-LSTM.
problem Challenges in applying LSTM to structured source code.
method Extended Tree-LSTM for abstract syntax trees (ASTs).
result Multi-way Tree-LSTM achieves better results than state-of-the-art techniques.
The paper explores tensor decompositions in deep learning models.
problem Compressing parameter space and creating richer representations.
method Tensor decompositions applied to deep learning models.
result Tensor methods can yield richer adaptive representations of complex data.
We consider the problem of consistently matching multiple sets of elements to each other, which is a common task in fields such as computer vision. To solve the underlying NP-hard objective, existing methods often relax or approximate it, but end up with unsatisfying empirical performance due to a misaligned objective.…
The Yarowsky algorithm is a rule-based semi-supervised learning algorithm that has been successfully applied to some problems in computational linguistics. The algorithm was not mathematically well understood until (Abney 2004) which analyzed some specific variants of the algorithm, and also proposed some new algorithm…
A new tree method for tensor data improves regression accuracy.
problem Efficiently modeling tensor data for regression problems.
method Scalar-output regression tree models for scalar-on-tensor problems, and tensor-on-tensor problems using additive tree ensemble approaches.
result The tensor-input tree (TT) method outperforms tensor-input GP models in efficiency and accuracy.
Bayesian model identifies outliers and determines tensor rank in streaming data.
problem Outliers and over-fitting in streaming tensor factorization.
method Variational Bayesian Inference for robust tensor rank determination and outlier identification.
result Model accurately identifies sparse outliers and determines tensor rank.
TRNN combines tensor geometry with neural network nonlinearity for HD data.
problem Modeling high-dimensional data with preserved tensor geometry and nonlinear interactions.
method Introduces TRNN that integrates tensor geometry and neural network nonlinearity.
result TRNN preserves tensor geometry while offering nonlinearity.
A new algorithm speeds up CP decomposition for large tensors.
problem Efficiently processing large-scale tensors in real-time.
method Randomized online CP decomposition (ROCP) algorithm.
result ROCP reduces computing time and memory usage significantly.
Spatially constrained Gaussian mixture models reduce covariance complexity.
problem High dimensionality in finite mixture models for spatial data.
method Spatial covariance constraint with only four free parameters.
result Improves clustering of multi-way spatial data and inference of spatial patterns.
New hypergraph method improves scRNA-seq clustering.
problem Loss of higher-order information and overestimation in coexpression networks.
method Conceptualizing scRNA-seq data as hypergraphs and proposing novel clustering methods.
result Proposed methods outperform existing methods on simulated and real datasets.