Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We consider practical data characteristics underlying federated learning, where unbalanced and non-i.i.d. data from clients have a block-cyclic structure: each cycle contains several blocks, and each client's training data follow block-specific and non-i.i.d. distributions. Such a data structure would introduce client …
Machine learning predicts Atlantic blocking using limited data.
This work is devoted to elaboration on the idea to use block term decomposition for group data analysis and to raise the possibility of modelling group activity with (Lr, 1) and Tucker blocks. A new generalization of block tensor decomposition was considered in application to group data analysis. Suggested approach was…
The tree reconstruction problem is to collect and analyze massive data at the th level of the tree, to identify whether there is non-vanishing information of the root, as goes to infinity. Its connection to the clustering problem in the setting of the stochastic block model, which has wide applications in machin…
BSTabDiff: Block-Subunit Diffusion Priors for HDLSS Tabular Data Generation
This paper presents a novel Block Iterative Bayesian Algorithm (Block-IBA) for reconstructing block-sparse signals with unknown block structures. Unlike the existing algorithms for block sparse signal recovery which assume the cluster structure of the nonzero elements of the unknown signal to be independent and identic…
Proposes a Latent Block Model for analyzing missing data.
New model for community detection with side information improves recovery accuracy.
We generalize the stochastic block model to the important case in which edges are annotated with weights drawn from an exponential family distribution. This generalization introduces several technical difficulties for model estimation, which we solve using a Bayesian approach. We introduce a variational algorithm that …
Study community detection in multi-view data with various types of information.
Extends co-clustering to mixed numerical and binary data.
This paper speeds up Gaussian process regression for autocorrelated data.
Method estimates number of clusters in Block Markov Chain trajectories.
The proliferation of models for networks raises challenging problems of model selection: the data are sparse and globally dependent, and models are typically high-dimensional and have large numbers of latent variables. Together, these issues mean that the usual model-selection criteria do not work properly for networks…
We propose a mixture of latent trait models with common slope parameters (MCLT) for model-based clustering of high-dimensional binary data, a data type for which few established methods exist. Recent work on clustering of binary data, based on a -dimensional Gaussian latent variable, is extended by incorporating com…
This note corrects a mistake in the paper "consistent cross-validatory model-selection for dependent data: -block cross-validation" by Racine (2000). In his paper, he implied that the therein proposed -block cross-validation is consistent in the sense of Shao (1993). To get this intuition, he relied on the spec…
Latent block models are used for probabilistic biclustering, which is shown to be an effective method for analyzing various relational data sets. However, there has been no statistical test method for determining the row and column cluster numbers of latent block models. Recent studies have constructed statistical-test…
Graphical lasso models ASR utterance dependencies for consistent WER estimation.
The paper uses deep learning to detect financial market regimes from correlation matrices.
A novel MCMC method clusters data faster and more accurately.
A main task in data analysis is to organize data points into coherent groups or clusters. The stochastic block model is a probabilistic model for the cluster structure. This model prescribes different probabilities for the presence of edges within a cluster and between different clusters. We assume that the cluster ass…
Proposes a method to derive knowledge graphs from EHR data.
Recurrent Neural Networks (RNNs) are used in state-of-the-art models in domains such as speech recognition, machine translation, and language modelling. Sparsity is a technique to reduce compute and memory requirements of deep learning models. Sparse RNNs are easier to deploy on devices and high-end server processors. …
A new method selects features for clustering using a block model.
Deep learning has shown promising results on many machine learning tasks but DL models are often complex networks with large number of neurons and layers, and recently, complex layer structures known as building blocks. Finding the best deep model requires a combination of finding both the right architecture and the co…
The excellent performance of representation learning of autoencoders have attracted considerable interest in various applications. However, the structure and multi-local collaborative relationships of unlabeled data are ignored in their encoding procedure that limits the capability of feature extraction. This paper pre…
This paper introduces GLT for better input data representation in BNN and proposes a compact topology with block pruning.
Paper proposes ABDR for convex subspace clustering with adaptive block diagonal representation.
We consider the problem of identifying multiway block structure from a large noisy tensor. Such problems arise frequently in applications such as genomics, recommendation system, topic modeling, and sensor network localization. We propose a tensor block model, develop a unified least-square estimation, and obtain the t…
Spectral clustering achieves strong consistency in the stochastic block model under certain conditions.
Spectral embedding is a popular technique for the representation of graph data. Several regularization techniques have been proposed to improve the quality of the embedding with respect to downstream tasks like clustering. In this paper, we explain on a simple block model the impact of the complete graph regularization…
There exist various types of network block models such as the Stochastic Block Model (SBM), the Degree Corrected Block Model (DCBM), and the Popularity Adjusted Block Model (PABM). While this leads to a variety of choices, the block models do not have a nested structure. In addition, there is a substantial jump in the …
Many statistical methods for network data parameterize the edge-probability by attributing latent traits to the vertices such as block structure and assume exchangeability in the sense of the Aldous-Hoover representation theorem. Empirical studies of networks indicate that many real-world networks have a power-law dist…
Two spectral algorithms for community detection in graphs with covariates are compared.
BlockEcho method improves imputation of block-wise missing data.
Two new methods improve block-sparse signal recovery from noisy data.
The Mallows model, introduced in the seminal paper of Mallows 1957, is one of the most fundamental ranking distribution over the symmetric group . To analyze more complex ranking data, several studies considered the Generalized Mallows model defined by Fligner and Verducci 1986. Despite the significant research in…
New model for clustering graphs with multiple data sources.
We present a method based on the orthogonal symmetric non-negative matrix tri-factorization of the normalized Laplacian matrix for community detection in complex networks. While the exact factorization of a given order may not exist and is NP hard to compute, we obtain an approximate factorization by solving an optimiz…
Improved model for grouping nodes in bipartite networks.
We introduce block-tree graphs as a framework for deriving efficient algorithms on graphical models. We define block-tree graphs as a tree-structured graph where each node is a cluster of nodes such that the clusters in the graph are disjoint. This differs from junction-trees, where two clusters connected by an edge al…
Unified multilinear model for causal factor disentanglement.
Tests if vertices in graphs have the same latent positions.
Paper proposes efficient methods for clustering and signal recovery in high-dimensional data with block structures.
Develops a new model to better predict corporate bond yields.
A central problem in analyzing networks is partitioning them into modules or communities. One of the best tools for this is the stochastic block model, which clusters vertices into blocks with statistically homogeneous pattern of links. Despite its flexibility and popularity, there has been a lack of principled statist…
New algorithm optimally clusters networks with side information.