Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

1223 · Sep 202019922001200920182026
48 results for Biclustering

New test determines appropriate number of biclusters in relational data.

problem Determining the correct number of biclusters in relational data matrices.
method Proposes a new statistical test that does not require regular-grid assumptions.
result Derives asymptotic behavior of the test statistic for both null and alternative cases.

RIn-Close_CVC2 improves biclustering efficiency by reducing memory usage.

problem Mining maximal biclusters in numerical datasets efficiently and without redundancy.
method Proposes RIn-Close_CVC2, a new version of RIn-Close_CVC that eliminates redundant biclusters without a symbol table.
result RIn-Close_CVC2 reduces memory usage and improves runtime compared to RIn-Close_CVC.

In the biclustering problem, we seek to simultaneously group observations and features. While biclustering has applications in a wide array of domains, ranging from text mining to collaborative filtering, the problem of identifying structure in high dimensional genomic data motivates this work. In this context, biclust…

2014-08-05abs ↗pdf ↗

Being an unsupervised machine learning and data mining technique, biclustering and its multimodal extensions are becoming popular tools for analysing object-attribute data in different domains. Apart from conventional clustering techniques, biclustering is searching for homogeneous groups of objects while keeping their…

2017-02-17abs ↗pdf ↗

Sparse Convex Biclustering improves accuracy and robustness in high-dimensional datasets.

problem Challenges in clustering rows and columns of large-scale datasets due to noise and computational complexity.
method Sparse Convex Biclustering (SpaCoBi) using convex optimization and stability-based tuning.
result Significantly outperforms state-of-the-art methods in accuracy for high-dimensional datasets.

Multimodal clustering is an unsupervised technique for mining interesting patterns in nn-adic binary relations or nn-mode networks. Among different types of such generalized patterns one can find biclusters and formal concepts (maximal bicliques) for 2-mode case, triclusters and triconcepts for 3-mode case, closed $n…

2017-02-27abs ↗pdf ↗

Motivation: Modelling methods that find structure in data are necessary with the current large volumes of genomic data, and there have been various efforts to find subsets of genes exhibiting consistent patterns over subsets of treatments. These biclustering techniques have focused on one data source, often gene expres…

2015-12-29abs ↗pdf ↗

In many situations it is desirable to identify clusters that differ with respect to only a subset of features. Such clusters may represent homogeneous subgroups of patients with a disease, such as cancer or chronic pain. We define a bicluster to be a submatrix U of a larger data matrix X such that the features and obse…

2014-07-11abs ↗pdf ↗

Perceptrons are neuronal devices capable of fully discriminating linearly separable classes. Although straightforward to implement and train, their applicability is usually hindered by non-trivial requirements imposed by real-world classification problems. Therefore, several approaches, such as kernel perceptrons, have…

2016-03-22abs ↗pdf ↗

Unified method for discovering biclusters and triclusters in longitudinal data.

problem High-dimensional, sparsely sampled, irregularly observed longitudinal data.
method Tri-SfSVD, a unified sparse functional Singular Value Decomposition framework.
result Identified localized structures at the subject, subject-feature, and subject-feature-time levels.

The problem of biclustering consists of the simultaneous clustering of rows and columns of a matrix such that each of the submatrices induced by a pair of row and column clusters is as uniform as possible. In this paper we approximate the optimal biclustering by applying one-way clustering algorithms independently on t…

2007-12-17abs ↗pdf ↗

A biclustering algorithm finds dense disjoint subgraphs in weighted bipartite graphs.

problem Finding dense disjoint bicliques in a weighted bipartite graph.
method Semidefinite programming-based branch-and-cut algorithm with upper and lower bounds.
result The algorithm can solve much larger instances than general-purpose solvers.

Biclustering, the process of simultaneously clustering the rows and columns of a data matrix, is a popular and effective tool for finding structure in a high-dimensional dataset. Many biclustering procedures appear to work well in practice, but most do not have associated consistency guarantees. To address this shortco…

2012-06-29abs ↗pdf ↗

EBIC is a biclustering tool for big genomic data, achieving significant speedup.

problem Mining genetic data for high-dimensional and big data challenges.
method EBIC is a biclustering algorithm enhanced for big data, including support for missing values and integration with R.
result EBIC achieves over 6.6 fold speedup on large datasets, demonstrating high scalability.

In this paper we propose two new algorithms based on biclustering analysis, which can be used at the basis of a recommender system for educational orientation of Russian School graduates. The first algorithm was designed to help students make a choice between different university faculties when some of their preference…

2012-02-13abs ↗pdf ↗

Learning the "blocking" structure is a central challenge for high dimensional data (e.g., gene expression data). Recently, a sparse singular value decomposition (SVD) has been used as a biclustering tool to achieve this goal. However, this model ignores the structural information between variables (e.g., gene interacti…

2016-03-19abs ↗pdf ↗

New computational lower bounds for clustering and related problems.

problem Statistical-computational gaps in high-dimensional clustering problems.
method Investigation of low-degree polynomials in latent space models to derive lower bounds.
result New and sharper computational lower bounds for clustering, sparse clustering, and biclustering.

New group-sparse SVD models improve biclustering of gene expression data.

problem Identifying block patterns with similar expressions in high-dimensional gene expression data.
method Proposed GL1-SVD, GL0-SVD, OGL1-SVD, and OGL0-SVD models with group Lasso and L0-norm penalties, using alternating iterative strategies and ADMM.
result Effective in identifying biologically interpretable gene modules with gene prior group knowledge.

Biarchetype analysis identifies extreme instances of observations and features.

problem Representing complex data structures in a more interpretable form.
method Solves biarchetype analysis through an algorithm that identifies biarchetypes as mixtures of observations and features.
result Biarchetypes enhance interpretability of data structures compared to traditional methods.

The principal submatrix localization problem deals with recovering a K×KK\times K principal submatrix of elevated mean μμ in a large n×nn\times n symmetric matrix subject to additive standard Gaussian noise. This problem serves as a prototypical example for community detection, in which the community corresponds to the …

2015-10-30abs ↗pdf ↗

We provide a unified treatment of a broad class of noisy structure recovery problems, known as structured normal means problems. In this setting, the goal is to identify, from a finite collection of Gaussian distributions with different means, the distribution that produced some observed data. Recent work has studied s…

2015-06-25abs ↗pdf ↗

Many modern big data applications feature large scale in both numbers of responses and predictors. Better statistical efficiency and scientific insights can be enabled by understanding the large-scale response-predictor association network structures via layers of sparse latent factors ranked by importance. Yet sparsit…

2017-04-26abs ↗pdf ↗

Develops methods to analyze feature-outcome associations in subpopulations.

problem Challenges in understanding feature-outcome associations in high-dimensional data.
method Geometric decomposition framework using gradient flow and co-monotonicity decomposition.
result Identifies context-dependent patterns and improves statistical power and interpretability.

New spectral clustering method improves community detection in sparse networks.

problem Community detection in sparse networks using spectral clustering.
method Data-driven regularization and novel spectral truncation for adjacency matrix.
result Consistency results for community detection in general SBM and beyond.

Paper proposes ABDR for convex subspace clustering with adaptive block diagonal representation.

problem Subspace clustering with block diagonal structure for noisy data.
method ABDR explicitly pursues block diagonality without sacrificing convexity, using a specially designed convex regularizer.
result Experimental results show ABDR outperforms state-of-the-arts.

Divide-and-conquer method speeds sparse factorization for large matrices.

problem Sparse factorization of large matrices for statistical learning.
method Statistical problem formulation, divide-and-conquer approach, stagewise learning.
result Efficient algorithm with lower complexity than existing methods.