A tutorial on using MDL for graph analysis, focusing on clique size.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Sublinear algorithms detect cliques in graphs with high probability.
Paper explores embedding methods for detecting pseudo-cliques in random graphs, showing limitations and potential.
Cliques, or fully connected subgraphs, are among the most important and well-studied graph motifs in network science. We consider the problem of finding a statisti- cally anomalous clique hidden in a large network. There are two parts to this problem: (1) detection, i.e., determining whether an anomalous clique is pres…
Deep nets learn structured densities without dimensionality issues.
MFCF algorithm learns conditional dependency structure from sparse data.
Maximal knotless graphs have at least 74% of their vertices' edges.
Develops a method to efficiently learn causal DAGs using directed clique trees.
Estimates log-concave densities in graphical models using tent functions.
We say a graph has property when it is an induced subgraph of the curve graph of a surface of genus with punctures. Two well-known graph invariants, the chromatic and clique numbers, can provide obstructions to . We introduce a new invariant of a graph, the 'nested complex…
The maximum number of maximum cliques in a graph is determined for graphs with at least 15 vertices.
Solves partial assignment problems using random clique complexes.
Infinite clique of rays in plane minus Cantor set.
Graph clustering involves the task of dividing nodes into clusters, so that the edge density is higher within clusters as opposed to across clusters. A natural, classic and popular statistical setting for evaluating solutions to this problem is the stochastic block model, also referred to as the planted partition model…
Reduces average-case complexity of sparse PCA from weak PC conjectures.
Research uses machine learning to find central nodes and cliques in YouTube social networks.
Tackles the computational hardness of HPC detection, conjecturing equivalence to PC detection.
We introduce Clique Matrices as an alternative representation of undirected graphs, being a generalisation of the incidence matrix representation. Here we use clique matrices to decompose a graph into a set of possibly overlapping clusters, de ned as well-connected subsets of vertices. The decomposition is based on a s…
New insights link diverse statistical problems via secret leakage planted clique.
Proposes clique pooling for graph classification.
Brain Electroencephalography (EEG) classification is widely applied to analyze cerebral diseases in recent years. Unfortunately, invalid/noisy EEGs degrade the diagnosis performance and most previously developed methods ignore the necessity of EEG selection for classification. To this end, this paper proposes a novel m…
The problem of categorical data analysis in high dimensions is considered. A discussion of the fundamental difficulties of probability modeling is provided, and a solution to the derivation of high dimensional probability distributions based on Bayesian learning of clique tree decomposition is presented. The main contr…
Artin groups not free of infinity are shown to have finite centers.
Study evaluates methods for expanding communities in hypergraphs using random walks.
In financial markets, abnormal trading behaviors pose a serious challenge to market surveillance and risk management. What is worse, there is an increasing emergence of abnormal trading events that some experienced traders constitute a collusive clique and collaborate to manipulate some instruments, thus mislead other …
New method tests causal relationships from data without needing to learn the entire graph.
Let be an orientable surface with negative Euler characteristic. For , let denote the , whose vertices are isotopy classes of essential simple closed curves on , and whose edges correspond to pairs of curves that can be realized to intersect at most …
New algorithm finds minimal causal models for complex latent variables.
We introduce a new embarrassingly parallel parameter learning algorithm for Markov random fields with untied parameters which is efficient for a large class of practical models. Our algorithm parallelizes naturally over cliques and, for graphs of bounded degree, its complexity is linear in the number of cliques. Unlike…
Label assignment problems with large state spaces are important tasks especially in computer vision. Often the pairwise interaction (or smoothness prior) between labels assigned at adjacent nodes (or pixels) can be described as a function of the label difference. Exact inference in such labeling tasks is still difficul…
These notes review six lectures given by Prof. Andrea Montanari on the topic of statistical estimation for linear models. The first two lectures cover the principles of signal recovery from linear measurements in terms of minimax risk. Subsequent lectures demonstrate the application of these principles to several pract…
Proposes CLIQUE for improved local variable importance in multi-class classification.
We construct a partial order relation which acts on the set of 3-cliques of a maximal planar graph G and defines a unique hierarchy. We demonstrate that G is the union of a set of special subgraphs, named `bubbles', that are themselves maximal planar graphs. The graph G is retrieved by connecting these bubbles in a tre…
In many domains, there is significant interest in capturing novel relationships between time series that represent activities recorded at different nodes of a highly complex system. In this paper, we introduce multipoles, a novel class of linear relationships between more than two time series. A multipole is a set of t…
The study shows acylindrical hyperbolicity for Artin groups not associated with joins or cones.
We introduce the concept of community trees that summarizes topological structures within a network. A community tree is a tree structure representing clique communities from the clique percolation method (CPM). The community tree also generates a persistent diagram. Community trees and persistent diagrams reveal topol…
New algorithm identifies near-optimal policies in adversarial distributed RL settings.
Efficient algorithm for self-directed learning of convex clusters on graphs.
This work sets a universal lower bound for learning causal DAGs with atomic interventions.
In this paper we study speaker linking (a.k.a.\ partitioning) given constraints of the distribution of speaker identities over speech recordings. Specifically, we show that the intractable partitioning problem becomes tractable when the constraints pre-partition the data in smaller cliques with non-overlapping speakers…
As a model problem for clustering, we consider the densest k-disjoint-clique problem of partitioning a weighted complete graph into k disjoint subgraphs such that the sum of the densities of these subgraphs is maximized. We establish that such subgraphs can be recovered from the solution of a particular semidefinite re…
This paper studies the problem of detecting the presence of a small dense community planted in a large Erdős-Rényi random graph , where the edge probability within the community exceeds by a constant factor. Assuming the hardness of the planted clique detection problem, we show that the computatio…
Paper shows statistical-computational gaps in learning sparse mixtures and robust estimation.
Paper calculates Gromov-Hausdorff distance between simplexes and 2-distance spaces.
Modern data acquisition routinely produces massive amounts of network data. Though many methods and models have been proposed to analyze such data, the research of network data is largely disconnected with the classical theory of statistical learning and signal processing. In this paper, we present a new framework for …
A new hypergraph-based active learning scheme reduces query complexity.
The Cartesian subgroup in graph products of groups is studied with bounds and algorithms.
Given a large data matrix , we consider the problem of determining whether its entries are i.i.d. with some known marginal distribution , or instead contains a principal submatrix whose entries have marginal distribution . As …