SplitNN-driven Vertical Partitioning enables distributed learning from diverse data sources.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Asynchronous federated learning for vertically partitioned data improves efficiency and privacy.
We investigate the minimal number of links and knots in complete partite graphs. We provide exact values or bounds on the minimal number of links for all complete partite graphs with all but 4 vertices in one partition, or with 9 vertices in total. In particular, we find that the minimal number of links for …
A new framework for federated learning tackles challenges with horizontally partitioned labels and stragglers.
Differentially private method for synthetic data generation from vertically partitioned data.
FDSKL algorithm trains vertically partitioned data with kernels securely and efficiently.
A new framework for clustering high-dimensional data using vertical shards.
We find the minimal number of links in an embedding of any complete -partite graph on 7 vertices (including , which has at least 21 links). We give either exact values or upper and lower bounds for the minimal number of links for all complete -partite graphs on 8 vertices. We also look at larger complete bip…
Graphs can model interactions between vertices, but how well depends on graph structure.
Preserving differential privacy has been well studied under centralized setting. However, it's very challenging to preserve differential privacy under multiparty setting, especially for the vertically partitioned case. In this work, we propose a new framework for differential privacy preserving multiparty learning in t…
We study the localization of a cluster of activated vertices in a graph, from adaptively designed compressive measurements. We propose a hierarchical partitioning of the graph that groups the activated vertices into few partitions, so that a top-down sensing procedure can identify these partitions, and hence the activa…
Secure Multiparty Computation protects data privacy in Symbolic Regression.
New invariant extends curvature estimates to noncompact manifolds.
We introduce new sufficient conditions for intrinsic knotting and linking. A graph on n vertices with at least 4n-9 edges is intrinsically linked. A graph on n vertices with at least 5n-14 edges is intrinsically knotted. We also classify graphs that are 0, 1, or 2 edges short of being complete partite graphs with respe…
We introduce a method for creating a special type of tree, called a tree position, from a weighted graph. Leaves of the tree correspond to vertices of the original graph, and the tree edges contain information which can be used to partition these vertices. By repeatedly applying reducing operations to the tree position…
We establish a necessary and sufficient condition for a heptagonal knot to be figure-8 knot. The condition is described by a set of Radon partitions formed by vertices of the heptagon. In addition we relate this result to the number of nontrivial heptagonal knots in linear embeddings of the complete graph into $\…
In earlier work the Kauffman bracket polynomial was extended to an invariant of marked graphs, i.e., looped graphs whose vertices have been partitioned into two classes (marked and not marked). The marked-graph bracket polynomial is readily modified to handle graphs with weighted vertices. We present formulas that simp…
Hypergraph partitioning is an important problem in machine learning, computer vision and network analytics. A widely used method for hypergraph partitioning relies on minimizing a normalized sum of the costs of partitioning hyperedges across clusters. Algorithmic solutions based on this approach assume that different p…
Suppose that one particular block in a stochastic block model is of interest, but block labels are only observed for a few of the vertices in the network. Utilizing a graph realized from the model and the observed block labels, the vertex nomination task is to order the vertices with unobserved block labels into a rank…
We classify graphs that are 0, 1, or 2 edges short of being complete partite graphs with respect to intrinsic linking and intrinsic knotting. In addition, we classify intrinsic knotting of graphs on 8 vertices. For graphs in these families, we verify a conjecture presented in Adams' "The Knot Book": If a vertex is remo…
The binary symmetric stochastic block model deals with a random graph of vertices partitioned into two equal-sized clusters, such that each pair of vertices is connected independently with probability within clusters and across clusters. In the asymptotic regime of and for fixe…
New study shows limits of low-degree algorithms in finding large independent sets in sparse hypergraphs.
Privacy is crucial in many applications of machine learning. Legal, ethical and societal issues restrict the sharing of sensitive data making it difficult to learn from datasets that are partitioned between many parties. One important instance of such a distributed setting arises when information about each record in t…
Graph embedding is a popular algorithmic approach for creating vector representations for individual vertices in networks. Training these algorithms at scale is important for creating embeddings that can be used for classification, ranking, recommendation and other common applications in industry. While industrial syst…
SDP approach recovers communities in multilayer hypergraphs from aggregated similarity matrices.
Combining data from varied sources has considerable potential for knowledge discovery: collaborating data parties can mine data in an expanded feature space, allowing them to explore a larger range of scientific questions. However, data sharing among different parties is highly restricted by legal conditions, ethical c…
HDP-VFL hybridizes DP for VFL, reducing privacy costs.
Locally Optimal Block Preconditioned Conjugate Gradient (LOBPCG) is demonstrated to efficiently solve eigenvalue problems for graph Laplacians that appear in spectral clustering. For static graph partitioning, 10-20 iterations of LOBPCG without preconditioning result in ~10x error reduction, enough to achieve 100% corr…
Paper solves graph matching for correlated Erdős--Rényi graphs.
Fitting statistical models is computationally challenging when the sample size or the dimension of the dataset is huge. An attractive approach for down-scaling the problem size is to first partition the dataset into subsets and then fit using distributed algorithms. The dataset can be partitioned either horizontally (i…
Spectral algorithm recovers community structure in sparse hypergraphs.
Network representation learning, as an approach to learn low dimensional representations of vertices, has attracted considerable research attention recently. It has been proven extremely useful in many machine learning tasks over large graph. Most existing methods focus on learning the structural representations of ver…
Paper tackles scalable VFL with data augmentation and amortized inference.
The Gibbs sampler is a particularly popular Markov chain used for learning and inference problems in Graphical Models (GMs). These tasks are computationally intractable in general, and the Gibbs sampler often suffers from slow mixing. In this paper, we study the Swendsen-Wang dynamics which is a more sophisticated Mark…
A new method learns graph compression from data.
A new algorithm reduces CI tests for causal graph recovery.
Bipartite networks are a common type of network data in which there are two types of vertices, and only vertices of different types can be connected. While bipartite networks exhibit community structure like their unipartite counterparts, existing approaches to bipartite community detection have drawbacks, including im…
New conjectures link SU(r) Vafa-Witten invariants to Ramanujan's continued fractions.
A meander of order n is a simple closed curve in the plane which intersects a horizontal line transversely at 2n points. (Meanders which differ by an isotopy of the line and plane are considered equivalent.) Let Gamma_n be the Cayley graph of the symmetric group S_n as generated by all (n choose 2) transpositions. Let …
There is a recent surge of interest in identifying the sharp recovery thresholds for cluster recovery under the stochastic block model. In this paper, we address the more refined question of how many vertices that will be misclassified on average. We consider the binary form of the stochastic block model, where ver…
A central problem in analyzing networks is partitioning them into modules or communities. One of the best tools for this is the stochastic block model, which clusters vertices into blocks with statistically homogeneous pattern of links. Despite its flexibility and popularity, there has been a lack of principled statist…
The community detection problem for graphs asks one to partition the n vertices V of a graph G into k communities, or clusters, such that there are many intracluster edges and few intercluster edges. Of course this is equivalent to finding a permutation matrix P such that, if A denotes the adjacency matrix of G, then P…
We consider the problem of identifying underlying community-like structures in graphs. Towards this end we study the Stochastic Block Model (SBM) on -clusters: a random model on vertices, partitioned in equal sized clusters, with edges sampled independently across clusters with probability and within …
Efficiently learns monophonic halfspaces in graph vertices.
SpaPool combines dense and sparse techniques for efficient graph pooling.
Paper tackles community recovery in binary symmetric SBM graphs.
VFGNN tackles privacy-preserving node classification with federated GNN.
We consider the Degree-Corrected Stochastic Block Model (DC-SBM): a random graph on nodes, having i.i.d. weights (possibly heavy-tailed), partitioned into asymptotically equal-sized clusters. The model parameters are two constants and the finite second moment of the weights $Φ^{…