PSimGNN partitions graphs into subgraphs for efficient graph similarity computation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper analyzes indices based on counting object pairs for assessing partition agreement in unsupervised learning.
Exchangeable graphs arise via a sampling procedure from measurable functions known as graphons. A natural estimation problem is how well we can recover a graphon given a single graph sampled from it. One general framework for estimating a graphon uses step-functions obtained by partitioning the nodes of the graph accor…
Partition functions of probability distributions are important quantities for model evaluation and comparisons. We present a new method to compute partition functions of complex and multimodal distributions. Such distributions are often sampled using simulated tempering, which augments the target space with an auxiliar…
Paper tackles clustering with ordinal comparisons, achieving near-optimal results.
In a series of recent works, we have generalised the consistency results in the stochastic block model literature to the case of uniform and non-uniform hypergraphs. The present paper continues the same line of study, where we focus on partitioning weighted uniform hypergraphs---a problem often encountered in computer …
Bayesian approach for multifile record linkage and duplicate detection.
WHOMP optimizes randomized controlled trials by minimizing subgroup bias.
Develops comparison-based hierarchical clustering algorithms without object representations.
New loss function calibrates WW-hinge loss for multiclass SVM.
Partition functions arise in a variety of settings, including conditional random fields, logistic regression, and latent gaussian models. In this paper, we consider semistochastic quadratic bound (SQB) methods for maximum likelihood inference based on partition function optimization. Batch methods based on the quadrati…
A common problem in machine learning is to rank a set of n items based on pairwise comparisons. Here ranking refers to partitioning the items into sets of pre-specified sizes according to their scores, which includes identification of the top-k items as the most prominent special case. The score of a given item is defi…
A new method for density estimation using mixture discrepancy and moments.
Greedy training of recursive partitioning estimators faces a computational barrier when the true function doesn't satisfy a specific property.
Improved supervised EM learning for shared kernel models with feature space partitioning.
Improved spectral clustering via Gromov-Wasserstein Learning.
The Restricted Boltzmann Machines (RBM) can be used either as classifiers or as generative models. The quality of the generative RBM is measured through the average log-likelihood on test data. Due to the high computational complexity of evaluating the partition function, exact calculation of test log-likelihood is ver…
We present an approach to deep estimation of discrete conditional probability distributions. Such models have several applications, including generative modeling of audio, image, and video data. Our approach combines two main techniques: dyadic partitioning and graph-based smoothing of the discrete space. By recursivel…
We propose using five data-driven community detection approaches from social networks to partition the label space for the task of multi-label classification as an alternative to random partitioning into equal subsets as performed by RAkELd: modularity-maximizing fastgreedy and leading eigenvector, infomap, walktrap an…
Novel hybrid method for Bayesian network structure learning reduces computational time without sacrificing accuracy.
Locally Optimal Block Preconditioned Conjugate Gradient (LOBPCG) is demonstrated to efficiently solve eigenvalue problems for graph Laplacians that appear in spectral clustering. For static graph partitioning, 10-20 iterations of LOBPCG without preconditioning result in ~10x error reduction, enough to achieve 100% corr…
A new procedure aggregates models to predict data from multiple clusters.
We consider sequential or active ranking of a set of n items based on noisy pairwise comparisons. Items are ranked according to the probability that a given item beats a randomly chosen item, and ranking refers to partitioning the items into sets of pre-specified sizes according to their scores. This notion of ranking …
Efficient algorithm for matching graphs with community structure.
We explore the performance of several automatic bandwidth selectors, originally designed for density gradient estimation, as data-based procedures for nonparametric, modal clustering. The key tool to obtain a clustering from density gradient estimators is the mean shift algorithm, which allows to obtain a partition not…
New method targets relative risk heterogeneity in clinical trials.
Optimized parallel algorithms for identifying strong ties in data.
We propose a new algorithm called PLUTO for building logistic regression trees to binary response data. PLUTO can capture the nonlinear and interaction patterns in messy data by recursively partitioning the sample space. It fits a simple or a multiple linear logistic regression model in each partition. PLUTO employs th…
CW-EDMD improves prediction accuracy by learning local Koopman models for different state-space regions.
In several application domains, high-dimensional observations are collected and then analysed in search for naturally occurring data clusters which might provide further insights about the nature of the problem. In this paper we describe a new approach for partitioning such high-dimensional data. Our assumption is that…
Breaks the hardness conjecture for batch RL with a novel tournament-based approach.
We show that the link invariants derived from 3-dimensional quantum hyperbolic geometry can be defined by means of planar state sums based on link diagrams and a new family of enhanced Yang-Baxteroperators (YBO) that we compute explicitly. By a local comparison of the respective YBO's we show that these invariants coin…
Multirate training speeds up neural network fine-tuning.
ProtoBandit uses bandits to find prototypes efficiently.
This paper tackles deep clustering evaluation challenges in high-dimensional data.
ECG improves graph clustering and resolves resolution limit issues.
Bayesian method finds patterns of mutual independence in data.
Active learning optimizes correlation clustering by querying the most informative pairwise comparisons.
The study limits how many parts regular simplicial partitions can overlap.
Partial soft-matching distance improves neural representation comparison by allowing some neurons to remain unmatched.
Hypergraph partitioning lies at the heart of a number of problems in machine learning and network sciences. Many algorithms for hypergraph partitioning have been proposed that extend standard approaches for graph partitioning to the case of hypergraphs. However, theoretical aspects of such methods have seldom received …
New method improves nearest neighbor search using neural networks and graph partitioning.
Clustering ensemble is one of the most recent advances in unsupervised learning. It aims to combine the clustering results obtained using different algorithms or from different runs of the same clustering algorithm for the same data set, this is accomplished using on a consensus function, the efficiency and accuracy of…
Rectangular Bounding Process (RBP) improves partitioning efficiency in multi-dimensional spaces.
GAP uses deep learning to efficiently partition graphs.
In this paper, we propose a family of graph partition similarity measures that take the topology of the graph into account. These graph-aware measures are alternatives to using set partition similarity measures that are not specifically designed for graph partitions. The two types of measures, graph-aware and set parti…
The study examines the balancedness of random partition models and finds the rich-get-richer characteristic is a result of model assumptions.
The paper develops mixed-integer formulations for neural networks using partitioning.