Active seriation recovers item order from noisy pairwise similarity measurements.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We propose a novel parameterized family of Mixed Membership Mallows Models (M4) to account for variability in pairwise comparisons generated by a heterogeneous population of noisy and inconsistent users. M4 models individual preferences as a user-specific probabilistic mixture of shared latent Mallows components. Our k…
Develops a hypothesis testing framework for generalized Thurstone models.
Study shows attention-style models learn pairwise interactions efficiently.
In spectral clustering and spectral image segmentation, the data is partioned starting from a given matrix of pairwise similarities S. the matrix S is constructed by hand, or learned on a separate training set. In this paper we show how to achieve spectral clustering in unsupervised mode. Our algorithm starts with a se…
ETM identifies field-specific keywords in text classification.
New framework extends ICA for non-independent variables, identifying pairwise mean independence.
Given a set of pairwise comparisons, the classical ranking problem computes a single ranking that best represents the preferences of all users. In this paper, we study the problem of inferring individual preferences, arising in the context of making personalized recommendations. In particular, we assume that there are …
We study exact recovery conditions for convex relaxations of point cloud clustering problems, focusing on two of the most common optimization problems for unsupervised clustering: -means and -median clustering. Motivations for focusing on convex relaxations are: (a) they come with a certificate of optimality, and…
New clustering method ensures fairness and community preservation.
EPFGNN models graph connections for better node classification.
Similarity-based clustering and semi-supervised learning methods separate the data into clusters or classes according to the pairwise similarity between the data, and the pairwise similarity is crucial for their performance. In this paper, we propose a novel discriminative similarity learning framework which learns dis…
In this paper, we unify the Markov theory of a variety of different types of graphs used in graphical Markov models by introducing the class of loopless mixed graphs, and show that all independence models induced by -separation on such graphs are compositional graphoids. We focus in particular on the subclass of rib…
The paper extends Johnson's result on Torelli group homology.
Surrogate-based analysis of interactions via local effect smooths
EM algorithm achieves optimal sample complexity for well-separated Gaussian mixtures.
Signed pairwise interactions conflate uniqueness, redundancy, and synergy
The classification of shapes is of great interest in diverse areas ranging from medical imaging to computer vision and beyond. While many statistical frameworks have been developed for the classification problem, most are strongly tied to early formulations of the problem - with an object to be classified described as …
We decompose the squared price-of-risk premium into three components: intervention-stable premium, confounding wedge, and information loss.
We consider the problem of learning a mixture of Random Utility Models (RUMs). Despite the success of RUMs in various domains and the versatility of mixture RUMs to capture the heterogeneity in preferences, there has been only limited progress in learning a mixture of RUMs from partial data such as pairwise comparisons…
For a certain class of distributions, we prove that the linear programming relaxation of -medoids clustering---a variant of -means clustering where means are replaced by exemplars from within the dataset---distinguishes points drawn from nonoverlapping balls with high probability once the number of points drawn a…
Algorithm learns Gaussian mixtures robust to outliers.
NucleusDiff models atomic nuclei interactions to prevent separation violations in drug design.
We prove that the group D^r(R) of C^r diffeomorphisms of the real line, endowed with the compact-open and Whitney C^r topologies, is bihomeomorphic to the group H(R) of homeomorphisms of the real line endowed with the compact-open and Whitney topologies. This implies that the diffeomorphism group D^r(R) endowed with th…
Decomposes Goldman-Turaev Lie bialgebra via cutting a surface.
Study on torsion in homology of Torelli group for surfaces.
In the context of clustering, we consider a generative model in a Euclidean ambient space with clusters of different shapes, dimensions, sizes and densities. In an asymptotic setting where the number of points becomes large, we obtain theoretical guaranties for a few emblematic methods based on pairwise distances: a si…
A new method for name disambiguation in academic networks using multi-view attention and recurrent neural networks.
Exact simulation of correlated binary outcomes using PMF constraints and linear programming.
Polynomial-time algorithm for clustering mixtures with separation Δ=Ω(√(log k)).
In a standard cluster analysis, such as k-means, in addition to clusters locations and distances between them, it's important to know if they are connected or well separated from each other. The main focus of this paper is discovering the relations between the resulting clusters. We propose a new method which is based …
In the context of clustering, we assume a generative model where each cluster is the result of sampling points in the neighborhood of an embedded smooth surface; the sample may be contaminated with outliers, which are modeled as points sampled in space away from the clusters. We consider a prototype for a higher-order …
Unified framework for binary responses using AUC loss and low-rank constraint.
Algorithm clusters mixtures with bounded covariances under specific separation conditions.
Theory predicts neural scaling exponents from language statistics.
Semi-supervised learning (SSL) has become important in current data analysis applications, where the amount of unlabeled data is growing exponentially and user input remains limited by logistics and expense. Constrained clustering, as a subclass of SSL, makes use of user input in the form of relationships between data …
We study the problem of ranking from crowdsourced pairwise comparisons. Answers to pairwise tasks are known to be affected by the position of items on the screen, however, previous models for aggregation of pairwise comparisons do not focus on modeling such kind of biases. We introduce a new aggregation model factorBT …
We study the classification performance of Kronecker-structured models in two asymptotic regimes and developed an algorithm for separable, fast and compact K-S dictionary learning for better classification and representation of multidimensional signals by exploiting the structure in the signal. First, we study the clas…
A new method clusters intersecting lines using hypergraphs.
Metric learning for classification has been intensively studied over the last decade. The idea is to learn a metric space induced from a normed vector space on which data from different classes are well separated. Different measures of the separation thus lead to various designs of the objective function in the metric …
In this study, a pairwise comparison matrix is generalized to the case when coefficients create Lie group , non necessarily abelian. A necessary and sufficient criterion for pairwise comparisons matrices to be consistent is provided. Basic criteria for finding a nearest consistent pairwise comparisons matrix (extend…
Paper introduces differential pairwise privacy for secure metric learning.
Improved sample complexity for Gaussian Mixture Models using Pair Correlation Factor.
Develops a smooth operator framework for analyzing neural network representations.
Algorithm identifies best item from subsets with random utility model feedback.
This paper uses Factored Latent Analysis (FLA) to learn a factorized, segmental representation for observations of tracked objects over time. Factored Latent Analysis is latent class analysis in which the observation space is subdivided and each aspect of the original space is represented by a separate latent class mod…
Cross-entropy loss linked to metric learning, outperforming complex pairwise losses.
Dynamic Vine Copulas detect and quantify time-varying higher-order interactions in multivariate systems.