New framework extends ICA for non-independent variables, identifying pairwise mean independence.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper extends Median-of-Means to new learning problems involving pairwise comparisons.
A new k-means variant minimizes pairwise distances within clusters.
Exact pairwise ranking is achievable but not possible under noisy comparisons.
SAEs struggle with feature consistency across runs, hindering MI reliability.
PNN-smoothing improves -means clustering by merging subsets' clusterings.
In this paper, we unify the Markov theory of a variety of different types of graphs used in graphical Markov models by introducing the class of loopless mixed graphs, and show that all independence models induced by -separation on such graphs are compositional graphoids. We focus in particular on the subclass of rib…
Paper introduces differential pairwise privacy for secure metric learning.
Study on deep neural networks for reward modeling with pairwise comparison data.
Identifying components and estimating mixing weights in unlabeled finite mixtures under marginal independence.
Study shows attention-style models learn pairwise interactions efficiently.
New algorithm ranks players from partial comparisons with optimal rate.
Many machine learning problems can be formulated as predicting labels for a pair of objects. Problems of that kind are often referred to as pairwise learning, dyadic prediction or network inference problems. During the last decade kernel methods have played a dominant role in pairwise learning. They still obtain a stat…
This work maps Boltzmann distributions to ARNNs for better physics-based model approximations.
Three supervised learning methods for selecting logratios in compositional data analysis.
Pairwise quantile regression tackles similarity scoring in biometric systems.
EPFGNN models graph connections for better node classification.
A new method for binary ICA using non-stationary sources.
Recommender systems are one of the most pervasive applications of machine learning in industry, with many services using them to match users to products or information. As such it is important to ask: what are the possible fairness risks, how can we quantify them, and how should we address them? In this paper we offer …
We examine several aspects of explicability of a classification system built from neural networks. The first aspect is the pairwise explicability, which is the ability to provide the most accurate prediction when the range of possibilities is narrowed to just two. Next we consider explicability in development, which me…
Study the structure of international trade through hypergraphs.
There is a growing need for discrete choice models that account for the complex nature of human choices, escaping traditional behavioral assumptions such as the transitivity of pairwise preferences. Recently, several parametric models of intransitive comparisons have been proposed, but in all cases the maximum likeliho…
In a standard cluster analysis, such as k-means, in addition to clusters locations and distances between them, it's important to know if they are connected or well separated from each other. The main focus of this paper is discovering the relations between the resulting clusters. We propose a new method which is based …
Several classification methods assume that the underlying distributions follow tree-structured graphical models. Indeed, trees capture statistical dependencies between pairs of variables, which may be crucial to attain low classification errors. The resulting classifier is linear in the log-transformed univariate and b…
Exact simulation of correlated binary outcomes using PMF constraints and linear programming.
This note has been updated (April, 2020) to respond to "Towards Clarifying the Theory of the Deconfounder" by Yixin Wang, David M. Blei (arXiv:2003.04948). This original note, posted in January, 2020, is meant to complement our previous comment on "The Blessings of Multiple Causes" by Wang and Blei (2019). We provide a…
The study examines the independence of GKM manifolds and symmetric spaces.
We introduce a family of pairwise stochastic gradient estimators for gradients of expectations, which are related to the log-derivative trick, but involve pairwise interactions between samples. The simplest example of our new estimator, dubbed the fundamental trick estimator, is shown to arise from either a) introducin…
The functions of proteins and RNAs are determined by a myriad of interactions between their constituent residues, but most quantitative models of how molecular phenotype depends on genotype must approximate this by simple additive effects. While recent models have relaxed this constraint to also account for pairwise in…
The paper addresses monotonicity in machine learning models for fairness and accountability.
We consider data in the form of pairwise comparisons of n items, with the goal of precisely identifying the top k items for some value of k < n, or alternatively, recovering a ranking of all the items. We analyze the Copeland counting algorithm that ranks the items in order of the number of pairwise comparisons won, an…
This paper presents a distance-based discriminative framework for learning with probability distributions. Instead of using kernel mean embeddings or generalized radial basis kernels, we introduce embeddings based on dissimilarity of distributions to some reference distributions denoted as templates. Our framework exte…
The independence clustering problem is considered in the following formulation: given a set of random variables, it is required to find the finest partitioning of into clusters such that the clusters are mutually independent. Since mutual independence is the target, pairwise …
This paper optimizes the number of comparisons needed to find the best k items from pairwise comparisons.
Deep model learns protein interfaces from high-order interactions.
GTMs model complex multivariate data with varying conditional independencies.
The paper analyzes systemic risk in an insurance model with multiple business lines and heterogeneous claims.
We decompose the squared price-of-risk premium into three components: intervention-stable premium, confounding wedge, and information loss.
New method recovers causal order from dependent data.
Observational data usually comes with a multimodal nature, which means that it can be naturally represented by a multi-layer graph whose layers share the same set of vertices (users) with different edges (pairwise relationships). In this paper, we address the problem of combining different layers of the multi-layer gra…
In this paper, we present a simple non-parametric method for learning the structure of undirected graphs from data that drawn from an underlying unknown distribution. We propose to use Brownian distance covariance to estimate the conditional independences between the random variables and encodes pairwise Markov graph. …
We consider a system of three surfaces, graphs over a bounded domain in , intersecting along a time-dependent curve and moving by mean curvature while preserving the pairwise angles at the curve of intersection (equal to .) For the corresponding two-dimensional parabolic free boundary problem we pr…
Local Clustering improves semi-supervised learning models.
Algorithm learns Gaussian mixtures robust to outliers.
We introduce an evolutionary algorithm called recombinator--means for optimizing the highly non-convex kmeans problem. Its defining feature is that its crossover step involves all the members of the current generation, stochastically recombining them with a repurposed variant of the -means++ seeding algorithm. Th…
Rank aggregation systems collect ordinal preferences from individuals to produce a global ranking that represents the social preference. Rank-breaking is a common practice to reduce the computational complexity of learning the global ranking. The individual preferences are broken into pairwise comparisons and applied t…
Clustering is inherently ill-posed: there often exist multiple valid clusterings of a single dataset, and without any additional information a clustering system has no way of knowing which clustering it should produce. This motivates the use of constraints in clustering, as they allow users to communicate their interes…
A new framework for clustering with uncertainty quantification.