New method extracts hidden phases in binary mixtures using tubular tilings.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New insights into identifying mixtures of product distributions using Hadamard extensions.
Activity coefficients, which are a measure of the non-ideality of liquid mixtures, are a key property in chemical engineering with relevance to modeling chemical and phase equilibria as well as transport processes. Although experimental data on thousands of binary mixtures are available, prediction methods are needed t…
The paper optimizes hyperplanes for binary classification in high-dimensional data with latent Gaussian mixtures.
Dasgupta and Shulman showed that a two-round variant of the EM algorithm can learn mixture of Gaussian distributions with near optimal precision with high probability if the Gaussian distributions are well separated and if the dimension is sufficiently high. In this paper, we generalize their theory to learning mixture…
BIND removes background noise from binary matrices, improving detection accuracy and fairness.
New method estimates hidden binary mixture model centers efficiently.
In this paper we propose a mixture model, SparseMix, for clustering of sparse high dimensional binary data, which connects model-based with centroid-based clustering. Every group is described by a representative and a probability distribution modeling dispersion from this representative. In contrast to classical mixtur…
We propose a new blind source separation algorithm based on mixtures of alpha-stable distributions. Complex symmetric alpha-stable distributions have been recently showed to better model audio signals in the time-frequency domain than classical Gaussian distributions thanks to their larger dynamic range. However, infer…
Paper analyzes adversarial training's performance in binary classification.
Our paper examines binary linear classification under Gaussian mixtures, revealing conditions for optimal performance.
We compare two statistical models of three binary random variables. One is a mixture model and the other is a product of mixtures model called a restricted Boltzmann machine. Although the two models we study look different from their parametrizations, we show that they represent the same set of distributions on the int…
Self-training improves weak classifiers in mixture models.
We study the problem of finding the smallest such that every element of an exponential family can be written as a mixture of elements of another exponential family. We propose an approach based on coverings and packings of the face lattice of the corresponding convex support polytopes and results from coding th…
Archetypal analysis is an exploratory tool that explains a set of observations as mixtures of pure (extreme) patterns. If the patterns are actual observations of the sample, we refer to them as archetypoids. For the first time, we propose to use archetypoid analysis for binary observations. This tool can contribute to …
Study on recovering sparse linear classifiers from mixed binary responses.
A new method detects outliers using ensembles of Dirichlet process mixtures.
A Bernoulli Mixture Model (BMM) is a finite mixture of random binary vectors with independent dimensions. The problem of clustering BMM data arises in a variety of real-world applications, ranging from population genetics to activity analysis in social networks. In this paper, we analyze the clusterability of BMMs from…
Binary PheNorm extends phenotype labeling for EHRs using binary silver labels.
Defines diversification as a binary relationship between financial portfolios.
We propose a mixture of latent trait models with common slope parameters (MCLT) for model-based clustering of high-dimensional binary data, a data type for which few established methods exist. Recent work on clustering of binary data, based on a -dimensional Gaussian latent variable, is extended by incorporating com…
Phase segregation, the process by which the components of a binary mixture spontaneously separate, is a key process in the evolution and design of many chemical, mechanical, and biological systems. In this work, we present a data-driven approach for the learning, modeling, and prediction of phase segregation. A direct …
In many real life problems, objects are described by large number of binary features. For instance, documents are characterized by presence or absence of certain keywords; cancer patients are characterized by presence or absence of certain mutations etc. In such cases, grouping together similar objects/profiles based o…
A new method for density estimation using mixture discrepancy and moments.
In binary-transaction data-mining, traditional frequent itemset mining often produces results which are not straightforward to interpret. To overcome this problem, probability models are often used to produce more compact and conclusive results, albeit with some loss of accuracy. Bayesian statistics have been widely us…
Optimal classifiers derived from GMMs are approximated by deep neural networks.
In this paper, we implement multi-label neural networks with optimal thresholding to identify gas species among a multi gas mixture in a cluttered environment. Using infrared absorption spectroscopy and tested on synthesized spectral datasets, our approach outperforms conventional binary relevance - partial least squar…
Study shows optimal self-distillation improves model performance on noisy data.
The paper analyzes the score field of diffusion models using Burgers dynamics.
Beta-SOD detects and corrects noisy object re-identification using cosine similarity and Beta mixtures.
Hierarchical structure is ubiquitous in data across many domains. There are many hierarchical clustering methods, frequently used by domain experts, which strive to discover this structure. However, most of these methods limit discoverable hierarchies to those with binary branching structure. This limitation, while com…
Many different methods to train deep generative models have been introduced in the past. In this paper, we propose to extend the variational auto-encoder (VAE) framework with a new type of prior which we call "Variational Mixture of Posteriors" prior, or VampPrior for short. The VampPrior consists of a mixture distribu…
Latent variable models with hidden binary units appear in various applications. Learning such models, in particular in the presence of noise, is a challenging computational problem. In this paper we propose a novel spectral approach to this problem, based on the eigenvectors of both the second order moment matrix and t…
Deep neural networks classify unbounded Gaussian mixture data without dimensionality issues.
ETM models improve efficiency in semi-supervised logistic regression.
This paper optimizes retraining models using their own predictions and noisy labels.
VBphenoR uses variational Bayes for EHR-based patient phenotyping.
This article investigates the eigenspectrum of the inner product-type kernel matrix under a binary mixture model in the high dimensional regime where the number of data and their dimension are both large and comparable. Based on…
We present a theoretical analysis of Gaussian-binary restricted Boltzmann machines (GRBMs) from the perspective of density models. The key aspect of this analysis is to show that GRBMs can be formulated as a constrained mixture of Gaussians, which gives a much better insight into the model's capabilities and limitation…
The paper confirms two groups of gamma-ray bursts using a new nonparametric metric.
Federated learning for Bayesian clustering of large datasets.
Algorithm identifies probability distributions from noisy moments with minimal samples.
Regularized linear regression improves binary classification performance, especially with ridge and regularization.
Transformers for binary decisions are sensitive to evidence order, leading to unreliable outcomes.
We introduce the problem of learning mixtures of subcubes over , which contains many classic learning theory problems as a special case (and is itself a special case of others). We give a surprising -time learning algorithm based on higher-order multilinear moments. It is not possible to l…
A matrix completion problem, which aims to recover a complete matrix from its partial observations, is one of the important problems in the machine learning field and has been studied actively. However, there is a discrepancy between the mainstream problem setting, which assumes continuous-valued observations, and some…
We present a new approach for learning compact and intuitive distributed representations with binary encoding. Rather than summing up expert votes as in products of experts, we employ for each variable the opinion of the most reliable expert. Data points are hence explained through a partitioning of the variables into …
Method reveals dissimilarity in alloys' Curie temperatures.