New method predicts activity coefficients for binary mixtures without using physical descriptors.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New method extracts hidden phases in binary mixtures using tubular tilings.
The paper optimizes hyperplanes for binary classification in high-dimensional data with latent Gaussian mixtures.
New method estimates hidden binary mixture model centers efficiently.
Dasgupta and Shulman showed that a two-round variant of the EM algorithm can learn mixture of Gaussian distributions with near optimal precision with high probability if the Gaussian distributions are well separated and if the dimension is sufficiently high. In this paper, we generalize their theory to learning mixture…
New insights into identifying mixtures of product distributions using Hadamard extensions.
In this paper we propose a mixture model, SparseMix, for clustering of sparse high dimensional binary data, which connects model-based with centroid-based clustering. Every group is described by a representative and a probability distribution modeling dispersion from this representative. In contrast to classical mixtur…
We propose a new blind source separation algorithm based on mixtures of alpha-stable distributions. Complex symmetric alpha-stable distributions have been recently showed to better model audio signals in the time-frequency domain than classical Gaussian distributions thanks to their larger dynamic range. However, infer…
Paper analyzes adversarial training's performance in binary classification.
BIND removes background noise from binary matrices, improving detection accuracy and fairness.
We compare two statistical models of three binary random variables. One is a mixture model and the other is a product of mixtures model called a restricted Boltzmann machine. Although the two models we study look different from their parametrizations, we show that they represent the same set of distributions on the int…
Our paper examines binary linear classification under Gaussian mixtures, revealing conditions for optimal performance.
Self-training improves weak classifiers in mixture models.
A new method detects outliers using ensembles of Dirichlet process mixtures.
A Bernoulli Mixture Model (BMM) is a finite mixture of random binary vectors with independent dimensions. The problem of clustering BMM data arises in a variety of real-world applications, ranging from population genetics to activity analysis in social networks. In this paper, we analyze the clusterability of BMMs from…
Study on recovering sparse linear classifiers from mixed binary responses.
We propose a mixture of latent trait models with common slope parameters (MCLT) for model-based clustering of high-dimensional binary data, a data type for which few established methods exist. Recent work on clustering of binary data, based on a -dimensional Gaussian latent variable, is extended by incorporating com…
In binary-transaction data-mining, traditional frequent itemset mining often produces results which are not straightforward to interpret. To overcome this problem, probability models are often used to produce more compact and conclusive results, albeit with some loss of accuracy. Bayesian statistics have been widely us…
Binary PheNorm extends phenotype labeling for EHRs using binary silver labels.
Phase segregation, the process by which the components of a binary mixture spontaneously separate, is a key process in the evolution and design of many chemical, mechanical, and biological systems. In this work, we present a data-driven approach for the learning, modeling, and prediction of phase segregation. A direct …
Archetypal analysis helps understand binary data sets.
Defines diversification as a binary relationship between financial portfolios.
We study the problem of finding the smallest such that every element of an exponential family can be written as a mixture of elements of another exponential family. We propose an approach based on coverings and packings of the face lattice of the corresponding convex support polytopes and results from coding th…
In many real life problems, objects are described by large number of binary features. For instance, documents are characterized by presence or absence of certain keywords; cancer patients are characterized by presence or absence of certain mutations etc. In such cases, grouping together similar objects/profiles based o…
Study shows optimal self-distillation improves model performance on noisy data.
This paper optimizes retraining models using their own predictions and noisy labels.
Optimal classifiers derived from GMMs are approximated by deep neural networks.
Latent variable models with hidden binary units appear in various applications. Learning such models, in particular in the presence of noise, is a challenging computational problem. In this paper we propose a novel spectral approach to this problem, based on the eigenvectors of both the second order moment matrix and t…
The paper analyzes the score field of diffusion models using Burgers dynamics.
ETM models improve efficiency in semi-supervised logistic regression.
Deep neural networks classify unbounded Gaussian mixture data without dimensionality issues.
Many different methods to train deep generative models have been introduced in the past. In this paper, we propose to extend the variational auto-encoder (VAE) framework with a new type of prior which we call "Variational Mixture of Posteriors" prior, or VampPrior for short. The VampPrior consists of a mixture distribu…
Beta-SOD detects and corrects noisy object re-identification using cosine similarity and Beta mixtures.
We present a theoretical analysis of Gaussian-binary restricted Boltzmann machines (GRBMs) from the perspective of density models. The key aspect of this analysis is to show that GRBMs can be formulated as a constrained mixture of Gaussians, which gives a much better insight into the model's capabilities and limitation…
VBphenoR uses variational Bayes for EHR-based patient phenotyping.
Hierarchical structure is ubiquitous in data across many domains. There are many hierarchical clustering methods, frequently used by domain experts, which strive to discover this structure. However, most of these methods limit discoverable hierarchies to those with binary branching structure. This limitation, while com…
A new method for density estimation using mixture discrepancy and moments.
Federated learning for Bayesian clustering of large datasets.
In this paper, we implement multi-label neural networks with optimal thresholding to identify gas species among a multi gas mixture in a cluttered environment. Using infrared absorption spectroscopy and tested on synthesized spectral datasets, our approach outperforms conventional binary relevance - partial least squar…
Algorithm identifies probability distributions from noisy moments with minimal samples.
This article investigates the eigenspectrum of the inner product-type kernel matrix under a binary mixture model in the high dimensional regime where the number of data and their dimension are both large and comparable. Based on…
The paper confirms two groups of gamma-ray bursts using a new nonparametric metric.
In regression tasks the distribution of the data is often too complex to be fitted by a single model. In contrast, partition-based models are developed where data is divided and fitted by local models. These models partition the input space and do not leverage the input-output dependency of multimodal-distributed data,…
We present a new approach for learning compact and intuitive distributed representations with binary encoding. Rather than summing up expert votes as in products of experts, we employ for each variable the opinion of the most reliable expert. Data points are hence explained through a partitioning of the variables into …
Method reveals dissimilarity in alloys' Curie temperatures.
Regularized linear regression improves binary classification performance, especially with ridge and regularization.
Transformers for binary decisions are sensitive to evidence order, leading to unreliable outcomes.
MMM model clusters mixed-type longitudinal data efficiently.