Archetypal analysis helps understand binary data sets.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new method for binary ICA using non-stationary sources.
Discovering causal relations among observed variables in a given data set is a major objective in studies of statistics and artificial intelligence. Recently, some techniques to discover a unique causal model have been explored based on non-Gaussianity of the observed data distribution. However, most of these are limit…
A new weighted MCC measure improves classifier performance evaluation.
Study on double descent behavior in two-layer neural networks for binary classification.
Study shows how adjusting for a binary proxy can bound causal effects.
Combines BO with context to optimize binary feedback.
Study identifies conditions for proxy adjustment in confounded binary treatment outcomes.
New method corrects skewed confidence for PbN classification.
A matrix completion problem, which aims to recover a complete matrix from its partial observations, is one of the important problems in the machine learning field and has been studied actively. However, there is a discrepancy between the mainstream problem setting, which assumes continuous-valued observations, and some…
This paper investigates whether the gravity model (GM) can explain the statistical properties of the International Trade Network (ITN). We fit data on international-trade flows with a GM specification using alternative fitting techniques and we employ GM estimates to build a weighted predicted ITN, whose topological pr…
Bayesian VAR model discovers Granger causality with uncertainty-aware binary graphs.
Parity calibration aims to predict increase-decrease events, not values.
New method extracts hidden phases in binary mixtures using tubular tilings.
Regularized linear regression improves binary classification performance, especially with ridge and regularization.
New binary loss functions improve density ratio estimation accuracy.
Archetypal analysis represents a set of observations as convex combinations of pure patterns, or archetypes. The original geometric formulation of finding archetypes by approximating the convex hull of the observations assumes them to be real valued. This, unfortunately, is not compatible with many practical situations…
In standard clustering problems, data points are represented by vectors, and by stacking them together, one forms a data matrix with row or column cluster structure. In this paper, we consider a class of binary matrices, arising in many applications, which exhibit both row and column cluster structure, and our goal is …
In this paper, we consider the matrix completion problem when the observations are one-bit measurements of some underlying matrix M, and in particular the observed samples consist only of ones and no zeros. This problem is motivated by modern applications such as recommender systems and social networks where only "like…
We consider analysis of relational data (a matrix), in which the rows correspond to subjects (e.g., people) and the columns correspond to attributes. The elements of the matrix may be a mix of real and categorical. Each subject and attribute is characterized by a latent binary feature vector, and an inferred matrix map…
New method learns signals from binary measurements, surpassing existing techniques.
A new method combines simple binary classifiers to build complex multiclass classifiers, achieving performance limits in a Gaussian setting.
This paper analyzes the training dynamics of binary neural networks using information bottleneck.
Discovering causal relations among observed variables in a given data set is a main topic in studies of statistics and artificial intelligence. Recently, some techniques to discover an identifiable causal structure have been explored based on non-Gaussianity of the observed data distribution. However, most of these are…
A gamma process dynamic Poisson factor analysis model is proposed to factorize a dynamic count matrix, whose columns are sequentially observed count vectors. The model builds a novel Markov chain that sends the latent gamma random variables at time as the shape parameters of those at time , which are linked …
The restricted Boltzmann machine is a graphical model for binary random variables. Based on a complete bipartite graph separating hidden and observed variables, it is the binary analog to the factor analysis model. We study this graphical model from the perspectives of algebraic statistics and tropical geometry, starti…
A new clustering algorithm for functional data using binary trees.
Unhinged loss minimization fails to improve classifier accuracy for simple data.
Two quandles from Coxeter groups studied, showing similarities in automorphism groups.
We present a scalable Bayesian model for low-rank factorization of massive tensors with binary observations. The proposed model has the following key properties: (1) in contrast to the models based on the logistic or probit likelihood, using a zero-truncated Poisson likelihood for binary data allows our model to scale …
We investigate the optimization of two probabilistic generative models with binary latent variables using a novel variational EM approach. The approach distinguishes itself from previous variational approaches by using latent states as variational parameters. Here we use efficient and general purpose sampling procedure…
BELIEF framework interprets GLMs using binary linear models.
Binary operations on algebras of observables are studied in the quantum as well as in the classical case. It is shown that certain natural compatibility conditions with the associative product imply the properties which usually are additionally required. In particular, it is proved that locality of a Loday bracket on s…
We provide a scheme for inferring causal relations from uncontrolled statistical data based on tools from computational algebraic geometry, in particular, the computation of Groebner bases. We focus on causal structures containing just two observed variables, each of which is binary. We consider the consequences of imp…
Binary Neural Networks (BNNs) have been garnering interest thanks to their compute cost reduction and memory savings. However, BNNs suffer from performance degradation mainly due to the gradient mismatch caused by binarizing activations. Previous works tried to address the gradient mismatch problem by reducing the disc…
Probit regression was first proposed by Bliss in 1934 to study mortality rates of insects. Since then, an extensive body of work has analyzed and used probit or related binary regression methods (such as logistic regression) in numerous applications and fields. This paper provides a fresh angle to such well-established…
We propose a probabilistic graphical model realizing a minimal encoding of real variables dependencies based on possibly incomplete observation and an empirical cumulative distribution function per variable. The target application is a large scale partially observed system, like e.g. a traffic network, where a small pr…
We give an explicit algorithm and source code for computing optimal weights for combining a large number N of alphas. This algorithm does not cost O(N^3) or even O(N^2) operations but is much cheaper, in fact, the number of required operations scales linearly with N. We discuss how in the absence of binary or quasi-bin…
Proposes a differentiable structure learning framework for general binary data.
Observational data hints at a finite universe, with spherical manifolds such as the Poincare dodecahedral space tentatively providing the best fit. Simulating the physics of a model universe requires knowing the eigenmodes of the Laplace operator on the space. The present article provides explicit polynomial eigenmodes…
ENTED efficiently decomposes binary and count tensors using nonparametric Gaussian processes.
Sensitivity analysis for individualized effects in OTRs with binary risk factors.
Proposes a method to classify binary data from multiple unlabeled datasets.
Binary data matrices can represent many types of data such as social networks, votes, or gene expression. In some cases, the analysis of binary matrices can be tackled with nonnegative matrix factorization (NMF), where the observed data matrix is approximated by the product of two smaller nonnegative matrices. In this …
BO algorithms improve binary and preferential optimization by distinguishing between types of uncertainty.
Estimates binary labels from dependent data using Markov Random Fields.
Proposes MRIV framework for unbiased CATE estimation using binary IVs.
RBMs model binary interactions with hidden node activation effects.