We consider analysis of relational data (a matrix), in which the rows correspond to subjects (e.g., people) and the columns correspond to attributes. The elements of the matrix may be a mix of real and categorical. Each subject and attribute is characterized by a latent binary feature vector, and an inferred matrix map…
Bayesian model for analyzing mixed data types.
problem Automatic exploratory analysis of heterogeneous datasets.
method General Bayesian non-parametric latent feature model.
result Automatic inference of model complexity and binary-valued latent features.
CSVAE learns interpretable latent subspaces for binary labels.
problem Learning interpretable latent representations correlated to specific labels.
method Conditional Subspace VAE (CSVAE) using mutual information minimization.
result CSVAE extracts interpretable latent subspaces for binary labels.
Generalizes latent feature models for mixed data types.
problem Lack of models for heterogeneous datasets with mixed data types.
method Bayesian nonparametric latent feature model for mixed data.
result Model automatically infers feature complexity and binary latent features.
The paper optimizes hyperplanes for binary classification in high-dimensional data with latent Gaussian mixtures.
problem Binary classification in high-dimensional data with latent Gaussian mixtures.
method Generalized least squares estimator for estimating the direction of the optimal separating hyperplane. Simple correction for intercept estimation.
result The procedure is minimax optimal in many scenarios and can retain the interpolation property.
Paper proposes a method to generate synthetic anomalies for robust anomaly detection.
problem Anomaly detection struggles with unbalanced data and rare anomalies.
method Two-level hierarchical latent space representation for feature distillation and synthesis.
result The method creates robust synthetic anomalies for training robust binary classifiers.
Latent feature models are attractive for image modeling, since images generally contain multiple objects. However, many latent feature models ignore that objects can appear at different locations or require pre-segmentation of images. While the transformed Indian buffet process (tIBP) provides a method for modeling tra…
Graphical models are commonly used tools for modeling multivariate random variables. While there exist many convenient multivariate distributions such as Gaussian distribution for continuous data, mixed data with the presence of discrete variables or a combination of both continuous and discrete variables poses new cha…
We propose a probabilistic model to infer supervised latent variables in the Hamming space from observed data. Our model allows simultaneous inference of the number of binary latent variables, and their values. The latent variables preserve neighbourhood structure of the data in a sense that objects in the same semanti…
Modern datasets are becoming heterogeneous. To this end, we present in this paper Mixed-Variate Restricted Boltzmann Machines for simultaneously modelling variables of multiple types and modalities, including binary and continuous responses, categorical options, multicategorical choices, ordinal assessment and category…
New model explains time-dependent latent factors in sensor data.
problem Understanding latent factors affecting sensor data over time.
method Developed new probabilistic models and inference methods.
result Models explain temporal dynamics of latent factors.
Model detects patterns in noisy binary data, explaining neuron activity in terms of cell assemblies.
problem Detecting structure in noisy or approximate repeats of patterns in sparse binary data.
method Probabilistic binary latent variable model based on Noisy-OR model, inferring sparse activity in latent variables.
result Model successfully extracts and explains latent structure in spiking neural data.
MaxMachine interprets product attributes in Amazon's large catalogue.
problem Predicting attribute applicability in large-scale product data with limited ground truth.
method Developed a probabilistic latent variable model that learns distributed binary representations.
result Improves over baseline in 17 out of 19 product groups.
Proposes a method to learn sparse and low-rank interactions in Ising models with latent variables.
problem Learning sparse interactions in Ising models with latent variables.
method Sparse + low-rank decomposition of Ising model parameters using convex regularized likelihood problem.
result Consistency properties in high-dimensional settings with growing number of variables and samples.
Gaussian CRFBC model for binary classification with latent variables.
problem Binary classification problems with undirected graphs.
method Gaussian conditional random fields with latent variables, local variational approximation, Newton-Cotes formulas.
result Improved prediction performance compared to unstructured predictors.
Paper extends FOFC algorithm to work with mixed data types.
problem Designing causal discovery algorithms for mixed data types.
method Proves tetrad constraint can be entailed for mixed data types and applies FOFC algorithm.
result FOFC algorithm can work on mixed data types.
ML4C uses binary classification to infer causal structures from latent vicinity.
problem Learning causal relations from observational data without ground truth.
method Two-phase paradigm with binary classifier and novel featurization.
result ML4C outperforms state-of-the-art algorithms in causal learning.
Paper proposes a novel spectral approach to learn binary latent variable models.
problem Learning binary latent variable models with hidden binary units in noisy data.
method Spectral approach based on eigenvectors of second and third order moment matrices.
result Consistently estimates model parameters at optimal rate under mild conditions.
Fourier analysis improves REINFORCE for binary models.
problem Improving gradient estimation for binary latent variable models.
method Connecting Fourier spectrum of Boolean functions to REINFORCE and developing low-variance unbiased gradient estimators.
result REINFORCE estimates degree-1 Fourier coefficients of a Boolean function.
We propose a new approach to inverse reinforcement learning (IRL) based on the deep Gaussian process (deep GP) model, which is capable of learning complicated reward structures with few demonstrations. Our model stacks multiple latent GP layers to learn abstract representations of the state feature space, which is link…
CausalEGM estimates causal effects by encoding confounders, improving performance in high-dimensional settings.
problem Challenges in estimating causal effects with high-dimensional confounders.
method CausalEGM framework using generative modeling to decouple confounders and estimate causal effects.
result CausalEGM outperforms existing methods in binary and continuous treatment settings, especially with large sample sizes and high-dimensional confounders.
Extends co-clustering to mixed numerical and binary data.
problem Co-clustering of mixed data types (numerical and binary).
method Latent block models for mixed data types.
result Effectiveness of the proposed approach on simulated data.
A Bayesian Boolean Matrix Factorization for cancer genomics
problem Identifying coordinated feature changes in cancer
method Bayesian Boolean Matrix Factorization
result Captures widespread, near-simultaneous chromosome-number changes
Symmetric binary matrices representing relations among entities are commonly collected in many areas. Our focus is on dynamically evolving binary relational matrices, with interest being in inference on the relationship structure and prediction. We propose a nonparametric Bayesian dynamic model, which reduces dimension…
Novel variational sampling improves generative model optimization.
problem Optimizing binary latent variable generative models efficiently.
method Truncated variational EM with efficient sampling.
result Efficiently increases variational free energy objective.
DisARM improves gradient estimation for binary latent variables.
problem Challenges in training models with discrete latent variables.
method Uses antithetic sampling over continuous augmentation.
result DisARM consistently outperforms ARM and baseline methods in log-likelihood and variance.
Direct optimization of binary latent VAEs achieves competitive results without sampling.
problem Training VAEs with discrete latent variables using standard methods is challenging.
method Applied evolutionary algorithms to directly optimize discrete latent distributions.
result Direct optimization is efficient and competitive in zero-shot learning.
Two binary Sine Cosine Algorithms improve feature selection in medical datasets.
problem Optimizing feature selection from medical datasets to enhance model accuracy.
method Proposed SBSCA and VBSCA algorithms using S-shaped and V-shaped transfer functions.
result SBSCA and VBSCA outperform four other binary optimization algorithms in medical datasets.
A new gradient estimator reduces variance near boundaries for binary latent variables.
problem Explosive gradient variance near boundaries in binary latent variable models.
method Introduces a new gradient estimator (bitflip-1) and an aggregated estimator (UGC) that uses either bitflip-1 or DisARM for each coordinate.
result UGC has uniformly lower variance than DisARM and achieves optimal optimization objectives.
Quantum circuits represent binary classification trees with binary features.
problem Classifying data using binary classification trees with binary features.
method Quantum circuits and probabilistic approach for traversing decision trees.
result First realization of a decision tree classifier on a quantum device.
Improves causal graph learning on dependent binary data.
problem Challenges in learning causal graphical models from dependent binary data.
method Decorrelation-based approach using latent utility model and EM-like algorithm.
result Significant improvement in accuracy of causal graph learning.
A common strategy for sparse linear regression is to introduce regularization, which eliminates irrelevant features by letting the corresponding weights be zeros. However, regularization often shrinks the estimator for relevant features, which leads to incorrect feature selection. Motivated by the above-mentioned issue…
A new method reduces variance in training discrete latent variable models.
problem High variance in stochastic gradient estimators for discrete latent variable models.
method Double control variates for score function estimators using Taylor expansions.
result Our method can have lower variance compared to other estimators.
Cubic predicts stock market indices by fusing stock latent embeddings and converting to binary classification.
problem Challenges in predicting stock market indices due to isolated time series treatment and simple regression.
method Fusion of stock latent embeddings, binary encoding classification, and confidence-guided prediction.
result Cubic outperforms state-of-the-art baselines in stock index prediction tasks.
PixelVAE++ improves generative models for natural images by combining VAE and PixelCNN.
problem Challenges in constructing powerful generative models for natural images.
method Introduces PixelVAE++, a VAE with three types of latent variables and a PixelCNN++ for the decoder, reusing a part of the decoder as an encoder.
result Achieves state-of-the-art performance on binary data sets and CIFAR-10.
New principle for disentangling latent factors using sparse regularization.
problem Disentangling latent factors from complex data.
method Sparse regularization of latent mechanisms to induce disentanglement.
result Recovery of latent variables up to permutation under certain conditions.
A new method for binary ICA using non-stationary sources.
problem Independent component analysis of binary data.
method Linear mixing model in latent space, followed by binary observation model with non-stationary sources.
result Proves non-identifiability with few observed variables but identifies with more variables.
Spectral method speeds fitting of binary time series models.
problem Modeling binary time series data with latent linear dynamical systems.
method Spectral learning of probit-Bernoulli latent linear dynamical systems.
result Spectral method provides robust, fixed-cost estimator.
New insights into BNN optimization redefine latent weights as inertia.
problem Optimizing Binarized Neural Networks (BNNs) with latent weights.
method Interpreted latent weights as inertia and introduced Binary Optimizer (Bop).
result Demonstrated improved performance of Bop on CIFAR-10 and ImageNet.
New hashing method improves document retrieval precision.
problem Efficiently retrieving similar documents from large text databases.
method Pairwise supervised hashing with Bernoulli VAE and unbiased gradient estimator.
result Superior performance compared to existing methods.
FSL-BM improves real-time classification with fuzzy logic and binary meta-features.
problem Real-time classification accuracy, memory consumption, and time complexity.
method FSL-BM integrates fuzzy logic, binary meta-features, Hamming Distance, and Hash function for efficient supervised learning.
result FSL-BM provides faster and more accurate real-time classification compared to existing algorithms.
Model predicts traffic speed using urban incidents.
problem Accurately predicting traffic speed in urban areas.
method Deep Incident-Aware Graph Convolutional Network (DIGC-Net).
result Model outperforms competing benchmarks in traffic speed prediction.
We propose a mixture of latent trait models with common slope parameters (MCLT) for model-based clustering of high-dimensional binary data, a data type for which few established methods exist. Recent work on clustering of binary data, based on a d-dimensional Gaussian latent variable, is extended by incorporating com…
A gamma process dynamic Poisson factor analysis model is proposed to factorize a dynamic count matrix, whose columns are sequentially observed count vectors. The model builds a novel Markov chain that sends the latent gamma random variables at time (t−1) as the shape parameters of those at time t, which are linked …
This study suggests replacing Ising distribution with Cox distribution.
problem Handling correlated binary data efficiently.
method Exploring conditions for replacing Ising distribution with Cox distribution as a latent variable model.
result The Ising distribution can be treated as a latent variable model with a quasi-normal distribution.
3D Adversarial Autoencoder learns compact binary descriptors from 3D point clouds.
problem Learning meaningful representations of 3D shapes for various tasks.
method End-to-end Adversarial Autoencoder model trained on 3D input and output.
result 3D Adversarial Autoencoder (3dAAE) generates state-of-the-art results for 3D points clustering and retrieval.
This paper addresses identifiability issues in HLAMs with hierarchical constraints.
problem Identifiability of HLAMs with hierarchical constraints.
method Developed sufficient and necessary identifiability conditions.
result Characterizes impacts of different attribute types in the graph on identifiability.
MeliusNet improves binary neural networks to match MobileNet-v1 accuracy.
problem Achieving high accuracy with binary neural networks on mobile devices.
method Alternating DenseBlocks and ImprovementBlocks to increase feature capacity and quality.
result MeliusNet matches MobileNet-v1 accuracy on ImageNet, improving binary network performance.