Socratic learning improves generative models by identifying and adapting to latent subsets in training data.
problem Lack of sufficient labeled training data for discriminative models.
method Uses feedback from a discriminative model to identify and adapt latent subsets in a generative model.
result Reduces error by up to 56.06% for relation extraction tasks compared to state-of-the-art techniques.
Paper finds a latent k−polytope in data efficiently.
problem Finding a latent k−polytope in data points. method Algorithm using subset smoothed polytope to estimate k−polytope. result Algorithm runs in O∗(k⋅extnnz) time, efficient even for sparse data. Subset LLDA improves scalability for large label sets in multi-label classification.
problem Scalability issues in Labeled Latent Dirichlet Allocation (LLDA) for large label sets.
method Subset LLDA, a simple variant of LLDA, addressing scalability issues.
result Subset LLDA outperforms LLDA and extreme multi-label classification algorithms on large label sets.
Algorithm uncovers latent attribute graph from molecular data.
problem Learning latent representations and interpreting them for limited data.
method Perturbation experiments on latent codes of a generative autoencoder.
result Effective graphical model of latent codes and attributes.
New method identifies latent causal graphs without parametric assumptions.
problem Identifying latent causal graphs without parametric assumptions.
method Constructive proofs with new graphical concepts.
result Conditions for nonparametric identification of latent causal graphs.
A novel ABC method for high-dimensional inverse problems using generative modeling and subset simulation.
problem Solving inverse-problems with high-dimensional inputs and expensive forward mappings.
method Joint deep generative modeling, Approximate Bayesian Computation (ABC) with Subset Simulation, and likelihood-free inference.
result Our method delivers promising performance without prior knowledge of the forward or noise distributions.
We study the problem of learning a latent tree graphical model where samples are available only from a subset of variables. We propose two consistent and computationally efficient algorithms for learning minimal latent trees, that is, trees without any redundant hidden nodes. Unlike many existing methods, the observed …
Extends linear structural causal models to include deterministic relations and latent confounders for causal discovery.
problem Causal discovery in linear SCMs with deterministic relations and latent confounders.
method Extended existing results to include deterministic relations and latent confounders, derived necessary and sufficient conditions for unique identifiability, proposed an algorithm for recovery.
result First work on identifiability results for causal discovery under latent confounding and deterministic relationships.
GCAE uses density estimation to achieve reliable disentanglement in latent space.
problem Disentangled learning representations suffer from reliability issues.
method GCAE uses Gaussian Channel Autoencoder with Dual Total Correlation (DTC) to avoid the curse of dimensionality.
result GCAE achieves highly competitive and reliable disentanglement scores.
New method identifies causal variables from partially observed data.
problem Learning from unpaired observations with instance-dependent partial observability.
method Proposes two methods enforcing sparsity in the inferred representation.
result Establishes two identifiability results for linear and piecewise linear mixing functions.
New method predicts unobserved interactions between sets of elements.
problem Limited access to interactions between sets of elements.
method Generalized Synthetic Interventions (GSI) estimator.
result GSI estimator outperforms existing methods on synthetic and real data.
Latent tree models are used in various fields like phylogenetics and computer vision.
problem Representing and analyzing complex data structures with latent variables.
method Graphical models defined on trees, focusing on tree metrics.
result Latent tree models encompass various well-known models and contain fundamental limits of what can be learned.
This paper uses deep neural networks for one-class classification by splitting normal data into typical and atypical subsets.
problem Training deep neural networks with only one class of data for one-class classification.
method Intra-class splitting to create typical and atypical subsets, using binary loss and auxiliary subnetworks.
result The method outperformed seven baselines and had comparable performance to state-of-the-art methods on image datasets.
Confidence-based filtering reveals latent structure in diffusion models.
problem Unclear latent structure in diffusion models.
method Confidence scores from a classifier.
result Class-relevant latent structure emerges under confidence-based filtering.
Unified framework for gradient estimation in combinatorial spaces.
problem Scaling relaxed gradient estimators to large combinatorial distributions.
method Introducing stochastic softmax tricks within the perturbation model framework.
result Stochastic softmax tricks improve model performance and discover more latent structure.
We consider learning a causal ordering of variables in a linear non-Gaussian acyclic model called LiNGAM. Several existing methods have been shown to consistently estimate a causal ordering assuming that all the model assumptions are correct. But, the estimation results could be distorted if some assumptions actually a…
Study identifies influences in VAR models with latent processes.
problem Identify influences among observed and latent processes in VAR models.
method Identify support of transition matrix and lengths of latent paths.
result Support of transition matrix and lengths of latent paths can be identified successfully under certain conditions.
The Gaussian process latent variable model (GP-LVM) is a popular approach to non-linear probabilistic dimensionality reduction. One design choice for the model is the number of latent variables. We present a spike and slab prior for the GP-LVM and propose an efficient variational inference procedure that gives a lower …
SubTab turns tabular data into a multi-view problem for better representation learning.
problem Lack of structure in tabular data makes it hard to apply effective self-supervised learning methods.
method Divides tabular features into subsets and uses autoencoder-like reconstruction for collaborative inference.
result SubTab achieves state-of-the-art performance on tabular datasets, matching or surpassing CNN-based methods.
Efficiently searches ancestral graphs using multivariate information.
problem Discovering causal relationships in graphs with latent variables.
method Greedy search-and-score algorithm with two-step approach.
result Outperforms existing methods on benchmark datasets.
Identifying latent structure in large data matrices is essential for exploring biological processes. Here, we consider recovering gene co-expression networks from gene expression data, where each network encodes relationships between genes that are locally co-regulated by shared biological mechanisms. To do this, we de…
Novel graphical models for time series with latent confounders improve causal inference.
problem Causal relationships and independencies in multivariate time series with unobserved confounders.
method Introduced a novel class of graphical models and characterized their properties.
result Novel graphs provide stronger causal inferences without additional assumptions.
HCL learns shared and modality-specific latent representations for multimodal data.
problem Binary shared-private decomposition inadequately represents shared information across subsets of modalities.
method Hierarchical Contrastive Learning framework combining latent-variable formulation, structural sparsity, and contrastive objective.
result HCL accurately recovers hierarchical structure and improves predictive performance on multimodal data.
New method identifies latent causal variables from observed data, overcoming indeterminacies.
problem Identifying latent causal variables from observed data, especially when latent variables are weight-variant.
method Introduces a novel identifiability condition for latent causal models, proposing SuaVE method.
result Identifies latent causal variables up to trivial permutation and scaling, demonstrating consistency and efficacy.
Sparse VAE learns latent factors from high-dimensional data.
problem Unsupervised representation learning on high-dimensional data.
method Sparse VAE model that learns latent factors summarizing data associations.
result Sparse VAE can recover true model parameters with infinite data.
New model explains time-dependent latent factors in sensor data.
problem Understanding latent factors affecting sensor data over time.
method Developed new probabilistic models and inference methods.
result Models explain temporal dynamics of latent factors.
Proposes a differentiable hypergeometric distribution for learning group importance.
problem Learning the sizes of subsets in applications like clustering and weakly-supervised learning.
method Introduces a reparameterizable hypergeometric distribution to model group sizes and learn their relative importance.
result Outperforms previous methods in weakly-supervised learning and clustering.
A new gradient estimator reduces variance near boundaries for binary latent variables.
problem Explosive gradient variance near boundaries in binary latent variable models.
method Introduces a new gradient estimator (bitflip-1) and an aggregated estimator (UGC) that uses either bitflip-1 or DisARM for each coordinate.
result UGC has uniformly lower variance than DisARM and achieves optimal optimization objectives.
T-JEPA learns tabular data representations without augmentations, outperforming traditional methods.
problem Challenges in self-supervised learning for tabular data due to lack of data augmentations.
method T-JEPA uses a Joint Embedding Predictive Architecture (JEPA) to predict latent representations of different subsets of features within the same sample.
result Significant improvement in classification and regression tasks, outperforming traditional methods.
Latent factor models are the canonical statistical tool for exploratory analyses of low-dimensional linear structure for an observation matrix with p features across n samples. We develop a structured Bayesian group factor analysis model that extends the factor model to multiple coupled observation matrices; in the cas…
Unified framework for accurate coresets in latent variable models and regularized regression.
problem Efficiently training models on large datasets.
method Unified framework for constructing accurate coresets for latent variable models and ℓp-regularized regression. result Unified framework reduces coreset size for latent variable models and ℓp-regularized regression. Study the averaging estimator on graphs with labeled nodes.
problem Understanding the quality of averaging estimators on graph data.
method Rigorously study concentration properties, variance bounds, and risk bounds.
result Contributes to theoretical understanding of graph learning.
A new model speeds up MRF learning from small datasets.
problem Intractable MRF learning and high computational cost.
method Characterized MRF subset with Lattice, Homogeneity, and Inertia; designed a non-Markov model.
result Learning algorithm is much faster (O(U log U) vs. general-purpose MRF's time complexity).
Improves change-point detection for high-dimensional time-series.
problem Uncertainty in latent variable estimation affects change-point detection.
method Proposes multinomial sampling to improve detection rate and reduce delay.
result Results outperform baseline method in experiments.
The paper analyzes unsupervised learning using contrastive methods and introduces a theoretical framework.
problem Learning useful feature representations from unlabeled data.
method Introduces latent classes and contrasts similar vs. non-similar data points.
result Proves guarantees on the performance of learned representations on downstream tasks.
The study addresses negative transfer in multi-output Gaussian processes by proposing latent structures.
problem Negative transfer in multi-output Gaussian processes leading to decreased performance.
method Defining negative transfer, deriving conditions for avoiding it, proposing latent structures.
result Latent structures can avoid negative transfer and scale to large datasets.
In latent Dirichlet allocation (LDA), topics are multinomial distributions over the entire vocabulary. However, the vocabulary usually contains many words that are not relevant in forming the topics. We adopt a variable selection method widely used in statistical modeling as a dimension reduction tool and combine it wi…
Proposes a flexible method for learning latent causal representations.
problem Limited applicability of existing causal representation learning methods.
method Imposes constraints on function classes and relaxes identifiability conditions.
result Establishes partial identifiability results under weaker conditions.
Efficient algorithms decide algebraic constraints of causal graphs.
problem Distinguish causal graphs with latent confounders.
method Study algebraic constraints and propose efficient algorithms.
result Decide equivalence or subset of algebraic constraints.
ATLAS separates invariant and transferable latent factors across diverse environments.
problem Transfer learning and robust prediction in heterogeneous environments.
method ATLAS leverages invariance principle to disentangle latent factors and uses auxiliary labels for robust prediction.
result Near-oracle performance and robust transferable prediction in new environments.
New method identifies stable latent variables across different domains using weak distributional invariances.
problem Learning causal representations for multi-domain datasets.
method Autoencoders incorporating weak distributional invariances.
result Autoencoders can identify stable latent variables across different domains.
Extends random dot product graph model to handle multiple graphs.
problem Modeling and analyzing multiple graphs with shared nodes.
method Jointly embed adjacency matrices into a latent space.
result Node representations converge to latent positions with Gaussian error.
Conditions for uniquely identifying parameters of deep ReLU networks.
problem Characterizing networks whose parameters can be uniquely identified.
method Conditions on deep fully-connected feedforward ReLU neural networks.
result Parameters of the network are uniquely identified under certain conditions.
New framework identifies causal models with arbitrary interventions, improving realism.
problem Identify causal models with realistic interventions.
method Theoretical framework for identifying causal models with arbitrary interventions.
result Identify causal models with arbitrary interventions, up to a higher-level abstraction.
New algorithm learns causal graph to minimize regret in bandits without full structure.
problem Learning optimal decisions in bandits with unknown causal graph and latent confounders.
method Two-stage approach: first learns ancestors and necessary confounders, second applies standard bandit algorithm.
result No full causal structure needed for optimal decisions; only necessary confounders are crucial.
Approach for training deep nets with unlabeled patches.
problem Training deep neural networks with detailed expert annotations.
method Cluster-based learning from weakly labeled bags in latent space.
result Improved performance on Camelyon dataset.
Given a set of experiments in which varying subsets of observed variables are subject to intervention, we consider the problem of identifiability of causal models exhibiting latent confounding. While identifiability is trivial when each experiment intervenes on a large number of variables, the situation is more complic…
CAE models learn complex manifold structures in data.
problem Flat latent spaces in auto-encoders fail to capture manifold structures.
method Proposes Chart Auto-Encoders (CAE) with a multi-chart latent space.
result CAE provides better data representation with manifold properties.