New approach learns shared architecture for multi-task learning.
problem Finding optimal shared layers, weights, and task losses in MTL.
method Latent multi-task architecture learning that jointly addresses sharing, weights, and task losses.
result Consistently outperforms previous approaches to multi-task learning, achieving up to 15% error reduction.
Researchers analyze neural process architectures and their representational capacities.
problem Understanding what functions can be represented by different neural process architectures.
method Analyzing four types of neural process architectures: CNPs, ANPs, TNPs, and their latent variants.
result Prove these architectures form a strict hierarchy and characterize their representational capabilities.
CrossNet uses cross-consistency to improve unpaired image translation.
problem Image-to-image translation without paired data.
method Introduced a novel architecture with cross-translators and latent cross-consistency constraints.
result CrossNet outperforms state-of-the-art on various image translation tasks.
Study reveals latent state computation in stochastic volatility models.
problem Understanding latent stochastic dynamics in noisy, partially observed observations.
method Multivariate stochastic volatility setting, controlled experiments on various architectures.
result Evidence of a two-stage computation: latent state encoding and output head mapping.
STCN combines TCNs with stochastic latent variables for sequence modeling.
problem Performance gap between TCNs and stochastic RNNs, especially with multiple layers of random variables.
method Proposes a hierarchy of stochastic latent variables in a modular architecture.
result Achieves state-of-the-art log-likelihoods across various tasks.
New framework models neural systems with random architecture on manifolds.
problem Complex, uncertain systems with non-Gaussian outputs.
method Latent random field on compact manifold generates neural architecture and weights.
result Synthetic neural systems can produce stochastic outputs for deterministic inputs.
Model learns cancer tissue images onto a low-dimensional space revealing tissue characteristics.
problem Improving cancer diagnosis through high-fidelity digital pathology.
method Deep generative model using PathologyGAN to map real images onto a latent space.
result Latent space encodes morphological characteristics and reveals distinct tissue clusters.
A fair autoencoder removes sensitive factors while preserving useful latent representations.
problem Learning fair representations that ignore sensitive factors.
method Variational autoencoder with priors and MMD penalty.
result The method effectively removes unwanted sources of variation.
A new architecture improves VAE disentanglement without explicit supervision.
problem Learning disentangled representations in VAEs.
method Spatial Broadcast Decoder: tiling latent vector, concatenating coordinates, fully convolutional network.
result Improves disentangling, reconstruction accuracy, and generalization.
GLSR-VAE enhances VAE latent space for better attribute control.
problem Fine-tuning VAE latent space for continuous data attributes.
method Geodesic Latent Space Regularization (GLSR) for VAEs.
result Controls latent space changes to reflect data attributes.
FisherNet extends Autoencoder using Fisher information for better data reconstruction.
problem Data reconstruction accuracy and model scalability in high-dimensional latent spaces.
method Introduces FisherNet architecture that uses Fisher information to quantify and account for latent space uncertainty.
result FisherNet produces more accurate reconstructions and scales better with latent space dimensions compared to VAE.
Latent MoS learns multiple symmetries for efficient dynamic learning.
problem Efficiently learning dynamics from limited system measurements.
method Latent Mixture of Symmetries (Latent MoS) with hierarchical architecture.
result Latent MoS outperforms baselines in interpolation and extrapolation tasks.
A graph VAE framework optimizes neural architectures in a continuous space.
problem Discovering efficient neural architectures in a discrete space.
method Graph VAE framework with VAE and GNN components, joint learning of predictors and decoders.
result The framework discovers powerful neural architectures with both excellent performance and high computational efficiency.
New theory explains how noisy, high-dimensional data can still lead to robust predictions.
problem Modern machine learning models achieve high performance with noisy, high-dimensional data.
method Synthesizes principles from Information Theory, Latent Factor Models, and Psychometrics to clarify predictive robustness.
result Predictive robustness arises from data architecture and model capacity, not just data cleanliness.
Researchers analyze and improve latent space in NAR models.
problem Latent space resolution and range issues in NAR models.
method Detailed analysis of GNN latent space structure, proposing and testing solutions.
result Improvements in majority of algorithms on CLRS-30 benchmark.
A novel circuit motif uses sister cells for inference with correlated priors.
problem Structured priors in neural systems pose architectural challenges.
method Proposes a novel circuit motif using sister cells to implement correlated priors without direct interactions.
result Demonstrates the efficacy of correlated priors for inference in noisy environments.
Study neural architectures on learned latent graphs using Schrödinger dynamics.
problem Understanding neural architectures on learned latent graphs.
method Optimizes over stratified moduli space of weighted graphs with Kähler-Hessian metric.
result Multilayer stationary networks are equivalent to global stationary problems on supra-graphs.
New method controls posterior collapse in VAEs without network architecture constraints.
problem Posterior collapse in VAEs reduces diversity of generated samples.
method Introduces Latent Reconstruction (LR) loss to control posterior collapse.
result Controls posterior collapse on various datasets without architectural constraints.
Stochastic WaveNet models sequential data with latent variables and dilated convolutions.
problem Modeling distribution of sequential data like speech and motions.
method Combines stochastic latent variables and dilated convolutions in WaveNet architecture.
result Obtains state-of-the-art performances on speech and handwriting datasets.
New model generates realistic single-cell gene expression data.
problem Generating realistic single-cell gene expression profiles is challenging.
method scLDM, a latent diffusion model using Diffusion Transformers and linear interpolants.
result Superior performance in generating realistic single-cell gene expression data.
Deep learning uses complex networks for high-dimensional data.
problem Computational inefficiency in training deep learning models.
method Use of hierarchical latent variables, efficient linear algebra, SGD optimization, and batch sampling.
result Efficient training and inference possible with optimized algorithms.
Model compresses event-like contexts using gated surprise signals.
problem Perceiving a dynamic world as organized events.
method Hierarchical, surprise-gated recurrent neural network architecture.
result Achieves best performance on multiple event processing tasks.
Unsupervised learning of latent object properties from interactions.
problem Discovering latent physical properties of objects without labeled data.
method Perception-Prediction Network (PPN) framework.
result PPN can accurately simulate and translate latent object properties.
Deep Discrete Encoders (DDEs) tackle interpretable generative models for rich data with discrete latent layers.
problem Overparametrized, non-identifiable, and uninterpretable deep generative models in high-stakes applications.
method Directed graphical model with multiple binary latent layers, transparent identifiability conditions, scalable estimation pipeline.
result Transparent identifiability conditions and scalable estimation pipeline for interpretable DDEs.
SEMASIA provides a large dataset of latent representations for model comparison.
problem Difficulty in comparing semantic structures across different neural network models.
method Collection of latent representations from 1700 pretrained models across various benchmarks.
result Consistent semantic organization across models and datasets.
Unsupervised framework captures acquisition variability in structural connectomes.
problem Acquisition differences across sites, scanners, and protocols complicate structural connectome analysis.
method An unsupervised framework using architectural annealing to balance discrete and continuous latent variables.
result Architectural annealing produces stronger site learning than baseline models.
Future autonomous systems need reliable world models and complex action sequences.
problem Current automated systems lack reliable world models and complex action sequences.
method Introduce energy-based and latent variable models combined in a hierarchical joint embedding predictive architecture (H-JEPA).
result Combining energy-based and latent variable models in H-JEPA can lead to reliable world models and complex action sequences.
Improved generalization in abstract reasoning tasks using disentangled latent representations.
problem Improving generalization in unsupervised representation learning for abstract reasoning.
method Used disentangled VAEs to learn latent representations from relational reasoning problems.
result Disentangled latent representations outperform supervised learning in generalization.
Efficient neural network invariant to symmetry subgroups.
problem Designing neural networks invariant to symmetry subgroups for computational efficiency.
method A new G-invariant transformation module and multi-layer perceptron. result The proposed architecture is computationally and memory efficient, and universal.
Pixel-space diffusion models outperform latent models on high-resolution image synthesis.
problem Efficiency and quality trade-off in high-resolution image synthesis.
method Sigmoid loss-weighting, simplified architecture, and resolution scaling.
result Achieved 1.5 FID on ImageNet512, new SOTA results on other datasets.
New autoencoder learns structured representations without regularization.
problem Learning structured representations without relying on regularization.
method Proposes a novel autoencoder architecture that learns a hierarchy of latent variables.
result Improves results in generation, disentanglement, and extrapolation tasks.
Generative model improves latent space convexity through adversarial training on interpolations.
problem Improving latent space convexity in generative models.
method Adversarial training on latent space interpolations within an AE-GAN architecture.
result Convex latent distribution of generated images, preserving realistic resemblances.
Retina-VAE models macular disease spectrum using clinical data.
problem Representing the spectrum of macular diseases clinically.
method Variational autoencoder (VAE) model trained on patient profiles.
result Latent vectors cluster into 14 subtypes, suggesting treatment responses.
New approaches improve uncertainty quantification in autoregressive models for sequence data.
problem Uncertainty quantification in autoregressive models for exchangeable sequences.
method Study of inferential and architectural biases for autoregressive models, focusing on multi-step inference.
result Custom architectures are necessary for multi-step inference to ensure exchangeability.
Paper improves deep learning models for cardiac potential reconstruction.
problem Improving generalization of sequence models for cardiac potential reconstruction.
method Constrained stochasticity and global aggregation of temporal information in latent space.
result Improved generalization of inverse reconstruction networks.
CGNP embeds functional processes into latent vectors using graph neural networks.
problem Embedding and sampling functional processes over arbitrary domains.
method CGNP employs graph neural networks to embed and decode functional processes.
result CGNP effectively samples encoded functions over any domain.
New findings show disentangled latent representations are not enough for robust compositional generalization.
problem Deep learning models struggle with compositional generalization, especially in out-of-distribution samples.
method Investigated a 2D Gaussian generation task with fully disentangled inputs, then forced disentangled latent representations into full-dimensional output space.
result Forcing disentangled latent representations into full-dimensional output space enables robust compositional generalization.
InvGAN combines generative and inference models for photo-realistic image manipulation.
problem GANs lack an inference model for image editing and downstream tasks.
method Train inference and generative models together to adapt and converge.
result InvGAN embeds real images into a high-quality generative model's latent space.
Study improves interpretability in generative models by disentangling latent variables in scientific datasets.
problem Extracting generative factors from complex, high-dimensional datasets in unsupervised or semi-supervised settings.
method Introducing Aux-VAE, a novel architecture within the VAE framework, which disentangles latent variables by guiding them with auxiliary variables.
result Aux-VAE achieves disentanglement with minimal modifications to the standard VAE loss function, validated on multiple datasets.
New metric improves latent space distances and interpolants.
problem Distortion in latent space of deep generative models.
method Characterized latent space curvature with stochastic Riemannian metric.
result Significant improvement in distances and interpolants under new metric.
A neural network model tackles high-dimensional data with latent structures.
problem Modeling high-dimensional data with latent low-dimensional structures.
method Integrates PCA and Soft PCA layers into neural network architecture for factor modeling and non-linear transformations.
result Demonstrates improved performance in forecasting and nowcasting with real-world data.
A new generator uses kernel distance to avoid GAN weaknesses.
problem Stability and mode collapse in GANs and autoencoders.
method LCW generator (Latent Cramer-Wold generator) using kernel distance.
result Very competitive FID values.
We develop VAE-DLM for dynamics with geometric flows in latent space.
problem Learning latent geometric properties for dynamics in high-dimensional data.
method Riemannian approaches to VAEs with a geometric flow in latent space, reformulating ELBO loss.
result Improved performance and robust learning for external dynamics, reducing OOD error.
VAE improves protein dynamics modeling by maximizing latent autocorrelation.
problem Modeling slow processes in protein dynamics.
method Maximizing latent autocorrelation in VAE architecture.
result VDE framework optimizes latent coordinate for slow processes.
A new neural network models discourse relations with latent variables.
problem Jointly modeling discourse relations and word sequences.
method Latent variable recurrent neural network for discourse relations.
result Model outperforms state-of-the-art alternatives on discourse classification tasks.
SUMO provides unbiased log marginal likelihood estimation for latent variable models.
problem Biased estimates of log marginal likelihood in latent variable models.
method Randomized truncation of infinite series for unbiased estimation.
result Models trained with SUMO give better test-set likelihoods than standard methods.
LatentTrack generates model parameters online for nonstationary data.
problem Online probabilistic prediction under nonstationary dynamics.
method Sequential neural architecture with latent filtering and amortized inference.
result Consistently lower negative log-likelihood and mean squared error than baselines.
GD-VAEs learn dynamics from observations using geometric and topological information.
problem Learning parsimonious representations of nonlinear dynamics from observations.
method Develops data-driven methods incorporating geometric and topological information using Variational Autoencoders (VAEs).
result GD-VAEs provide methods for learning reduced dimensional representations of nonlinear dynamics.