Plug-and-play multimodal controller improves class-conditional image generation.
problem Generating class-conditional images from user-specified labels.
method Introduces a `multimodal controller` to generate multimodal data without additional learning parameters.
result Multimodal controlled generative models produce higher quality class-conditional images and novel modalities.
Paper learns multimodal transition dynamics for RL using conditional VI.
problem Learning stochastic, multimodal transition dynamics in RL.
method Conditional variational inference (VI) for complex stochasticity.
result VI successfully predicts multimodal outcomes but ignores deterministic parts.
The paper tackles conditional multimodal learning using variational methods.
problem Learning conditional distributions between modalities.
method Variational methods for maximizing conditional log-likelihood.
result Generated faces are more representative of the attributes.
Unified approach for multimodal data prediction using synthetic data generation.
problem Challenges in integrating heterogeneous data types for accurate predictive performance.
method Generative Distribution Prediction (GDP) framework that uses multimodal synthetic data generation.
result Empirical validation across four tasks demonstrates versatility and effectiveness of GDP.
This paper proposes a model to learn multimodal representations robust to missing data.
problem Learning multimodal representations from heterogeneous sources of information.
method Optimizes a joint generative-discriminative objective across multimodal data and labels, factorizing representations into multimodal discriminative and modality-specific generative factors.
result The proposed model achieves state-of-the-art performance on six multimodal datasets and can reconstruct missing modalities without significant performance drop.
Graph Mixture Density Networks model multimodal data on graphs.
problem Challenging conditional density estimation problems with structured data.
method Combining mixture models and graph representation learning.
result Significant improvement in likelihood of epidemic outcomes.
This paper strengthens the computational separation between multimodal and unimodal learning, showing unimodal learning is hard on typical instances.
problem Theoretical justification for empirical success of multimodal machine learning.
method Introduced a stronger average-case computational separation between unimodal and multimodal learning.
result For typical instances, unimodal learning is computationally hard, while multimodal learning is easy.
PVAE learns disentangled representations from multimodal data.
problem Learning disentangled representations from multimodal sensory data.
method Partitioned Variational Autoencoder (PVAE) with multimodal generative model and training objectives.
result PVAE achieves over 99% accuracy on both modalities for semantic units.
Improves VAEs for generating data from mixed distributions.
problem Inability of VAEs to generate from individual data modalities.
method Conditional Prior VAE (CP-VAE) with two-level generative process.
result Generations from individual mixture components of multimodal data.
Simple baselines improve multimodal utterance learning.
problem Learning rich multimodal utterance representations.
method Conditional factorization of utterances into unimodal factors; extending to bimodal and trimodal factors.
result Optimal embeddings can be derived in closed form.
New model improves multimodal autoencoders by learning joint and conditional distributions.
problem Limitations in recent multimodal autoencoders restrict their quality on complex datasets.
method Proposes a multistage training process with variational inference and Normalizing Flows, leveraging shared modality information.
result Achieves state-of-the-art results on benchmark datasets.
A framework for uncertainty-aware multimodal learning using conformal Shapley intervals.
problem Uncertainty and modality level importance in multimodal learning.
method Introduces conformal Shapley intervals to quantify modality level importance and uncertainty.
result Demonstrates meaningful uncertainty quantification and strong predictive performance.
New method for risk allocation under multimodality of loss distribution.
problem Risk assessment under multimodal conditional loss distribution.
method Maximum Likelihood Allocation (MLA) and multimodality adjustment.
result Multimodality adjustment improves soundness of risk allocations.
MDNs offer a data-efficient alternative to diffusion and flow models for multimodal scientific learning.
problem Capturing multimodal conditional uncertainty in scientific inverse problems.
method Mixture Density Networks (MDNs) as explicit parametric density estimators.
result MDNs achieve superior generalization, interpretability, and sample efficiency in scientific tasks.
Generative Score Inference improves uncertainty quantification for multimodal data.
problem Accurate uncertainty quantification in multimodal learning tasks.
method Generative Score Inference (GSI) uses synthetic samples to approximate conditional score distributions.
result GSI achieves state-of-the-art performance in hallucination detection and image captioning uncertainty estimation.
GGMPs improve non-Gaussian conditional density estimation.
problem Multimodality, heteroscedasticity, and strong non-Gaussianity in conditional density estimation.
method GGMP combines local Gaussian mixture fitting, cross-input component alignment, and per-component heteroscedastic GP training.
result GGMPs improve distributional approximation on synthetic and real-world datasets.
A new mathematical framework for multimodal learning.
problem Linking different data modalities for better understanding.
method Interpreting contrastive learning as optimizing encoders for conditional probability distributions.
result Novel probabilistic loss functions and metrics for alignment in latent space.
Develops a deep multimodal classifier for medical images and reports.
problem Challenges of small datasets in medical deep learning.
method Transfer learning, integrated gradients, unsupervised clustering, tuning parameter.
result Improves classification accuracy by 4% and 7% on average.
New method handles missing data and multiple data types in time series models.
problem Handling missing data and multiple data modalities in time series models.
method Factorized inference method for Multimodal Deep Markov Models (MDMMs).
result Method performs well even with high levels of missing data and outperforms existing approaches.
Deep model predicts shapes of curves with multiple covariates.
problem Predicting shapes of planar curves with various covariates.
method Deep learning model using complex-valued functions, conditional covariance smoother with modality-specific encoders.
result Model accurately predicts shapes of curves with multimodal covariates.
Model learns multiple tasks using visual and textual representations.
problem Training visual navigation agents for multiple tasks.
method Dual-Attention unit for task-invariant alignment of visual and textual representations.
result Model outperforms baselines on semantic goal navigation and embodied question answering.
Paper proposes Seq2Seq models for multimodal sentiment analysis.
problem Learning representations from multiple modalities in machine learning.
method Two unsupervised Seq2Seq models for multimodal sentiment analysis.
result Seq2Seq models improve F1 Score by twelve points in Bimodal sentiment analysis.
Study analyzes bias interactions in multimodal models using simulation-based methods.
problem Analyzing dynamic bias interactions in multimodal models to ensure fairness and equity.
method Simulation-based heuristic approach to compute bias scores for text-only, image-only, and multimodal embeddings.
result Multimodal bias interactions can be amplification, mitigation, or neutral, with text bias often dominant.
TCT learns multimodal sequence representations by translating from related sequences.
problem Challenges in learning semantic representations from multimodalities.
method Transformer based Cross-modal Translator (TCT) combined with Multimodal Transformer Network (MTN).
result Proposed method achieves new state-of-the-art performance on video-grounded dialogue.
Develops a contrastive framework for data-efficient multimodal learning.
problem Expensive training of multimodal generative models requiring related multimodal data.
method Contrastive framework for multimodal learning, distinguishing related from unrelated data.
result Data-efficient multimodal learning on challenging datasets for various VAE models.
FMT model improves multimodal sequential learning across language, vision, and acoustic data.
problem Modeling spatio-temporal dynamics across multiple modalities.
method Factorized Multimodal Transformer (FMT) that models intramodal and intermodal dynamics in a factorized manner.
result FMT outperforms existing models on 3 datasets and 21 labels, setting new state of the art.
CQNPs enhance predictive performance and distribution modeling using quantile regression.
problem Limited predictive likelihood of Gaussian models for complex distributions.
method Introducing Conditional Quantile Neural Processes (CQNPs) that focus on estimating informative quantiles.
result Significant improvements in predictive performance and better modeling of multimodal distributions.
New method identifies shared components from unpaired multimodal mixtures.
problem Identify shared components from unpaired multimodal mixtures.
method Distribution divergence minimization-based loss with sufficient conditions for identifiability.
result Sufficient conditions for shared component identifiability from unaligned multimodal mixtures.
Enhances multimodal generation with Normalizing Flows and correlation analysis.
problem Generating coherent cross-modal data from multiple sources.
method Uses Deep Canonical Correlation Analysis for shared information, Normalizing Flows for diversity, and Product of Experts for scalability.
result Improves likelihood, diversity, and coherence in conditional generation.
Multimodal bitransformer boosts image-text classification.
problem Combining text and image modalities for improved classification.
method Supervised multimodal bitransformer model integrating text and image encoders.
result State-of-the-art performance on multimodal classification benchmarks.
Dp-CLIP preserves privacy in multimodal AI training.
problem Privacy concerns in multimodal AI, especially in vision-language tasks.
method Differentially private adaptation of CLIP model.
result Dp-CLIP retains accuracy while ensuring privacy.
Survey of multimodal deep generative models for diverse data types.
problem Inference of shared representations and cross-modal generation from heterogeneous multimodal data.
method Variational autoencoders and other deep generative models.
result A comprehensive survey of multimodal deep generative models.
Paper proposes multimodal contrastive learning for EHR data.
problem Separate treatment of structured and unstructured EHR data.
method Proposes a multimodal feature embedding generative model and a multimodal contrastive loss.
result Multimodal learning yields better feature representation than single-modality learning.
EmbraceNet improves robustness in multimodal classification.
problem Robust multimodal classification with partial data loss.
method Deep learning architecture for multimodal fusion.
result EmbraceNet outperforms other models in partial data scenarios.
Framework translates images between domains without supervision.
problem Challenges in unsupervised image-to-image translation, especially handling multimodality.
method Proposes a Multimodal Unsupervised Image-to-Image Translation (MUNIT) framework, decomposing images into content and style codes.
result Demonstrates improved generation of diverse outputs from a single source image.
Paper proposes a harmonized approach to multimodal learning using GPLVMs.
problem Modality heterogeneity in multimodal data.
method Develops a novel learning scheme called Harmonization to jointly learn latent model parameters from different modalities.
result Experimental results show superior performance in cross-modal retrieval tasks.
New results show contrastive learning can recover shared factors in multimodal data.
problem Understanding when contrastive learning can recover shared latent factors in multimodal data.
method New identifiability results for multimodal contrastive learning, distinguishing between multi-view and multimodal settings.
result Contrastive learning can block-identify shared latent factors in multimodal data, even with dependencies.
Improved sentiment analysis with multimodal data.
problem Cross-modal sentiment analysis in social media, customer service, and video blogs.
method Gated mechanism for attention-based learning of cross-modal interactions, with experiments on CMU-MOSI and CMU-MOSEI datasets.
result 1.6% and 1.34% absolute improvement over state-of-the-art.
New method uses MRI data to improve PET tomography uncertainty quantification.
problem Improving uncertainty quantification in emission tomography with multimodal data.
method Nonparametric posterior learning technique adapted for Poisson-type data.
result Sampling algorithms are scalable, parallelizable, and easy to implement.
Improved multimodal learning with Gated Multimodal Units.
problem Finding an intermediate representation from multiple data sources.
method Gated neural networks for multimodal fusion.
result GMU outperformed single-modality approaches and other fusion strategies.
New model uses financial filings to predict bankruptcy, even without MDA sections.
problem Lack of complete MDA data limits traditional bankruptcy prediction models.
method Conditional Multimodal Discriminative (CMMD) model learns from accounting, market, and textual data.
result Empirical results show superior classification performance compared to traditional models.
Improved software flaw detection using NAS on multimodal DL models.
problem Software flaw detection in multimodal deep learning models.
method Adapted NAS framework for multimodal learning, combined with multimodal deep learning models.
result Improved performance on the Juliet Test Suite.
Study quantifies interactions between unlabeled multimodal data.
problem Understanding how modalities combine in semi-supervised settings.
method Information-theoretic definitions and bounds derivation.
result Validated lower and upper bounds accurately track true interactions.
We improve conditional VAEs by incentivizing informative latent variables.
problem Structured-prediction tasks with one-to-many mappings.
method Modify latent variable model and introduce a multimodal prior.
result Significantly higher generalisation capability demonstrated on various datasets.
Paper presents a fast method to generate multimodal embeddings.
problem Integrating visual and linguistic information into a single representation.
method Learning a language-to-vision mapping to build multimodal embeddings.
result Mapped vectors outperform unimodal and multimodal baselines, especially in zero-shot settings.
Multimodal deep learning improves flaw detection in software programs.
problem Current flaw detection relies on single software representations.
method Adapted multimodal deep learning models for flaw detection.
result Multimodal models outperform traditional deep learning models.
Multimodal learning with deep Boltzmann machines (DBMs) is an generative approach to fuse multimodal inputs, and can learn the shared representation via Contrastive Divergence (CD) for classification and information retrieval tasks. However, it is a 2-fan DBM model, and cannot effectively handle multiple prediction tas…
Model learns tensor representations from imperfect multimodal data.
problem Learning from imperfect multimodal data with noise or missing entries.
method Tensor rank minimization to regularize rank of tensor representations.
result Model effectively learns tensor representations from imperfect data.