Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

25.0%50.0%75.0%100.0% · May 199319922001200920182026
48 results for conditional multimodal learning

Plug-and-play multimodal controller improves class-conditional image generation.

problem Generating class-conditional images from user-specified labels.
method Introduces a `multimodal controller` to generate multimodal data without additional learning parameters.
result Multimodal controlled generative models produce higher quality class-conditional images and novel modalities.

Unified approach for multimodal data prediction using synthetic data generation.

problem Challenges in integrating heterogeneous data types for accurate predictive performance.
method Generative Distribution Prediction (GDP) framework that uses multimodal synthetic data generation.
result Empirical validation across four tasks demonstrates versatility and effectiveness of GDP.

This paper proposes a model to learn multimodal representations robust to missing data.

problem Learning multimodal representations from heterogeneous sources of information.
method Optimizes a joint generative-discriminative objective across multimodal data and labels, factorizing representations into multimodal discriminative and modality-specific generative factors.
result The proposed model achieves state-of-the-art performance on six multimodal datasets and can reconstruct missing modalities without significant performance drop.

This paper strengthens the computational separation between multimodal and unimodal learning, showing unimodal learning is hard on typical instances.

problem Theoretical justification for empirical success of multimodal machine learning.
method Introduced a stronger average-case computational separation between unimodal and multimodal learning.
result For typical instances, unimodal learning is computationally hard, while multimodal learning is easy.

PVAE learns disentangled representations from multimodal data.

problem Learning disentangled representations from multimodal sensory data.
method Partitioned Variational Autoencoder (PVAE) with multimodal generative model and training objectives.
result PVAE achieves over 99% accuracy on both modalities for semantic units.

New model improves multimodal autoencoders by learning joint and conditional distributions.

problem Limitations in recent multimodal autoencoders restrict their quality on complex datasets.
method Proposes a multistage training process with variational inference and Normalizing Flows, leveraging shared modality information.
result Achieves state-of-the-art results on benchmark datasets.

A framework for uncertainty-aware multimodal learning using conformal Shapley intervals.

problem Uncertainty and modality level importance in multimodal learning.
method Introduces conformal Shapley intervals to quantify modality level importance and uncertainty.
result Demonstrates meaningful uncertainty quantification and strong predictive performance.

MDNs offer a data-efficient alternative to diffusion and flow models for multimodal scientific learning.

problem Capturing multimodal conditional uncertainty in scientific inverse problems.
method Mixture Density Networks (MDNs) as explicit parametric density estimators.
result MDNs achieve superior generalization, interpretability, and sample efficiency in scientific tasks.

Generative Score Inference improves uncertainty quantification for multimodal data.

problem Accurate uncertainty quantification in multimodal learning tasks.
method Generative Score Inference (GSI) uses synthetic samples to approximate conditional score distributions.
result GSI achieves state-of-the-art performance in hallucination detection and image captioning uncertainty estimation.

GGMPs improve non-Gaussian conditional density estimation.

problem Multimodality, heteroscedasticity, and strong non-Gaussianity in conditional density estimation.
method GGMP combines local Gaussian mixture fitting, cross-input component alignment, and per-component heteroscedastic GP training.
result GGMPs improve distributional approximation on synthetic and real-world datasets.

New method handles missing data and multiple data types in time series models.

problem Handling missing data and multiple data modalities in time series models.
method Factorized inference method for Multimodal Deep Markov Models (MDMMs).
result Method performs well even with high levels of missing data and outperforms existing approaches.

Deep model predicts shapes of curves with multiple covariates.

problem Predicting shapes of planar curves with various covariates.
method Deep learning model using complex-valued functions, conditional covariance smoother with modality-specific encoders.
result Model accurately predicts shapes of curves with multimodal covariates.

Study analyzes bias interactions in multimodal models using simulation-based methods.

problem Analyzing dynamic bias interactions in multimodal models to ensure fairness and equity.
method Simulation-based heuristic approach to compute bias scores for text-only, image-only, and multimodal embeddings.
result Multimodal bias interactions can be amplification, mitigation, or neutral, with text bias often dominant.

TCT learns multimodal sequence representations by translating from related sequences.

problem Challenges in learning semantic representations from multimodalities.
method Transformer based Cross-modal Translator (TCT) combined with Multimodal Transformer Network (MTN).
result Proposed method achieves new state-of-the-art performance on video-grounded dialogue.

Develops a contrastive framework for data-efficient multimodal learning.

problem Expensive training of multimodal generative models requiring related multimodal data.
method Contrastive framework for multimodal learning, distinguishing related from unrelated data.
result Data-efficient multimodal learning on challenging datasets for various VAE models.

FMT model improves multimodal sequential learning across language, vision, and acoustic data.

problem Modeling spatio-temporal dynamics across multiple modalities.
method Factorized Multimodal Transformer (FMT) that models intramodal and intermodal dynamics in a factorized manner.
result FMT outperforms existing models on 3 datasets and 21 labels, setting new state of the art.

CQNPs enhance predictive performance and distribution modeling using quantile regression.

problem Limited predictive likelihood of Gaussian models for complex distributions.
method Introducing Conditional Quantile Neural Processes (CQNPs) that focus on estimating informative quantiles.
result Significant improvements in predictive performance and better modeling of multimodal distributions.

New method identifies shared components from unpaired multimodal mixtures.

problem Identify shared components from unpaired multimodal mixtures.
method Distribution divergence minimization-based loss with sufficient conditions for identifiability.
result Sufficient conditions for shared component identifiability from unaligned multimodal mixtures.

Enhances multimodal generation with Normalizing Flows and correlation analysis.

problem Generating coherent cross-modal data from multiple sources.
method Uses Deep Canonical Correlation Analysis for shared information, Normalizing Flows for diversity, and Product of Experts for scalability.
result Improves likelihood, diversity, and coherence in conditional generation.

Paper proposes multimodal contrastive learning for EHR data.

problem Separate treatment of structured and unstructured EHR data.
method Proposes a multimodal feature embedding generative model and a multimodal contrastive loss.
result Multimodal learning yields better feature representation than single-modality learning.

Framework translates images between domains without supervision.

problem Challenges in unsupervised image-to-image translation, especially handling multimodality.
method Proposes a Multimodal Unsupervised Image-to-Image Translation (MUNIT) framework, decomposing images into content and style codes.
result Demonstrates improved generation of diverse outputs from a single source image.

Paper proposes a harmonized approach to multimodal learning using GPLVMs.

problem Modality heterogeneity in multimodal data.
method Develops a novel learning scheme called Harmonization to jointly learn latent model parameters from different modalities.
result Experimental results show superior performance in cross-modal retrieval tasks.

New results show contrastive learning can recover shared factors in multimodal data.

problem Understanding when contrastive learning can recover shared latent factors in multimodal data.
method New identifiability results for multimodal contrastive learning, distinguishing between multi-view and multimodal settings.
result Contrastive learning can block-identify shared latent factors in multimodal data, even with dependencies.

Improved sentiment analysis with multimodal data.

problem Cross-modal sentiment analysis in social media, customer service, and video blogs.
method Gated mechanism for attention-based learning of cross-modal interactions, with experiments on CMU-MOSI and CMU-MOSEI datasets.
result 1.6% and 1.34% absolute improvement over state-of-the-art.

New method uses MRI data to improve PET tomography uncertainty quantification.

problem Improving uncertainty quantification in emission tomography with multimodal data.
method Nonparametric posterior learning technique adapted for Poisson-type data.
result Sampling algorithms are scalable, parallelizable, and easy to implement.

New model uses financial filings to predict bankruptcy, even without MDA sections.

problem Lack of complete MDA data limits traditional bankruptcy prediction models.
method Conditional Multimodal Discriminative (CMMD) model learns from accounting, market, and textual data.
result Empirical results show superior classification performance compared to traditional models.

Improved software flaw detection using NAS on multimodal DL models.

problem Software flaw detection in multimodal deep learning models.
method Adapted NAS framework for multimodal learning, combined with multimodal deep learning models.
result Improved performance on the Juliet Test Suite.

Paper presents a fast method to generate multimodal embeddings.

problem Integrating visual and linguistic information into a single representation.
method Learning a language-to-vision mapping to build multimodal embeddings.
result Mapped vectors outperform unimodal and multimodal baselines, especially in zero-shot settings.

Multimodal learning with deep Boltzmann machines (DBMs) is an generative approach to fuse multimodal inputs, and can learn the shared representation via Contrastive Divergence (CD) for classification and information retrieval tasks. However, it is a 2-fan DBM model, and cannot effectively handle multiple prediction tas…

2015-03-26abs ↗pdf ↗

Model learns tensor representations from imperfect multimodal data.

problem Learning from imperfect multimodal data with noise or missing entries.
method Tensor rank minimization to regularize rank of tensor representations.
result Model effectively learns tensor representations from imperfect data.