Proposes MR-SNE for multimodal data visualization.
problem Visualizing data from multiple domains with relations across them.
method Extends t-SNE to compute augmented relations and jointly embed them in a low-dimensional space.
result Demonstrates promising performance in visualizing Flickr and Animal with Attributes 2 datasets.
MAESTRO improves multimodal learning for dynamic time series with adaptive attention and robustness.
problem Challenges in multimodal learning, especially in healthcare and daily living.
method Dynamic intra- and cross-modal interactions, symbolic tokenization, adaptive attention budgeting, sparse cross-modal attention, MoE mechanism.
result Average relative improvements of 4% and 8% over existing multimodal and multivariate approaches, respectively, under complete observations.
Pilot study shows multimodal signals improve chess expertise detection.
problem Detecting chess expertise through multimodal signals.
method Multimodal observation of chess players' eye-gaze, posture, emotion, and body language.
result Multimodal approach reaches up to 93% accuracy in detecting chess expertise compared to 86% with unimodal approach.
Paper discovers hidden common variables in nonlinear data.
problem Discover hidden common variables in nonlinear high-dimensional observations.
method Local CCA metric integrated with manifold learning.
result Metric discovers hidden common variables without rigid model assumptions.
Study predicts traffic congestion based on population mobility data.
problem Predicting traffic congestion in multimodal transport networks.
method Machine learning methods applied to population mobility data.
result Likely prediction of congestion based on population movements.
Model learns tensor representations from imperfect multimodal data.
problem Learning from imperfect multimodal data with noise or missing entries.
method Tensor rank minimization to regularize rank of tensor representations.
result Model effectively learns tensor representations from imperfect data.
ICYM2I corrects missingness bias in multimodal learning.
problem Missingness patterns between source and target environments affect multimodal learning performance.
method ICYM2I uses inverse probability weighting to correct missingness bias in predictive performance and information gain.
result ICYM2I improves multimodal learning performance by accounting for missingness.
The paper tackles conditional multimodal learning using variational methods.
problem Learning conditional distributions between modalities.
method Variational methods for maximizing conditional log-likelihood.
result Generated faces are more representative of the attributes.
Improved multimodal variational models capture more complex joint distributions.
problem Limited expressiveness of multimodal variational models.
method Used normalizing flows to approximate and transform a simple parametric joint posterior into a more complex one.
result The model improves on state-of-the-art multimodal variational methods on various tasks.
New method clusters multimodal data with consistency.
problem Multimodal clustering with unaligned data.
method Conjugate mixture models and EM algorithm.
result Consistent multimodal clustering achieved.
MKBE embeds multimodal data for knowledge base completion.
problem Missing multimodal data in knowledge bases.
method Multimodal encoders and decoders for text, images, and numerical values.
result State-of-the-art link prediction with 5-7% improvement.
The paper introduces multimodal generative models to improve data marginal likelihood.
problem Improving data marginal likelihood in multimodal settings.
method Derives variational bounds on the evidence for multimodal deep generative models, generalizes objectives for different model types, and benchmarks across various datasets.
result Multimodal VAEs excel in image, label, and text datasets with and without weak supervision.
New framework estimates graph from multimodal functional data.
problem Estimating graph from joint multimodal functional data.
method Integrative framework using partial correlation operator.
result Estimator converges to stationary point with quantifiable error.
A framework for uncertainty-aware multimodal learning using conformal Shapley intervals.
problem Uncertainty and modality level importance in multimodal learning.
method Introduces conformal Shapley intervals to quantify modality level importance and uncertainty.
result Demonstrates meaningful uncertainty quantification and strong predictive performance.
New method disentangles shared and private latent factors in multimodal data.
problem Challenges in disentangling shared and private latent factors in multimodal data.
method Proposes a modification to existing multimodal Variational Autoencoders (MMVAE) to better handle modality-specific variation.
result Demonstrates improved robustness of modified MMVAE to modality-specific variation.
Paper presents a method to align unpaired samples across different modalities.
problem Challenges in collecting paired samples for multimodal representation learning.
method Uses propensity score alignment based on Rubin's framework to estimate a common space for unpaired samples.
result Optimal transport matching significantly improves alignment in real-world data.
Nonlinear filtering extracts relevant variables from multimodal sleep data.
problem Recover relevant variables from multiple sensor data.
method Diffusion-based manifold learning for nonlinear filtering.
result Method gives robust data-driven representation correlated with sleep process.
New method for risk allocation under multimodality of loss distribution.
problem Risk assessment under multimodal conditional loss distribution.
method Maximum Likelihood Allocation (MLA) and multimodality adjustment.
result Multimodality adjustment improves soundness of risk allocations.
New algorithm learns switching dynamics from multiple neural signals.
problem Learning accurate switching dynamical system models from multimodal neural data.
method Unsupervised learning algorithm for multiscale switching dynamical system models.
result Switching multiscale dynamical system models outperform single-scale models in behavior decoding.
Graph-based multimodal federated learning for HAR improves accuracy and privacy.
problem Challenges in HAR due to noisy data, incomplete measurements, and privacy concerns.
method Proposes GraMFedDHAR, a Graph-based Multimodal Federated Learning framework for HAR tasks, using modality-specific graphs, residual GCNs, and attention-based fusion.
result Experimental results show up to 13 percent performance improvement for MultiModalGCN under differential privacy constraints.
This paper investigates uncertainty calibration in multimodal large language models.
problem Challenges in properly calibrating uncertainty in multimodal large language models.
method Investigation of representative MLLMs across various scenarios, including visual fine-tuning and multimodal training.
result MLLMs tend to give answers rather than admit uncertainty, but this self-assessment improves with proper prompt adjustments.
Bayesian model updating uses VAEs to approximate likelihood with small data.
problem Approximating likelihood for small data sets in structural analysis.
method Uses multimodal VAEs to approximate likelihood, suitable for high-dimensional correlated observations.
result Demonstrates computational efficiency and accuracy compared to original VAE approach.
HMC and RWM have similar performance on multimodal densities.
problem Comparing the performance of Hamiltonian Monte Carlo and Random-Walk Metropolis on multimodal distributions.
method Computed spectral gaps for both algorithms on specific multimodal target densities.
result HMC and RWM have identical spectral gaps for multimodal targets.
Study improves document processing in banking with multimodal analytics.
problem Raising operational efficiency in banking through document-intensive processes.
method Comparative analysis of text classifiers and multimodal model (LayoutXLM) on company register extracts.
result Incorporating layout information in a model substantially increases performance.
Framework detects anomalies in industrial processes using deep learning.
problem Detect anomalies in complex industrial processes.
method Causal-based framework with unsupervised deep learning.
result Successfully validated abstract contexts of blast furnace assets.
New method optimizes portfolios by dynamically integrating ESG constraints.
problem Static ESG scores mismatch sequential portfolio decisions.
method MACF-X, a family of adapters that learns ESG costs from multimodal evidence.
result Reduces tail ESG budget pressure while maintaining financial performance.
WebGUM learns web navigation from multimodal data, outperforming previous methods.
problem Limited generalization from domain-specific models in web navigation.
method Instruction-following multimodal agent trained on vision-language foundation models.
result Significant improvement in web navigation performance on benchmarks.
New model improves multimodal autoencoders by learning joint and conditional distributions.
problem Limitations in recent multimodal autoencoders restrict their quality on complex datasets.
method Proposes a multistage training process with variational inference and Normalizing Flows, leveraging shared modality information.
result Achieves state-of-the-art results on benchmark datasets.
Paper analyzes Annealed Langevin Dynamics for multimodal sampling stability.
problem Ensuring stability of Annealed Langevin Dynamics across dimensions.
method Uniform-in-dimension analysis of ALD for Gaussian-mixture targets.
result ALD achieves prescribed accuracy in KL divergence with spectral conditions.
Paper proposes Seq2Seq models for multimodal sentiment analysis.
problem Learning representations from multiple modalities in machine learning.
method Two unsupervised Seq2Seq models for multimodal sentiment analysis.
result Seq2Seq models improve F1 Score by twelve points in Bimodal sentiment analysis.
Paper proposes a method to preserve multimodal sentiment analysis fidelity.
problem Lack of fidelity in multimodal fusion for sentiment analysis.
method Variational autoencoder-based approach for modality fusion.
result Empirically shows superior performance over state-of-the-art methods.
Graph matching is a challenging problem with very important applications in a wide range of fields, from image and video analysis to biological and biomedical problems. We propose a robust graph matching algorithm inspired in sparsity-related techniques. We cast the problem, resembling group or collaborative sparsity f…
TCT learns multimodal sequence representations by translating from related sequences.
problem Challenges in learning semantic representations from multimodalities.
method Transformer based Cross-modal Translator (TCT) combined with Multimodal Transformer Network (MTN).
result Proposed method achieves new state-of-the-art performance on video-grounded dialogue.
Efficient multimodal fusion reduces complexity and improves performance.
problem Multimodal data fusion with tensor transformations.
method Low-rank Multimodal Fusion using tensors.
result Significant reduction in computational complexity with competitive performance.
Paper proposes RMFN for multimodal language analysis.
problem Modeling interactions between language, visual, and acoustic modalities.
method Recurrent Multistage Fusion Network (RMFN) decomposes fusion into stages focusing on subsets of multimodal signals.
result RMFN achieves state-of-the-art performance across multimodal sentiment analysis, emotion recognition, and speaker traits recognition datasets.
Multimodal bitransformer boosts image-text classification.
problem Combining text and image modalities for improved classification.
method Supervised multimodal bitransformer model integrating text and image encoders.
result State-of-the-art performance on multimodal classification benchmarks.
UR-FUNNY dataset aids in understanding multimodal humor.
problem Understanding humor in a multimodal context is understudied.
method Developed a multimodal dataset (UR-FUNNY) for humor detection.
result UR-FUNNY opens the door to multimodal humor detection research.
Survey of multimodal deep generative models for diverse data types.
problem Inference of shared representations and cross-modal generation from heterogeneous multimodal data.
method Variational autoencoders and other deep generative models.
result A comprehensive survey of multimodal deep generative models.
This paper proposes a model to learn multimodal representations robust to missing data.
problem Learning multimodal representations from heterogeneous sources of information.
method Optimizes a joint generative-discriminative objective across multimodal data and labels, factorizing representations into multimodal discriminative and modality-specific generative factors.
result The proposed model achieves state-of-the-art performance on six multimodal datasets and can reconstruct missing modalities without significant performance drop.
FMT model improves multimodal sequential learning across language, vision, and acoustic data.
problem Modeling spatio-temporal dynamics across multiple modalities.
method Factorized Multimodal Transformer (FMT) that models intramodal and intermodal dynamics in a factorized manner.
result FMT outperforms existing models on 3 datasets and 21 labels, setting new state of the art.
Develops a contrastive framework for data-efficient multimodal learning.
problem Expensive training of multimodal generative models requiring related multimodal data.
method Contrastive framework for multimodal learning, distinguishing related from unrelated data.
result Data-efficient multimodal learning on challenging datasets for various VAE models.
Plug-and-play multimodal controller improves class-conditional image generation.
problem Generating class-conditional images from user-specified labels.
method Introduces a `multimodal controller` to generate multimodal data without additional learning parameters.
result Multimodal controlled generative models produce higher quality class-conditional images and novel modalities.
Paper presents a fast method to generate multimodal embeddings.
problem Integrating visual and linguistic information into a single representation.
method Learning a language-to-vision mapping to build multimodal embeddings.
result Mapped vectors outperform unimodal and multimodal baselines, especially in zero-shot settings.
Improved sentiment analysis with multimodal data.
problem Cross-modal sentiment analysis in social media, customer service, and video blogs.
method Gated mechanism for attention-based learning of cross-modal interactions, with experiments on CMU-MOSI and CMU-MOSEI datasets.
result 1.6% and 1.34% absolute improvement over state-of-the-art.
Paper proposes multimodal contrastive learning for EHR data.
problem Separate treatment of structured and unstructured EHR data.
method Proposes a multimodal feature embedding generative model and a multimodal contrastive loss.
result Multimodal learning yields better feature representation than single-modality learning.
EmbraceNet improves robustness in multimodal classification.
problem Robust multimodal classification with partial data loss.
method Deep learning architecture for multimodal fusion.
result EmbraceNet outperforms other models in partial data scenarios.
Convolutional neural networks cluster multimodal data without labels.
problem Clustering multimodal data without labeled examples.
method Three-stage framework: encoder, self-expressive layer, decoder. Uses distance between reconstruction and input for training.
result Proposed methods significantly outperform state-of-the-art methods on three datasets.
Study builds complex network from multimodal physiological data.
problem Understanding dynamic interactions in biological systems.
method Network-based multimodal data fusion using recurrence plots and temporal metrics.
result Model accurately characterizes emotional states through physiological responses.