Study builds complex network from multimodal physiological data.
problem Understanding dynamic interactions in biological systems.
method Network-based multimodal data fusion using recurrence plots and temporal metrics.
result Model accurately characterizes emotional states through physiological responses.
Convolutional neural networks cluster multimodal data without labels.
problem Clustering multimodal data without labeled examples.
method Three-stage framework: encoder, self-expressive layer, decoder. Uses distance between reconstruction and input for training.
result Proposed methods significantly outperform state-of-the-art methods on three datasets.
DMAN embeds social images with deep multimodal attention networks.
problem Sub-optimal social image representation for social media data.
method Deep Multimodal Attention Networks (DMAN) that jointly embed multimodal contents and link information.
result DMAN achieves significant improvement in multi-label classification and cross-modal search compared to state-of-the-art image embeddings.
Proposes deep multimodal fusion for biometric identification.
problem Improving biometric identification accuracy with multiple modalities.
method Joint optimization of multiple modality-specific CNNs at different feature abstraction levels.
result Significant improvement in multimodal person identification performance.
Paper proposes RMFN for multimodal language analysis.
problem Modeling interactions between language, visual, and acoustic modalities.
method Recurrent Multistage Fusion Network (RMFN) decomposes fusion into stages focusing on subsets of multimodal signals.
result RMFN achieves state-of-the-art performance across multimodal sentiment analysis, emotion recognition, and speaker traits recognition datasets.
TCT learns multimodal sequence representations by translating from related sequences.
problem Challenges in learning semantic representations from multimodalities.
method Transformer based Cross-modal Translator (TCT) combined with Multimodal Transformer Network (MTN).
result Proposed method achieves new state-of-the-art performance on video-grounded dialogue.
A new GAN model LDAGAN uses Latent Dirichlet Allocation to model multimodal images.
problem Ignoring the structure and multimodal characteristics of vision data in GANs.
method Introduced a Dirichlet prior for multimodal image generation leading to LDAGAN. LDAGAN defines generative modes for each sample and uses a VEM algorithm for adversarial training.
result Experimental results show LDAGAN outperforms other GANs on real-world datasets.
GWN improves multimodal data fusion accuracy for chronic pain patients.
problem Dynamic and unspecified uncertainties in multimodal data fusion.
method Inspired by Global Workspace Theory, GWN is a neural network architecture that dynamically attends to multiple modalities.
result GWN achieved higher F1 scores (0.92 and 0.75) for multimodal discrimination and classification tasks.
Survey of multimodal deep generative models for diverse data types.
problem Inference of shared representations and cross-modal generation from heterogeneous multimodal data.
method Variational autoencoders and other deep generative models.
result A comprehensive survey of multimodal deep generative models.
Develops Cyclical SG-MCMC for exploring multimodal posterior distributions in deep learning.
problem High-dimensional, multimodal posterior distributions in Bayesian deep learning.
method Cyclical stepsize schedule in SG-MCMC to discover and characterize modes.
result Non-asymptotic convergence of the proposed algorithm.
Chain of neural networks learns scale-specific features for fast multimodal image registration.
problem Non-rigid registration of multimodal images, especially in remote sensing.
method Chain of fully-convolutional neural networks designed to learn scale-specific features, predicting deformation directly.
result Global registration in linear time, outperforming current methods in remote sensing tasks.
MDNs offer a data-efficient alternative to diffusion and flow models for multimodal scientific learning.
problem Capturing multimodal conditional uncertainty in scientific inverse problems.
method Mixture Density Networks (MDNs) as explicit parametric density estimators.
result MDNs achieve superior generalization, interpretability, and sample efficiency in scientific tasks.
Bayesian VFLMSP improves multimodal survival prediction with privacy.
problem Privacy and reliability in multimodal time-to-event prediction.
method Bayesian Vertical Federated Learning (VFL) with differential privacy.
result Consistent improvements in C-index compared to existing methods.
Graph-based multimodal federated learning for HAR improves accuracy and privacy.
problem Challenges in HAR due to noisy data, incomplete measurements, and privacy concerns.
method Proposes GraMFedDHAR, a Graph-based Multimodal Federated Learning framework for HAR tasks, using modality-specific graphs, residual GCNs, and attention-based fusion.
result Experimental results show up to 13 percent performance improvement for MultiModalGCN under differential privacy constraints.
Graph Mixture Density Networks model multimodal data on graphs.
problem Challenging conditional density estimation problems with structured data.
method Combining mixture models and graph representation learning.
result Significant improvement in likelihood of epidemic outcomes.
Multimodal deep learning improves toxicity prediction accuracy.
problem Improving prediction accuracy of chemical compound toxicity.
method Combining multiple neural network types and data representations.
result Significantly better accuracy on a toxicity benchmark.
Push-forward models struggle to fit multimodal distributions due to high Lipschitz constants.
problem Expressivity of push-forward generative models in fitting multimodal distributions.
method Analyzing the Lipschitz constant and its relation to the total variation distance and Kullback-Leibler divergence.
result Push-forward models require high Lipschitz constants to approximate multimodal distributions, leading to a trade-off between expressivity and stability.
Proposes a novel network for CTR prediction by learning modality-specific and modality-invariant representations.
problem Learning good representation of items from multimodal features in E-commerce is challenging due to redundant information across modalities.
method Introduces a Multimodal Adversarial Representation Network (MARN) that calculates modality-specific weights and learns modality-invariant representations.
result Consistently achieves remarkable improvements over state-of-the-art methods in CTR prediction.
Study proposes a multimodal model for cardiovascular risk prediction using EHRs.
problem Lack of comprehensive risk prediction from EHRs due to unstructured text.
method Proposes a multimodal BiLSTM model integrating structured and unstructured EHR data.
result Proposed BiLSTM model outperforms other DNN architectures in cardiovascular risk prediction.
Proposes MR-SNE for multimodal data visualization.
problem Visualizing data from multiple domains with relations across them.
method Extends t-SNE to compute augmented relations and jointly embed them in a low-dimensional space.
result Demonstrates promising performance in visualizing Flickr and Animal with Attributes 2 datasets.
PNNs improve treatment outcomes in TAVR and liver trauma.
problem Improving treatment outcomes in medical procedures.
method Multimodal Prescriptive Neural Networks (PNNs) combining optimization and machine learning.
result PNNs significantly improve estimated outcomes in medical procedures.
This paper presents a novel model for multimodal learning based on gated neural networks. The Gated Multimodal Unit (GMU) model is intended to be used as an internal unit in a neural network architecture whose purpose is to find an intermediate representation based on a combination of data from different modalities. Th…
Proposes neural network for causal inference with multimodal data.
problem Estimating causal effects with text and image data as confounders.
method Double machine learning framework adapted to partially linear models, semi-synthetic dataset generation.
result Improved performance in causal effect estimation with multimodal data.
Bayesian neural networks benefit from fully marginalizing over all modes to improve generalization.
problem Bayesian neural networks suffer from multimodal posterior distributions that can lead to suboptimal generalization.
method Use appropriate Bayesian sampling tools to fully marginalize over all posterior modes.
result Training with full marginalization improves the ability of the network to reason between multiple candidate solutions.
Generative Stochastic Networks (GSNs) have been recently introduced as an alternative to traditional probabilistic modeling: instead of parametrizing the data distribution directly, one parametrizes a transition operator for a Markov chain whose stationary distribution is an estimator of the data generating distributio…
Proposes a new CNN approach for multimodal biometric identification.
problem Improving biometric identification accuracy across multiple modalities.
method Uses a bank of modality-specific CNNs, fuses their outputs, and optimizes the system.
result Significantly outperforms unimodal systems and demonstrates reduction in parameters.
MDNE embeds network structures and attributes for better analysis.
problem Preserving both structural and attribute features in network embedding.
method Multimodal Deep Network Embedding (MDNE) using deep model with multiple layers of non-linear functions.
result MDNE outperforms baselines on various tasks with real-world datasets.
New method captures multimodal disconnectivity in schizophrenia.
problem Misinterpretation of single modality data in schizophrenia research.
method Gaussian graphical model and modularity-based approach on multimodal data.
result Identifies missing links in schizophrenia's default mode network.
DeepSIP predicts network failures' impact using CNN from syslog and traffic data.
problem Predicting service impact from network failures.
method Temporal multimodal CNN for predicting time to recovery and traffic loss.
result DeepSIP reduced prediction error by approximately 50%.
Study predicts traffic congestion based on population mobility data.
problem Predicting traffic congestion in multimodal transport networks.
method Machine learning methods applied to population mobility data.
result Likely prediction of congestion based on population movements.
UniPhyNet improves cognitive load classification accuracy using EEG, ECG, and EDA signals.
problem Classifying cognitive load using multimodal physiological data.
method Unified network architecture integrating multiscale parallel convolutional blocks, ResNet-type blocks, and channel block attention module. Uses bidirectional gated recurrent unit for temporal dependencies.
result Improves raw signal classification accuracy from 70% to 80% (binary) and 62% to 74% (ternary) on CL-Drive dataset.
Current high-throughput data acquisition technologies probe dynamical systems with different imaging modalities, generating massive data sets at different spatial and temporal resolutions posing challenging problems in multimodal data fusion. A case in point is the attempt to parse out the brain structures and networks…
Paper introduces IGMM-GAN for multimodal anomaly detection in mobility data.
problem Lack of ground truth data and dependence on pre-processing for anomaly detection in human mobility.
method Coupled IGMM-GAN for generating realistic synthetic datasets and multimodal anomaly detection.
result IGMM-GAN improves anomaly detection performance over existing GAN methods.
New deep fusion methods improve human action recognition using depth and inertial sensor data.
problem Existing multimodal HAR frameworks lack mid-level feature fusion.
method Proposes three deep multilevel multimodal fusion frameworks, transforming depth and inertial sensor data into images and using convolution with Prewitt filter to create modality within modality.
result Supremacy of proposed fusion frameworks over existing methods on three publicly available datasets.
Study improves stock movement prediction using multimodal data.
problem Inaccurate stock movement prediction due to incomplete multimodal data integration.
method Introduces MSGCA framework for robust multimodal fusion.
result MSGCA framework outperforms existing methods by 21.7% on multimodal datasets.
DeepBioD combines feature engineering and learned representations for biodegradability prediction.
problem Limited labeled data for biodegradability prediction.
method Developed a multimodal CNN-MLP neural network that integrates engineered and learned representations.
result DeepBioD achieves an error rate of 0.125, 27% lower than state-of-the-art methods.
Deep learning predicts mental disorders from audio and text samples.
problem Predicting mental disorders from speech samples.
method Multimodal deep learning structure using various pre-trained models for audio and text embeddings, transfer learning, and auxiliary corpora.
result Acceptable accuracy in predicting mental disorders through multimodal analysis.
The paper introduces multimodal generative models to improve data marginal likelihood.
problem Improving data marginal likelihood in multimodal settings.
method Derives variational bounds on the evidence for multimodal deep generative models, generalizes objectives for different model types, and benchmarks across various datasets.
result Multimodal VAEs excel in image, label, and text datasets with and without weak supervision.
Enhances hand gesture recognition with separate networks and shared features.
problem Improving recognition accuracy of unimodal 3D-CNNs for dynamic hand gestures.
method Separate networks for each modality, collaborative learning, spatiotemporal semantic alignment loss, focal regularization.
result Improves test time recognition accuracy and state-of-the-art performance.
The study develops multimodal models to predict 1-year mortality risk from large clinical datasets.
problem Limited datasets in biomedical studies that do not generalize over large heterogeneous datasets.
method Develops multimodal models using a massive clinical dataset of 25 million videos and 2.9 million ECG traces.
result Extremely low-parameter models with optimized feature selection achieve AUC of 0.89.
Hybrid method improves sampling from multimodal distributions.
problem Sampling from multimodal posterior distributions efficiently.
method Jump-Diffusion Langevin Dynamics hybrid with Metropolis.
result Calibrated hybrid method outperforms pure methods.
Paper proposes ICCN to learn correlations between text, audio, and video for multimodal sentiment analysis.
problem Improving multimodal sentiment analysis by learning hidden correlations between text and audio/video features.
method Interaction Canonical Correlation Network (ICCN) using deep canonical correlation analysis (DCCA).
result Empirical results confirm the effectiveness of ICCN in capturing useful information from all three views.
BIL allows binary input data in CNNs, improving performance on multimodal datasets.
problem Efficient execution of CNNs on edge devices with reduced bit width.
method BIL concept that learns bit-specific binary weights for binary input data.
result BIL outperforms full precision weights by 1.92% on multimodal datasets.
This research improves multimodal systems by adding a second objective and regularisation methods.
problem Improving performance of multimodal systems with multiple objectives and regularisation.
method Introduces a second objective over multimodal fusion using variational inference and regularisation methods.
result Demonstrates potential for multiple objectives and probabilistic methods to lower variance and improve generalisation.
lamBERT learns language and actions using multimodal BERT.
problem Learning language and actions in complex environments.
method Extending BERT to multimodal representation and integrating with reinforcement learning.
result lamBERT model achieved higher rewards in multitask and transfer settings.
Bayesian neural networks reveal multimodal predictive distributions.
problem Uncertainty quantification and interpretability in neural networks.
method Discretized prior for inner layer weights, Gaussian mixture approximation of posterior predictive distribution.
result Distinct parameter realizations can produce the same training error but different posterior predictive distributions.
A new mutual information lower bound for multimodal regression active learning.
problem Lack of effective acquisition functions for multimodal regression active learning.
method Introduces a Two-Index framework for separating epistemic and aleatoric sources of uncertainty, deriving MI-LB as a closed-form approximation.
result MI-LB consistently outperforms baselines on multimodal regression tasks.
Generates multimodal safety-critical scenarios for robustness evaluation of decision-making algorithms.
problem Lack of comprehensive evaluation of neural network robustness under real-world scenarios.
method Proposes a flow-based multimodal scenario generator using weighted likelihood maximization and gradient-based sampling.
result Demonstrates improved testing efficiency and multimodal modeling capability compared to traditional methods.