This paper reviews deep learning for multi-modality medical image segmentation.
problem Improving segmentation accuracy in medical images using multiple modalities.
method Overview of deep learning and multi-modal medical image segmentation, analysis of different network architectures and fusion strategies.
result Later fusion of modalities can lead to more accurate segmentation results.
The paper analyzes and proposes an algorithm for multi-modal nonlinear embeddings with theoretical performance bounds.
problem Generalizability of multi-modal nonlinear embeddings to unseen data.
method Theoretical analysis and a multi-modal nonlinear representation learning algorithm motivated by performance bounds.
result The proposed algorithm yields promising performance in multi-modal image classification and cross-modal image-text retrieval applications.
Proposes LM3FE for multi-modal feature extraction in image classification.
problem High-dimensional features and multi-modal data challenges.
method Large margin multi-modal multi-task feature extraction (LM3FE) framework.
result LM3FE outperforms single-task feature extraction and multi-modal feature extraction.
MMVAE learns multi-modal data with shared and private latent spaces.
problem Learning useful representations across multiple data modalities.
method Mixture-of-experts variational autoencoder (MMVAE).
result MMVAE satisfies four criteria for multi-modal learning.
Develops a deep model for joint image-text learning.
problem Bidirectional joint image-text modeling.
method Variational hetero-encoder randomized GAN (VHE-GAN).
result Achieves state-of-the-art performance in image-text learning and generation.
Graph convolution model uses self-attention to predict diseases from multi-modal data.
problem Predicting diseases from diverse multi-modal data.
method Graph convolution with self-attention layer.
result Significantly outperforms state-of-the-art methods in disease prediction.
AI framework diagnoses Parkinson's disease with 100% accuracy.
problem Expertise-demanding medical imaging procedures for Parkinson's disease diagnosis.
method End-to-end, multi-modality diagnosis framework using T1-MRI and 11C-CFT PET.
result 100% accuracy in PD/NL classification.
Proposes a novel method for detecting novelty in multi-modal data.
problem Challenges in detecting novelty in high-dimensional, multi-modal data.
method Orthogonalized latent space for disentangling features and defining novelty score.
result Proposed method outperforms state-of-the-art algorithms in novelty detection.
Improved self-supervised learning for document images.
problem Performance of self-supervised pre-training on document images is poor.
method Proposed context-aware alternatives and a novel multi-modal method.
result Novel method outperforms other self-supervised methods on document image classification.
KD-Net transfers knowledge from multi-modal to mono-modal segmentation networks.
problem Limited acquisition of multiple imaging modalities in clinical settings.
method Generalized distillation framework adapted for mono-modal networks.
result The student network outperforms baseline mono-modal networks in brain tumor segmentation.
Study improves product categorization on Amazon using multi-modal fusion.
problem Multi-label product categorization in e-commerce.
method Late fusion of image, description, and title modalities using modified CNN and ResNet-50 models.
result Tri-modal late fusion model achieved an F1 score of 88.2%, significantly better than single modal models. GPCCA integrates multi-modal data with missing values, improving clustering accuracy.
problem Integrating and analyzing multi-modal data with missing values and partial observations.
method Generalized Probabilistic Canonical Correlation Analysis (GPCCA) for unsupervised multi-modal data integration and dimensionality reduction.
result GPCCA outperforms existing methods in capturing essential patterns across modalities and provides robust low-dimensional embeddings.
SAFE detects fake news by analyzing text and image similarities.
problem Detecting fake news with less focus on text-image similarity.
method SAFE uses neural networks to extract text and visual features, then learns their relationship to predict fake news.
result SAFE effectively recognizes fake news based on text, images, or mismatches.
Framework for handling long-tailed multi-modal data.
problem Class imbalance and long-tailed distributions in multi-modal data.
method Multi-expert architecture with modality-specific networks and dynamic fusion weights.
result Framework outperforms existing methods in long-tailed, class-imbalanced scenarios.
Adaptive anchor methods improve multi-modal learning by balancing intra-modal and inter-modal information.
problem Fixed anchor methods limit multi-modal learning by over-reliance on a single modality and inadequate cross-modal correlation.
method Adaptive anchor methods using centroid-based anchors from all modalities.
result Adaptive anchor methods like CentroBind consistently outperform fixed anchor methods across various datasets.
Diffusion models enhance robotic manipulation through probabilistic multi-modal learning.
problem Enhancing robotic manipulation through robust and multi-modal learning.
method Probabilistic diffusion models integrating imitation and reinforcement learning.
result Diffusion models improve grasp learning, trajectory planning, and data augmentation in robotics.
MMGAN stabilizes GANs for multi-modal data clustering.
problem Stability and data distribution loss in GANs for multi-modal data.
method Model latent space as Gaussian mixture model, clusters data manifolds, and trains with clustering network.
result MMGAN outperforms state-of-the-art models in clustering multi-modal data.
Framework for causal discovery using multi-modal data.
problem Failure of representation learning in causal tasks.
method Statistical and computational framework combining representation learning and causal inference.
result Effective use of observational and perturbational data for causal discovery.
Unified architecture for multi-modal multi-task learning using transformer.
problem Training multiple tasks concurrently with varying modalities.
method Spatio-temporal cache mechanism for multi-modal learning.
result Training multiple tasks together reduces model size by about three times.
A GAN variant synthesizes missing MRI sequences from available ones.
problem Missing MRI sequences due to various constraints.
method Multi-modal Generative Adversarial Network (GAN) that combines multiple available sequences to synthesize missing ones.
result The proposed GAN method outperforms competing approaches in synthesizing missing MRI sequences.
Proposes a method to generate diverse outputs in conditional GANs.
problem Mode-collapse in conditional GANs, where outputs are overly simplified.
method Explicit regularization to produce diverse outputs based on latent codes.
result Demonstrates improved diversity in image-to-image translation, inpainting, and future video prediction tasks.
Paper develops a theory explaining contrastive pre-training for multimodal AI.
problem Limited theoretical understanding of contrastive pre-training for multi-modal AI.
method Introduces approximate sufficient statistics and Joint Generative Hierarchical Model.
result Near-minimizers of contrastive loss are approximately sufficient, enabling diverse downstream tasks.
Image captioning has demonstrated models that are capable of generating plausible text given input images or videos. Further, recent work in image generation has shown significant improvements in image quality when text is used as a prior. Our work ties these concepts together by creating an architecture that can enabl…
Paper tackles cross-modal anomalies in multi-source data.
problem Detect anomalies in multi-modal data where patterns are inconsistent across different sources.
method Proposes a deep structured anomaly detection framework.
result Demonstrates effectiveness on real-world datasets.
Increasingly many real world tasks involve data in multiple modalities or views. This has motivated the development of many effective algorithms for learning a common latent space to relate multiple domains. However, most existing cross-view learning algorithms assume access to paired data for training. Their applicabi…
Generates realistic faces from detailed textual descriptions.
problem Face generation from fine-grained textual descriptions.
method Conditional GAN model with DC-GAN and GAN-CLS loss, using CelebA dataset with generated captions.
result Promising results in generating diverse face images from fine-grained textual descriptions.
New model learns from missing modalities and class labels.
problem Conflict between learning joint representations and modalities in multi-modal data.
method Introduces a novel conditional multi-modal discriminative model using an informative prior distribution and a likelihood-free objective function.
result Our model achieves state-of-the-art results in downstream classification, acoustic inversion, and image and annotation generation.
Study shows agents benefit from hearing in addition to vision.
problem Limited effectiveness of vision-only reinforcement learning agents.
method Used audio as complementary information to visual cues in state representation.
result Agents perform better when hearing is added to vision.
Multi-modal data collections, such as corpora of paired images and text snippets, require analysis methods beyond single-view component and topic models. For continuous observations the current dominant approach is based on extensions of canonical correlation analysis, factorizing the variation into components shared b…
Paper proposes a new hierarchical attention mechanism for multi-scale data.
problem Challenges in applying neural attention to multi-scale, multi-modal data.
method Developed a mathematical framework for multi-modal, multi-scale data and derived optimal neural attention mechanics.
result Proposed hierarchical attention mechanism improves transformer performance in multi-scale, multi-modal settings.
ROME improves density estimation for multi-modal, non-normal data.
problem Robust multi-modal density estimation in non-normal, highly correlated distributions.
method ROME uses clustering to segment multi-modal data into uni-modal clusters, then combines KDE estimates for each cluster.
result ROME outperforms state-of-the-art methods and is more robust to various distributions.
Proposes a new LSTM gate structure using bivariate Beta distribution.
problem Inflexibility of sigmoid gates in modeling multi-modality and skewness, and lack of modeling correlation between gates.
method Introduces a bivariate Beta distribution gate structure within LSTM cells.
result Empirically shows higher gradient values and improved model performance.
Urban2Vec combines street view imagery and POIs for better urban neighborhood embeddings.
problem Lack of comprehensive representation of urban neighborhoods using heterogeneous data.
method Unsupervised multi-modal framework using CNN for visual features and bag-of-words for POI data.
result Urban2Vec achieves better performance than baseline models and comparable to fully-supervised methods.
The group affect or emotion in an image of people can be inferred by extracting features about both the people in the picture and the overall makeup of the scene. The state-of-the-art on this problem investigates a combination of facial features, scene extraction and even audio tonality. This paper combines three addit…
O-GANs improve fashion image generation by conditioning on a hierarchical ontology.
problem Challenges in training GANs to generate images from text descriptions.
method Ontology Generative Adversarial Networks (O-GANs) that condition on a hierarchical fashion ontology.
result O-GANs achieve better image quality and conditioning results compared to standard GANs.
MoCA uses a novel autoencoder to analyze multi-modal health data.
problem Challenges in analyzing continuous multi-modal health data from wearable devices.
method Proposes MoCA, a self-supervised learning framework combining transformer and masked autoencoder methods.
result Demonstrates strong performance boosts across reconstruction and classification tasks.
The paper evaluates samplers on multi-modal targets, focusing on mode separation and recovery.
problem Handling multi-modality in sampling.
method Synthetic experimental setting focusing on mode relative importance recovery.
result Illustrates the challenges and potential of samplers in multi-modality.
Radiomics aims to extract and analyze large numbers of quantitative features from medical images and is highly promising in staging, diagnosing, and predicting outcomes of cancer treatments. Nevertheless, several challenges need to be addressed to construct an optimal radiomics predictive model. First, the predictive p…
CLIP learns joint image-text representations for zero-shot learning.
problem Understanding and improving zero-shot transfer performance in CLIP.
method Formal study of transferrable representation learning and analysis of zero-shot transfer performance.
result Proposes a new CLIP-type approach that outperforms existing methods.
Generates detailed fashion feedback from outfit images.
problem Creating informative and diverse fashion feedback from outfit images.
method Trained deep generative models with visual attention, then improved with Maximum Mutual Information objective function.
result Generated sentences are more diverse and detailed.
Deep learning enhances art market valuation by incorporating visual data.
problem Improving valuation accuracy in the art market, especially for first-time sales.
method Benchmarked classical and modern deep learning models using a large auction dataset.
result Visual embeddings add distinct economic value for first-time art sales.
AF improves sampling from high-dimensional, multi-modal distributions.
problem Sampling from high-dimensional, multi-modal distributions is challenging.
method Annealing Flow (AF) using Continuous Normalizing Flow (CNF) with dynamic Optimal Transport (OT) objective and annealing procedures.
result AF significantly improves training efficiency and stability, outperforming state-of-the-art methods.
Tool detects tax evasion on social media using multi-modal deep learning.
problem Detecting tax evasion on social media platforms.
method Developed a multi-modal deep neural network combining comments, hashtags, and images.
result Multi-modal deep neural network achieved AUC of 0.808 and F1 score of 0.762.
Multiple modalities often co-occur when describing natural phenomena. Learning a joint representation of these modalities should yield deeper and more useful representations. Previous generative approaches to multi-modal input either do not learn a joint distribution or require additional computation to handle missing …
Develops multi-modal neural network models for improved prediction and uncertainty quantification.
problem Improving prediction accuracy and uncertainty quantification for multi-modal data.
method Multi-modal Bayesian neural network models with conjugate last-layer estimation using SVI.
result Improved prediction accuracy and uncertainty quantification compared to uni-modal models.
METEOR learns efficient representations from multi-modal data streams.
problem Efficiently interpreting multi-modal information in complex environments.
method METEOR learns compact representations by sharing parameters within semantically meaningful groups and preserving domain-agnostic semantics.
result METEOR reduces memory usage by around 80% compared to conventional methods.
New framework shows cross-attention improves multi-modal in-context learning.
problem Understanding multi-modal in-context learning in neural networks.
method Mathematical framework and linearized cross-attention mechanism.
result Cross-attention mechanism is provably optimal for multi-modal in-context learning.
A novel multi-modal active learning approach using RL for engagement estimation.
problem Challenges in labeling multi-modal human data for accurate user state estimation.
method Deep reinforcement learning for optimal data selection and multi-modal data fusion.
result The proposed approach outperforms existing methods in engagement estimation.