Paper proposes multi-task learning for multi-modal video Q&A.
problem Expensive to create large-scale datasets for multi-modal video Q&A.
method Composed of three networks: video Q&A, temporal retrieval, and modality alignment.
result State-of-the-art results on TVQA dataset.
Method trains multi-modal policy from unlabeled mixed demonstrations.
problem Training policies from unlabeled mixed demonstrations.
method Variational autoencoder with categorical latent variable to discover latent factors of variation.
result Policy can reproduce specific behaviors by conditioning on categorical vectors.
Proposes LM3FE for multi-modal feature extraction in image classification.
problem High-dimensional features and multi-modal data challenges.
method Large margin multi-modal multi-task feature extraction (LM3FE) framework.
result LM3FE outperforms single-task feature extraction and multi-modal feature extraction.
ESE-FN improves elderly activity recognition accuracy.
problem Recognizing individual actions and human-object interactions in elderly activities.
method Exploits multi-modal features from RGB videos and skeleton sequences using ESE attentions and a new Multi-modal Loss.
result ESE-FN achieves best accuracy on ETRI-Activity3D dataset.
MEx dataset benchmarks HAR and multi-modal fusion for exercise quality.
problem Recognizing and evaluating exercise quality for Musculoskeletal Disorders patients.
method Multi-sensor, multi-modal dataset with four sensors (pressure mat, depth camera, accelerometers) for HAR and exercise quality assessment.
result Reference performance for each sensor identified, exposing their strengths and weaknesses.
This paper proposes MM-DAGs for analyzing traffic congestion, learning multiple DAGs jointly.
problem Analyzing multi-modal traffic data with overlapping and distinct variables.
method Developed MM-DAGs for multi-task, multi-modal DAG learning, using multi-modal regression and CD measure.
result Proved the effectiveness of MM-DAGs in traffic congestion analysis.
This paper reviews deep learning for multi-modality medical image segmentation.
problem Improving segmentation accuracy in medical images using multiple modalities.
method Overview of deep learning and multi-modal medical image segmentation, analysis of different network architectures and fusion strategies.
result Later fusion of modalities can lead to more accurate segmentation results.
Paper develops a theory explaining contrastive pre-training for multimodal AI.
problem Limited theoretical understanding of contrastive pre-training for multi-modal AI.
method Introduces approximate sufficient statistics and Joint Generative Hierarchical Model.
result Near-minimizers of contrastive loss are approximately sufficient, enabling diverse downstream tasks.
MoCA uses a novel autoencoder to analyze multi-modal health data.
problem Challenges in analyzing continuous multi-modal health data from wearable devices.
method Proposes MoCA, a self-supervised learning framework combining transformer and masked autoencoder methods.
result Demonstrates strong performance boosts across reconstruction and classification tasks.
Develops multi-modal neural network models for improved prediction and uncertainty quantification.
problem Improving prediction accuracy and uncertainty quantification for multi-modal data.
method Multi-modal Bayesian neural network models with conjugate last-layer estimation using SVI.
result Improved prediction accuracy and uncertainty quantification compared to uni-modal models.
DMAE learns shared latent space from unpaired data.
problem Learning shared latent space from unpaired multi-modal data.
method Formulates cross-domain representation learning and object matching problem, optimizes autoencoders and pairing.
result Promising results in image captioning and unsupervised classifier learning.
New method for multi-modal depth prediction challenges.
problem Challenges in applying unimodal coreset selection to multi-modal data.
method Adapted state-of-the-art coreset selection technique for multimodal data.
result Challenges in extending unimodal algorithms to multi-modal scenarios.
Unified architecture for multi-modal multi-task learning using transformer.
problem Training multiple tasks concurrently with varying modalities.
method Spatio-temporal cache mechanism for multi-modal learning.
result Training multiple tasks together reduces model size by about three times.
This paper improves sentiment classification by combining text, audio, and video data using DCCA.
problem Improving sentiment classification accuracy using multi-modal data.
method Deep Canonical Correlation Analysis (DCCA) for combining text, audio, and video embeddings.
result One-Step DCCA outperforms current state-of-the-art in multi-modal embedding learning.
Advances citation and subject label recommendation using multi-modal adversarial autoencoders.
problem Improving recommendation systems for citations and subject labels.
method Multi-modal adversarial autoencoders with adversarial regularization, sparsity, and input modality analysis.
result Adversarial regularization consistently improves recommendation performance.
New framework shows cross-attention improves multi-modal in-context learning.
problem Understanding multi-modal in-context learning in neural networks.
method Mathematical framework and linearized cross-attention mechanism.
result Cross-attention mechanism is provably optimal for multi-modal in-context learning.
This work tackles uncertainty in multi-agent multi-modal trajectory forecasting.
problem Measuring and ranking uncertainty in multi-agent multi-modal trajectory forecasting.
method Proposes collaborative uncertainty (CU) and a CU-aware regression framework.
result The CU-aware regression framework improves SOTA systems' performances.
Adaptive anchor methods improve multi-modal learning by balancing intra-modal and inter-modal information.
problem Fixed anchor methods limit multi-modal learning by over-reliance on a single modality and inadequate cross-modal correlation.
method Adaptive anchor methods using centroid-based anchors from all modalities.
result Adaptive anchor methods like CentroBind consistently outperform fixed anchor methods across various datasets.
Framework for causal discovery using multi-modal data.
problem Failure of representation learning in causal tasks.
method Statistical and computational framework combining representation learning and causal inference.
result Effective use of observational and perturbational data for causal discovery.
FinTMMBench benchmarks RAG systems for finance tasks across multiple data types and time periods.
problem Evaluating temporal-aware multi-modal retrieval augmented generation in finance.
method TMMHybridRAG method that converts and integrates data from various modalities and temporal information.
result Demonstrated effectiveness of TMMHybridRAG in diverse financial analysis tasks.
METEOR learns efficient representations from multi-modal data streams.
problem Efficiently interpreting multi-modal information in complex environments.
method METEOR learns compact representations by sharing parameters within semantically meaningful groups and preserving domain-agnostic semantics.
result METEOR reduces memory usage by around 80% compared to conventional methods.
A novel multi-modal active learning approach using RL for engagement estimation.
problem Challenges in labeling multi-modal human data for accurate user state estimation.
method Deep reinforcement learning for optimal data selection and multi-modal data fusion.
result The proposed approach outperforms existing methods in engagement estimation.
Proposes a flexible normalization method to handle multi-modal data.
problem Reduced effectiveness of batch normalization in multi-modal distributions.
method Extends normalization to multiple means and variances, detecting data modes on-the-fly.
result Outperforms batch normalization and other methods in various experiments.
This paper introduces NPR, a technique to improve Bayesian inference for multi-modal, high-dimensional simulations.
problem Challenges in Bayesian inference for multi-modal, high-dimensional simulations.
method Introduces Neural Posterior Regularization (NPR) to enforce exploration of input parameter space.
result Empirically validated that NPR significantly improves performance on various simulation tasks.
This paper improves ensemble learning for vision tasks by encouraging diversity in predictions.
problem Generating effective ensembles of neural networks for multi-modal data.
method Explicitly optimize a diversity inducing adversarial loss for learning stochastic latent variables.
result Significant improvements in classification accuracy and out-of-distribution detection compared to baselines.
New method trains neural samplers to sample from multi-modal distributions efficiently.
problem Mode-seeking behavior of reverse KL divergence hinders effective sampling from multi-modal target distributions.
method Minimizing reverse diffusive KL divergence along diffusion trajectories of model and target densities.
result Demonstrated enhanced sampling performance across various multi-modal distributions.
Unified Bayesian model for multi-modal, small sample size biomedical data classification.
problem Classifying high-dimensional, multi-modal biomedical data with small sample sizes.
method Combines multi-modal data views into a latent space, prunes irrelevant features, and uses dual kernels for small sample size scenarios.
result Outperforms state-of-the-art models and identifies features aligned with existing markers.
This paper tackles multi-modal label disentanglement in partition-based XMC.
problem Existing partition-based XMC methods create mutually exclusive clusters, which is sub-optimal for multi-modal labels.
method Formulates label assignment as an optimization problem to maximize precision rates, creating flexible and overlapped label clusters.
result Successfully disentangles multi-modal labels, leading to state-of-the-art results on XMC benchmarks.
DiGS improves sampling from multi-modal distributions.
problem Inadequate mixing in MCMC methods for multi-modal distributions.
method Integrates diffusion models and Gibbs sampling to create an auxiliary noisy distribution.
result DiGS exhibits better mixing for multi-modal distributions than state-of-the-art methods.
Paper presents ECL dataset for multi-modal bankruptcy prediction.
problem Developing models to predict corporate bankruptcy using textual and numerical data.
method Used ECL dataset to develop and evaluate classical and neural models.
result Textual and numerical data modalities complement each other for bankruptcy prediction.
Diffusion models learn multi-modal distributions with optimal efficiency.
problem Learning high-dimensional distributions with low-dimensional multi-modal structures.
method Score-based diffusion models, focusing on subgaussian distributions within subspaces.
result Diffusion models require O ~ ( ε − k ∨ 2 ) \widetilde{O}(\varepsilon^{-k \vee 2}) O ( ε − k ∨ 2 ) samples for 1-Wasserstein ε \varepsilon ε error, improving over prior guarantees. MultiPath predicts multi-modal future trajectories for better motion planning.
problem Predicting human behavior in uncertain real-world domains like autonomous driving.
method Leverages fixed future state-sequence anchors and regresses offsets with uncertainties.
result Achieves more accurate predictions with an order of magnitude fewer trajectories.
This work improves manifold learning for multi-modal data.
problem Distortions and modeling errors in multi-modal data.
method Isometrizing learned Riemannian structure and balancing regularity and expressivity.
result The synergy of proposed approaches enhances manifold learning.
Improved student engagement detection using contextual and visual data.
problem Detecting students' behavioral engagement in real-world settings.
method Two-phase approach: contextual logs for active use, appearance information for engagement inference.
result Improved F1-scores from 0.77 to 0.82 with contextual information.
FinVision uses LLM agents to predict stock markets by processing various financial data types.
problem Challenges in integrating diverse financial data for accurate stock market prediction.
method Multi-agent framework with LLMs specialized in different financial data types and a reflection module.
result The reflection module enhances decision-making capabilities for financial trading.
System diagnoses Alzheimer's disease from spoken language using multi-modal features.
problem Early diagnosis of Alzheimer's disease from spoken language.
method Classification system based on spoken language using three approaches (N-gram, i-vector, x-vector).
result Accuracy of 83.6% on the cookie picture description task from Pitt Corpus dementia bank.
A new density model using Fourier basis achieves better approximations and compression.
problem Approximating multi-modal 1D densities.
method Constrained Fourier basis model for end-to-end training.
result Lower cross entropy compared to deep factorized models.
Stein-Encoder isolates genetic signals in multi-modal biomedical data.
problem Integration of high-dimensional genomic data with clinical data obscures genetic predictive impact.
method White-box supervised framework using Stein's method and residualization.
result Stein-Encoder improves predictive accuracy and reveals specific biological mechanisms.
Paper proposes a multi-modal probabilistic prediction model for interactive behavior.
problem Predicting future motions of interacting entities in real-world scenarios.
method Generative model for joint prediction of sequential motions of interacting agents.
result Interpretable model capable of handling prediction uncertainties and multi-modal distributions.
A new method learns from multi-modal sequences with external memory.
problem Learning new modes in a dynamic environment without prior knowledge.
method Maintains a neural episodic memory with a Dirichlet Process prior to store mode descriptors and transfers knowledge through retrieval.
result Performs continual learning favorably compared to mainstream approaches.
New method improves sampling from multi-modal distributions using Langevin diffusion and simulated tempering.
problem Sampling from multi-modal distributions is challenging and slow.
method Combining Langevin diffusion with simulated tempering.
result The new method mixes more rapidly and samples from distributions close to mixtures of gaussians.
Study shows agents benefit from hearing in addition to vision.
problem Limited effectiveness of vision-only reinforcement learning agents.
method Used audio as complementary information to visual cues in state representation.
result Agents perform better when hearing is added to vision.
MTRGL learns temporal correlations from multi-modal data for improved pair trading.
problem Discerning temporal correlations among financial entities.
method Combines time series data and discrete features into a temporal graph, using a memory-based temporal graph neural network.
result MTRGL outperforms traditional methods in temporal graph link prediction and pair trading.
Paper tackles cross-modal anomalies in multi-source data.
problem Detect anomalies in multi-modal data where patterns are inconsistent across different sources.
method Proposes a deep structured anomaly detection framework.
result Demonstrates effectiveness on real-world datasets.
The paper analyzes and proposes an algorithm for multi-modal nonlinear embeddings with theoretical performance bounds.
problem Generalizability of multi-modal nonlinear embeddings to unseen data.
method Theoretical analysis and a multi-modal nonlinear representation learning algorithm motivated by performance bounds.
result The proposed algorithm yields promising performance in multi-modal image classification and cross-modal image-text retrieval applications.
Paper proposes a new framework for robust multi-modal data fusion under uncertainty.
problem Unexpected modality failures in nonlinear non-Gaussian dynamic processes.
method Dynamic model averaging (DMA) based particle filter (PF) algorithm.
result The proposed solution outperforms state-of-the-art methods in experiments.
C-qGAN learns multi-modal distributions efficiently.
problem Learning multi-modal distributions efficiently.
method Conditional Quantum Generative Adversarial Network (C-qGAN) within quantum circuits.
result C-qGAN outperforms current state preparation methods in efficiency.
New method minimizes f-divergence for better imitation learning.
problem Learning from multi-modal demonstrations, especially interpolating between modes.
method Minimizing reverse KL divergence or I-projection for any f-divergence.
result Our method reliably imitates multi-modal behaviors better than existing methods.