This research improves multimodal systems by adding a second objective and regularisation methods.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study analyzes bias interactions in multimodal models using simulation-based methods.
Unified model predicts stock and systemic risks from diverse financial data.
VAEs struggle with surjective multimodal data, especially class labels describing images.
End-to-end CAD system for thyroid nodule classification using multimodal data and expert guidance.
Computational modeling of human multimodal language is an emerging research area in natural language processing spanning the language, visual and acoustic modalities. Comprehending multimodal language requires modeling not only the interactions within each modality (intra-modal interactions) but more importantly the in…
Characterizing the dynamic interactive patterns of complex systems helps gain in-depth understanding of how components interrelate with each other while performing certain functions as a whole. In this study, we present a novel multimodal data fusion approach to construct a complex network, which models the interaction…
Study quantifies interactions between unlabeled multimodal data.
DAM improves cryptocurrency trend forecasting using multimodal data.
Estimates interactions between modalities for multimodal data.
New algorithm learns switching dynamics from multiple neural signals.
Paper proposes multimodal contrastive learning for EHR data.
AECF improves multimodal inference robustness and calibration.
In this paper we seek methods to effectively detect urban micro-events. Urban micro-events are events which occur in cities, have limited geographical coverage and typically affect only a small group of citizens. Because of their scale these are difficult to identify in most data sources. However, by using citizen sens…
A new mutual information lower bound for multimodal regression active learning.
Generates multimodal safety-critical scenarios for robustness evaluation of decision-making algorithms.
UniFinEval benchmarks financial models across text, images, and videos.
DeepSIP predicts network failures' impact using CNN from syslog and traffic data.
Study predicts traffic congestion based on population mobility data.
RAHMC improves sampling from multimodal distributions using dissipative dynamics.
Modified PCA algorithm with continual learning preserves features of previous modes for multimode process monitoring.
The problem of information fusion from multiple data-sets acquired by multimodal sensors has drawn significant research attention over the years. In this paper, we focus on a particular problem setting consisting of a physical phenomenon or a system of interest observed by multiple sensors. We assume that all sensors m…
Representation learning becomes especially important for complex systems with multimodal data sources such as cameras or sensors. Recent advances in reinforcement learning and optimal control make it possible to design control algorithms on these latent representations, but the field still lacks a large-scale standard …
FinAgent tackles financial trading with multimodal data and advanced AI.
Proposes a novel network for CTR prediction by learning modality-specific and modality-invariant representations.
Multimodal learning has shown promising performance in content-based recommendation due to the auxiliary user and item information of multiple modalities such as text and images. However, the problem of incomplete and missing modality is rarely explored and most existing methods fail in learning a recommendation model …
Paper reviews the evolution of alpha from human insight to AI-powered systems.
Current high-throughput data acquisition technologies probe dynamical systems with different imaging modalities, generating massive data sets at different spatial and temporal resolutions posing challenging problems in multimodal data fusion. A case in point is the attempt to parse out the brain structures and networks…
Multimodal machine learning is a core research area spanning the language, visual and acoustic modalities. The central challenge in multimodal learning involves learning representations that can process and relate information from multiple modalities. In this paper, we propose two methods for unsupervised learning of j…
Automatic analysis of teacher and student interactions could be very important to improve the quality of teaching and student engagement. However, despite some recent progress in utilizing multimodal data for teaching and learning analytics, a thorough analysis of a rich multimodal dataset coming for a complex real lea…
In this paper, we propose to employ a bank of modality-dedicated Convolutional Neural Networks (CNNs), fuse, train, and optimize them together for person classification tasks. A modality-dedicated CNN is used for each modality to extract modality-specific features. We demonstrate that, rather than spatial fusion at the…
Survey of multimodal deep generative models for diverse data types.
Driver vigilance estimation is an important task for transportation safety. Wearable and portable brain-computer interface devices provide a powerful means for real-time monitoring of the vigilance level of drivers to help with avoiding distracted or impaired driving. In this paper, we propose a novel multimodal archit…
New model detects Alzheimer's and severity from speech, cognitive, and language data.
Multimodal fusion is considered a key step in multimodal tasks such as sentiment analysis, emotion detection, question answering, and others. Most of the recent work on multimodal fusion does not guarantee the fidelity of the multimodal representation with respect to the unimodal representations. In this paper, we prop…
Multimodal research is an emerging field of artificial intelligence, and one of the main research problems in this field is multimodal fusion. The fusion of multimodal data is the process of integrating multiple unimodal representations into one compact multimodal representation. Previous research in this field has exp…
Benchmarking AutoML for tables with text fields, achieving top performance.
Develops a contrastive framework for data-efficient multimodal learning.
Plug-and-play multimodal controller improves class-conditional image generation.
We study an agent-based model of evolution of wealth distribution in a macro-economic system. The evolution is driven by multiplicative stochastic fluctuations governed by the law of proportionate growth and interactions between agents. We are mainly interested in interactions increasing wealth inequality that is in a …
Learning multimodal representations is a fundamentally complex research problem due to the presence of multiple heterogeneous sources of information. Although the presence of multiple modalities provides additional valuable information, there are two key challenges to address when learning from multimodal data: 1) mode…
Improved sentiment analysis with multimodal data.
Multimodalities provide promising performance than unimodality in most tasks. However, learning the semantic of the representations from multimodalities efficiently is extremely challenging. To tackle this, we propose the Transformer based Cross-modal Translator (TCT) to learn unimodal sequence representations by trans…
Humor is a unique and creative communicative behavior displayed during social interactions. It is produced in a multimodal manner, through the usage of words (text), gestures (vision) and prosodic cues (acoustic). Understanding humor from these three modalities falls within boundaries of multimodal language; a recent r…
VLM judges rank well but score poorly; task difficulty and annotation quality affect interval width.
AV-ASR system improves speech recognition with visual context.
Self-supervised bidirectional transformer models such as BERT have led to dramatic improvements in a wide variety of textual classification tasks. The modern digital world is increasingly multimodal, however, and textual information is often accompanied by other modalities such as images. We introduce a supervised mult…
The complex world around us is inherently multimodal and sequential (continuous). Information is scattered across different modalities and requires multiple continuous sensors to be captured. As machine learning leaps towards better generalization to real world, multimodal sequential learning becomes a fundamental rese…