A new method uses CML to speed up social signal annotation.
problem Manual annotation of social signals in large multi-modal corpora is time-consuming and exhausting.
method CML techniques to predict local parts first, followed by a session-independent classification model.
result The method has the potential to significantly reduce human labelling efforts.
Multi-modal data collections, such as corpora of paired images and text snippets, require analysis methods beyond single-view component and topic models. For continuous observations the current dominant approach is based on extensions of canonical correlation analysis, factorizing the variation into components shared b…
Research creates parallel sentences from comparable corpora for machine translation.
problem Scarce parallel sentences limit cross-lingual applications.
method Web crawling and analogy-based heuristics for building parallel corpora.
result Improved machine translation on Polish-English pairs.
Flexible SC framework converts voices from non-aligned corpora.
problem Limited practical applications of SC due to lack of parallel corpora.
method Variational auto-encoder framework for non-parallel corpora.
result Framework enables spectral conversion without parallel corpora or alignments.
Study develops methods for creating parallel corpora from comparable texts.
problem Lack of parallel texts for machine translation.
method Automatic web crawling and unsupervised methods for comparable corpora.
result Improved quality of parallel corpora for machine translation.
Modeling lead-lag relationship between two text corpora for improved topic modeling.
problem Recognizing the relationship between multiple text corpora for better topic modeling.
method Proposed a jointly dynamic topic model and embedding extension for large-scale text corpus.
result The proposed model can well recognize the lead-lag relationship between two text corpora and improve topic learning.
Deep models predict missing product attributes from text and images.
problem Incomplete or missing product attributes in e-commerce catalogs.
method Combining textual and visual data with a novel modality-merging method.
result Our approach improves attribute prediction on Rakuten-Ichiba and other datasets.
We develop the multilingual topic model for unaligned text (MuTo), a probabilistic model of text that is designed to analyze corpora composed of documents in two languages. From these documents, MuTo uses stochastic EM to simultaneously discover both a matching between the languages and multilingual latent topics. We d…
Enhanced SMT systems for diverse language pairs using comparable corpora.
problem Improving SMT quality on diverse language pairs.
method Translation model training, adaptation of training settings, comparable corpora, domain adaptation, symmetrized word alignment, unsupervised transliteration, KenLM language modeling.
result Our approach improved SMT quality on diverse language pairs.
Efficiently trains large corpora models without sampling.
problem Training neural network embedding models on very large corpora using SGD is expensive.
method Proposes new methods to train models without sampling unobserved pairs, using Gramian estimation and variance reduction schemes.
result Significant improvement in training time and generalization quality compared to traditional methods.
The paper analyzes and proposes an algorithm for multi-modal nonlinear embeddings with theoretical performance bounds.
problem Generalizability of multi-modal nonlinear embeddings to unseen data.
method Theoretical analysis and a multi-modal nonlinear representation learning algorithm motivated by performance bounds.
result The proposed algorithm yields promising performance in multi-modal image classification and cross-modal image-text retrieval applications.
Ultra-fast search algorithm for trillion-scale corpora with semantic flexibility.
problem Efficiently searching over large natural language corpora with semantic variations.
method String matching based on suffix arrays, vector representation of words, dynamic corpus-aware pruning, fast exact lookup.
result Substantially lower search latency compared to existing methods on FineWeb-Edu corpus.
Paper proposes continual learning for sentence encoders.
problem Optimize sentence encoders for new corpora while maintaining old corpus accuracy.
method Initialize encoders with corpus-independent features, update using Boolean operations of conceptor matrices.
result Proposed sentence encoder can continually learn features from new corpora.
ROME improves density estimation for multi-modal, non-normal data.
problem Robust multi-modal density estimation in non-normal, highly correlated distributions.
method ROME uses clustering to segment multi-modal data into uni-modal clusters, then combines KDE estimates for each cluster.
result ROME outperforms state-of-the-art methods and is more robust to various distributions.
DAPPER improves scalability of DAP topic model for large corpora.
problem Scaling complex models like DAP to large text corpora.
method Adapted approximate inference techniques for DAP, developing CVI-based EM.
result Significant improvements in model fit and training time without compromising structure.
Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our methodology for mining such data from previously obtained comparable corpora. The task is highly practical since non-parallel m…
MoCA uses a novel autoencoder to analyze multi-modal health data.
problem Challenges in analyzing continuous multi-modal health data from wearable devices.
method Proposes MoCA, a self-supervised learning framework combining transformer and masked autoencoder methods.
result Demonstrates strong performance boosts across reconstruction and classification tasks.
The paper evaluates samplers on multi-modal targets, focusing on mode separation and recovery.
problem Handling multi-modality in sampling.
method Synthetic experimental setting focusing on mode relative importance recovery.
result Illustrates the challenges and potential of samplers in multi-modality.
Improves topic modeling for multi-collection corpora.
problem Challenges in text mining from multi-collection corpora.
method Compound Latent Dirichlet Allocation (cLDA) model and MCMC methods.
result cLDA model identifies topic proportions across multiple collections.
Improved neural NER by optimizing large corpora for German.
problem Low-resource language named entity recognition.
method Optimized large corpora, lemmatization, part-of-speech tagging, and detailed optimization.
result Up to 11% improvement in F-score on German NER tasks.
AF improves sampling from high-dimensional, multi-modal distributions.
problem Sampling from high-dimensional, multi-modal distributions is challenging.
method Annealing Flow (AF) using Continuous Normalizing Flow (CNF) with dynamic Optimal Transport (OT) objective and annealing procedures.
result AF significantly improves training efficiency and stability, outperforming state-of-the-art methods.
This paper proposes an efficient method to train word embeddings for large corpora without synchronization.
problem Training word embeddings for large text corpora is computationally expensive and requires synchronization.
method Partition the input space instead of the vocabulary size, using asynchronous training without parameter synchronization.
result Comparable and up to 45% performance improvement in NLP benchmarks with 1/10 the training time.
Test evaluates NMF-based topic models for document corpora.
problem Violation of likelihood assumptions in NMF topic models.
method Double parametric bootstrap test based on KL divergence and Poisson ML.
result Correctly identifies reliable NMF-based topic models.
Novel fuzzy topic model improves health corpus retrieval.
problem Automatic retrieval of health and medical knowledge from text.
method Fuzzy Latent Semantic Analysis (FLSA) for topic discovery.
result FLSA outperforms LDA in estimating topics and improving retrieval.
Develops multi-modal neural network models for improved prediction and uncertainty quantification.
problem Improving prediction accuracy and uncertainty quantification for multi-modal data.
method Multi-modal Bayesian neural network models with conjugate last-layer estimation using SVI.
result Improved prediction accuracy and uncertainty quantification compared to uni-modal models.
METEOR learns efficient representations from multi-modal data streams.
problem Efficiently interpreting multi-modal information in complex environments.
method METEOR learns compact representations by sharing parameters within semantically meaningful groups and preserving domain-agnostic semantics.
result METEOR reduces memory usage by around 80% compared to conventional methods.
Framework for handling long-tailed multi-modal data.
problem Class imbalance and long-tailed distributions in multi-modal data.
method Multi-expert architecture with modality-specific networks and dynamic fusion weights.
result Framework outperforms existing methods in long-tailed, class-imbalanced scenarios.
New framework shows cross-attention improves multi-modal in-context learning.
problem Understanding multi-modal in-context learning in neural networks.
method Mathematical framework and linearized cross-attention mechanism.
result Cross-attention mechanism is provably optimal for multi-modal in-context learning.
A novel multi-modal active learning approach using RL for engagement estimation.
problem Challenges in labeling multi-modal human data for accurate user state estimation.
method Deep reinforcement learning for optimal data selection and multi-modal data fusion.
result The proposed approach outperforms existing methods in engagement estimation.
Inductive graph-based approach for disease classification with incomplete data.
problem Classifying patients with incomplete multi-modal data.
method Multi-modal graph fusion trained end-to-end for node-level classification.
result Outperforms single static graph approach in multi-modal disease classification.
Open-FinLLMs tackle financial tasks with multimodal capabilities.
problem Financial LLMs lack multimodal capabilities and real-world applicability.
method Developed Open-FinLLMs, an open-source multimodal financial LLM suite.
result Open-FinLLMs outperform advanced financial and general LLMs in diverse tasks.
This paper improves sentiment classification by combining text, audio, and video data using DCCA.
problem Improving sentiment classification accuracy using multi-modal data.
method Deep Canonical Correlation Analysis (DCCA) for combining text, audio, and video embeddings.
result One-Step DCCA outperforms current state-of-the-art in multi-modal embedding learning.
Contrastive learning adapts to data intrinsic dimensions, learning low-dimensional representations.
problem Learning high-dimensional representations from multi-modal data.
method Multi-modal contrastive learning with temperature optimization.
result Contrastive learning adapts to intrinsic dimensions of data, not specified dimensions.
Adaptive anchor methods improve multi-modal learning by balancing intra-modal and inter-modal information.
problem Fixed anchor methods limit multi-modal learning by over-reliance on a single modality and inadequate cross-modal correlation.
method Adaptive anchor methods using centroid-based anchors from all modalities.
result Adaptive anchor methods like CentroBind consistently outperform fixed anchor methods across various datasets.
PyTorch Frame simplifies multi-modal tabular learning with modular data and model handling.
problem Handling complex multi-modal tabular data in deep learning.
method A PyTorch-based framework that provides a data structure, model abstraction, and integration with external models.
result Demonstrated the effectiveness of PyTorch Frame in implementing and applying diverse tabular models to complex multi-modal tabular data.
Anomaly detection in multi-modal data using cyclostationary models and neural networks.
problem Detect anomalies in multi-modal data like CCTV imagery and social media posts.
method Deep neural network for object detection, cyclostationary model for regular patterns, sequential anomaly detection algorithms.
result Asymptotically efficient anomaly detection algorithms applied to NYC 5K run detection.
MEx dataset benchmarks HAR and multi-modal fusion for exercise quality.
problem Recognizing and evaluating exercise quality for Musculoskeletal Disorders patients.
method Multi-sensor, multi-modal dataset with four sensors (pressure mat, depth camera, accelerometers) for HAR and exercise quality assessment.
result Reference performance for each sensor identified, exposing their strengths and weaknesses.
Paper proposes multi-task learning for multi-modal video Q&A.
problem Expensive to create large-scale datasets for multi-modal video Q&A.
method Composed of three networks: video Q&A, temporal retrieval, and modality alignment.
result State-of-the-art results on TVQA dataset.
Extends Gaussian Processes for multi-modal, non-stationary data.
problem Modeling non-stationary multi-modal processes.
method Adds a latent variable to modulate covariance over training data.
result Shows improved modeling of multi-modal and non-stationary processes.
Improves comparable corpora mining for better translation quality.
problem Limited availability of high-quality bilingual data for translation.
method Reimplemented comparison algorithms, introduced tuning script, GPU acceleration.
result Positive impact on quality and quantity of mined data, improved translation quality.
Derives M2VAE objective from marginal joint log-likelihood.
problem Training Multi-Modal Variational Autoencoders (M2VAEs). method Derives trainable evidence lower bound from marginal joint log-likelihood.
result Derives M2VAE objective from marginal joint log-likelihood. Method trains multi-modal policy from unlabeled mixed demonstrations.
problem Training policies from unlabeled mixed demonstrations.
method Variational autoencoder with categorical latent variable to discover latent factors of variation.
result Policy can reproduce specific behaviors by conditioning on categorical vectors.
This work improves multi-modal generative models by using permutation-invariant neural networks.
problem Improving multi-modal generative models with tighter variational objectives.
method Developed more flexible aggregation schemes based on permutation-invariant neural networks.
result Our variational objective and flexible aggregation models can better approximate the true joint distribution.
MMVAE learns multi-modal data with shared and private latent spaces.
problem Learning useful representations across multiple data modalities.
method Mixture-of-experts variational autoencoder (MMVAE).
result MMVAE satisfies four criteria for multi-modal learning.
Two methods generate parallel data for GEC, improving neural models' performance.
problem Lack of parallel data for GEC.
method Two approaches to generate large parallel datasets from Wikipedia data.
result Neural GEC models trained on generated corpora perform similarly and surpass state-of-the-art.
This paper reviews deep learning for multi-modality medical image segmentation.
problem Improving segmentation accuracy in medical images using multiple modalities.
method Overview of deep learning and multi-modal medical image segmentation, analysis of different network architectures and fusion strategies.
result Later fusion of modalities can lead to more accurate segmentation results.
Method predicts NBA players' multi-modal movement trajectories.
problem Understanding NBA players' decision-making during games.
method LSTM-based architecture with multi-modal loss function.
result Method outperforms state-of-the-art in predicting realistic trajectories.
Improves MCMC sampling for multi-modal distributions.
problem Inefficient MCMC sampling in multi-modal posterior distributions.
method Pseudo-extended MCMC method using auxiliary variables.
result Improved MCMC sampling over Hamiltonian Monte Carlo.