This work investigates upper bounds for visual captioning without extra effort.
problem Establishing performance upper bounds for visual captioning datasets.
method Decomposing visual captioning into two steps and constructing upper bounds.
result Current state-of-the-art models fall short of learned upper bounds.
Paper develops a framework for generating coherent image captions using visual features and hierarchical topics.
problem Generating semantically coherent paragraphs to describe image content.
method Plug-and-play hierarchical-topic-guided image paragraph generation framework integrating visual extractor and deep topic model.
result Proposed models can distill interpretable multi-layer semantic topics and generate diverse and coherent captions.
New 2D maps improve image captioning models.
problem Captions generated by RNNs are often poor.
method Used 2D maps instead of vectors to represent latent states.
result 2D maps lead to better captioning performance.
Improved video and movie description using multitask learning.
problem Lack of training data and poor generalization in video captioning.
method Multitask learning encoder-decoder framework for video sequences.
result Improved performance on multi-caption and single-caption datasets.
Grad-CAM visualizes CNN predictions to explain model decisions.
problem Making CNN models transparent and understandable.
method Using class-specific gradient information to localize important regions.
result Improved understanding of CNN-based models, including image captioning and VQA.
Recent progress on automatic generation of image captions has shown that it is possible to describe the most salient information conveyed by images with accurate and meaningful sentences. In this paper, we propose an image caption system that exploits the parallel structures between images and sentences. In our model, …
Pretrained model improves visual dialog performance.
problem Improving performance in visual dialog tasks.
method Pretrained ViLBERT model on vision-language datasets, fine-tuned on VisDial.
result Best model outperforms prior work by more than 1% on NDCG and MRR.
Enhances image captioning with novel context combination methods.
problem Improving machine learning for image captioning with structured learning and meaningful interpretation.
method Combines Feature Distribution Composition (FDC), Multiple Role Representation Crossover (MRRC) attention layers, and language decoder.
result Significantly improved image captioning performance (35.3%) and established new standards.
This paper investigates uncertainty calibration in multimodal large language models.
problem Challenges in properly calibrating uncertainty in multimodal large language models.
method Investigation of representative MLLMs across various scenarios, including visual fine-tuning and multimodal training.
result MLLMs tend to give answers rather than admit uncertainty, but this self-assessment improves with proper prompt adjustments.
New model improves image captioning's ability to describe unseen concepts.
problem Image captioning models struggle with describing unseen combinations of concepts.
method Proposes a multi-task model combining caption generation and image-sentence ranking, with a decoding mechanism to re-rank captions based on image similarity.
result The model significantly outperforms state-of-the-art models in compositional generalization.
DBNet improves natural language image localization and detection.
problem Natural language-based visual entity localization with limited accuracy.
method Discriminative bimodal neural network (DBNet) trained with extensive negative samples.
result Significantly outperforms previous methods on Visual Genome dataset.
Improves text-to-image generation with bidirectional capabilities.
problem Generating realistic images from text descriptions.
method Integrates text and image modalities using MMVR architecture with n-gram cost function and multiple sentences.
result Significant improvement in image quality over existing methods (over 20%).
Survey of deep learning methods for image captioning.
problem Generating accurate and complex image descriptions.
method Comprehensive review of deep learning techniques for image captioning.
result Analysis of strengths, limitations, and popular datasets in deep learning image captioning.
Survey on GANs for generating visual arts, music, and literature.
problem Tackles the challenge of generating art using GANs.
method Uses generative adversarial networks (GANs) to generate visual arts, music, and literary text.
result Performance comparison and description of various GAN architectures presented.
Bayesian approach improves image captioning quality metrics.
problem Improving image captioning quality metrics.
method Bayesian Self-Critical Sequence Training (B-SCST) with Monte Carlo dropout approximate variational inference.
result B-SCST improves CIDEr-D scores on various image captioning datasets.
Generative model uses captions to generate images, improving semantic understanding.
problem Complex image generation models require large datasets and intricate learning.
method Adapts captioning models to generate images, using learned sentence and frame vectors.
result Images generated from multiple captions better capture semantic meaning.
C4Synth generates images from multiple captions to improve image quality.
problem Generating images from a single caption is limited; multiple captions are needed.
method Two deep generative models that ensure 'Cross-Caption Cycle Consistency'.
result Quantitative and qualitative validation on Caltech-UCSD Birds and Oxford-102 Flowers datasets.
Current image captioning systems miss spatial location details.
problem Capturing spatial location information in image captions.
method Evaluation of image captioning systems from literature.
result Language models alone are insufficient for capturing spatial location in image captions.
Improved image captions through adversarial semantic alignment.
problem Automatic evaluation and generalization to unseen compositions.
method Context-aware LSTM captioner and co-attentive discriminator for semantic alignment; SCST training method.
result SCST training method shows better performance in semantic score and human evaluation.
Model learns image-word associations from captions using contrastive learning.
problem Phrase grounding, associating image regions to caption words.
method Optimizing word-region attention to maximize mutual information, using language model guided word substitutions for negatives.
result Model achieves 76.7% accuracy on Flickr30K Entities benchmark, a 5.7% gain from weak supervision.
Transformer model estimates keywords for better audio captioning.
problem Indeterminacy in word selection for audio events/scenes.
method Transformer-based model with keyword estimation.
result Achieved state-of-the-art performance in AAC.
System tackles indeterminacies in automated audio captioning.
problem Word selection and sentence length indeterminacies in automated audio captioning.
method Solves caption generation and sub-indeterminacy problems through multi-task learning to estimate keywords and sentence length.
result Model achieved 20.7 SPIDEr score, significantly outperforming baseline.
Develops a variational autoencoder for image, label, and caption modeling.
problem Deep learning of images, labels, and captions.
method Uses a Deep Generative Deconvolutional Network (DGDN) and a Convolutional Neural Network (CNN) for image and latent feature encoding.
result Able to predict labels or captions for new images using latent code distributions.
The New Yorker publishes a weekly captionless cartoon. More than 5,000 readers submit captions for it. The editors select three of them and ask the readers to pick the funniest one. We describe an experiment that compares a dozen automatic methods for selecting the funniest caption. We show that negative sentiment, hum…
Improves text-to-image translation by using GANs and captioning networks.
problem Generating images that accurately reflect the meaning of a sentence.
method Uses cycle consistent adversarial networks and captioning networks to improve image generation.
result Significantly improved image quality compared to existing methods.
JECL clusters images and captions by jointly learning representations and assignments.
problem Clustering image-caption pairs with limited structured training data.
method Parallel encoders trained with clustering and alignment objectives, minimizing KL divergence and maximizing Jensen-Shannon divergence, with regularizers.
result JECL outperforms single-view and multi-view methods on large image-caption datasets.
Paper proposes method to generate images from text using GANs trained on uncaptioned images.
problem Limited captioned image datasets for text-to-image synthesis.
method Conditional GANs trained on uncaptioned images with an Image Captioning Module.
result Promising preliminary results compared to unconditional GANs.
Efficient video captioning model captures cross-modal interactions.
problem Capturing frame-level cross-modal interactions in video captioning.
method Proposes High-Order Cross-Modal Attention (HOCA) and Low-Rank HOCA.
result Low-Rank HOCA achieves state-of-the-art performance.
Study improves image-caption retrieval by quantifying feature and posterior uncertainty.
problem Improving reliability in image-caption retrieval tasks with deep learning models.
method Quantified feature and posterior uncertainty for model averaging and reliability measure in image-caption retrieval.
result Consistent improvement in retrieval performance with different datasets and architectures.
System automatically searches for scientific features in Mars rover images.
problem Manual task planning for Mars rover images is time-consuming and error-prone.
method Deep image captioning network and automated task prioritization.
result System can prioritize images for transmission based on search tasks.
Unified architecture for multi-modal multi-task learning using transformer.
problem Training multiple tasks concurrently with varying modalities.
method Spatio-temporal cache mechanism for multi-modal learning.
result Training multiple tasks together reduces model size by about three times.
Study finds CLIP's caption-based learning outperforms image-only methods under certain conditions.
problem Comparing CLIP's performance with traditional image-only methods in learning transferable representations.
method Controlled comparison of CLIP and image-only methods using a dataset with descriptive captions and specific criteria.
result CLIP's caption-based learning outperforms image-only methods when certain conditions are met, but can be detrimental in others.
Paper improves video feature learning for better downstream tasks.
problem Improving video feature learning for better performance on downstream tasks.
method Self-supervised learning approach using contrastive bidirectional transformer, extending BERT for real-valued feature vectors.
result Significantly improved performance on video classification, captioning, and segmentation tasks.
Improved satellite image captions enhance descriptiveness without large models.
problem Extracting meaningful text from satellite imagery.
method Evaluated seven models on a large benchmark, extended vocabulary, and introduced a novel confusion matrix.
result Reduced model size by 100x without sacrificing accuracy, offering new deployment opportunities.
RL for image captions improved with a language prior.
problem Learning biases and large sample space issues in RL image captioning.
method Added a language prior to constrain the action space.
result RL with the language prior module performs better in readability and speed.
Paper introduces Auto DeepVis to explain catastrophic forgetting in continual learning.
problem Catastrophic forgetting in continual learning of deep neural networks.
method Auto DeepVis and critical freezing techniques to address catastrophic forgetting.
result Critical freezing outperforms other methods on both past and future tasks.
Generates realistic faces from detailed textual descriptions.
problem Face generation from fine-grained textual descriptions.
method Conditional GAN model with DC-GAN and GAN-CLS loss, using CelebA dataset with generated captions.
result Promising results in generating diverse face images from fine-grained textual descriptions.
Generates detailed fashion feedback from outfit images.
problem Creating informative and diverse fashion feedback from outfit images.
method Trained deep generative models with visual attention, then improved with Maximum Mutual Information objective function.
result Generated sentences are more diverse and detailed.
Non-autoregressive model speeds up sequence generation tasks.
problem Efficiency in sequence generation tasks.
method Iterative refinement based on latent variable models and denoising autoencoders.
result Significant speedup in decoding with comparable quality.
Paper tackles hard attention training using variational inference.
problem Training hard attention models is difficult due to discrete latent variables.
method Uses variational inference methods (VIMCO, NVIL) and a novel adaptation.
result Method outperforms REINFORCE on phoneme recognition tasks.
LoRA-MCL improves language models by generating diverse sentence continuations.
problem Language models struggle with generating diverse, plausible sentence continuations.
method Low-Rank Adaptation combined with Multiple Choice Learning (MCL) to handle ambiguity.
result LoRA-MCL generates high-diversity and relevant outputs in various tasks.
Generative Score Inference improves uncertainty quantification for multimodal data.
problem Accurate uncertainty quantification in multimodal learning tasks.
method Generative Score Inference (GSI) uses synthetic samples to approximate conditional score distributions.
result GSI achieves state-of-the-art performance in hallucination detection and image captioning uncertainty estimation.
Bayesian attention modules improve model interpretability and performance.
problem Deterministic attention modules limit model interpretability and optimization.
method Proposes a scalable stochastic attention module using simplex-constrained distributions and Bayesian learning.
result Consistent improvements over baselines in various attention-based models.
Compact RNNs reduce parameters and improve efficiency.
problem High computational cost of RNNs with large inputs.
method Block-Term Tensor Decomposition (BT-TD) to reduce RNN parameters.
result BT-RNN achieves better accuracy and faster convergence than standard RNNs.
Unified approach for multimodal data prediction using synthetic data generation.
problem Challenges in integrating heterogeneous data types for accurate predictive performance.
method Generative Distribution Prediction (GDP) framework that uses multimodal synthetic data generation.
result Empirical validation across four tasks demonstrates versatility and effectiveness of GDP.
Fine-tunes language models with captions for differential equation solving.
problem Overreliance on function data in operator learning.
method Integrates human knowledge through captions and fine-tunes language models.
result Significantly enhanced performance and reduced function data requirements.
COBRA reduces modality gap in cross-modal tasks.
problem Joint embedding spaces fail to sufficiently reduce modality gap in multi-modal tasks.
method COBRA trains image and text modalities in a joint fashion using Contrastive Predictive Coding and Noise Contrastive Estimation.
result COBRA significantly reduces the modality gap and generates robust joint-embedding space.
Improves sequence generation by training a backward network.
problem Generating long-term dependencies in sequence models.
method Train a backward recurrent network to predict states of a forward model.
result Achieves 9% relative improvement in speech recognition and significant improvement in caption generation.