Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

2579 · Sep 201819922001200920182026
48 results for Captions

New model improves image captioning's ability to describe unseen concepts.

problem Image captioning models struggle with describing unseen combinations of concepts.
method Proposes a multi-task model combining caption generation and image-sentence ranking, with a decoding mechanism to re-rank captions based on image similarity.
result The model significantly outperforms state-of-the-art models in compositional generalization.

Generative model uses captions to generate images, improving semantic understanding.

problem Complex image generation models require large datasets and intricate learning.
method Adapts captioning models to generate images, using learned sentence and frame vectors.
result Images generated from multiple captions better capture semantic meaning.

C4Synth generates images from multiple captions to improve image quality.

problem Generating images from a single caption is limited; multiple captions are needed.
method Two deep generative models that ensure 'Cross-Caption Cycle Consistency'.
result Quantitative and qualitative validation on Caltech-UCSD Birds and Oxford-102 Flowers datasets.

Model learns image-word associations from captions using contrastive learning.

problem Phrase grounding, associating image regions to caption words.
method Optimizing word-region attention to maximize mutual information, using language model guided word substitutions for negatives.
result Model achieves 76.7% accuracy on Flickr30K Entities benchmark, a 5.7% gain from weak supervision.

Experiment shows that certain factors predict which New Yorker cartoon captions are funniest.

problem Determining which captions are funniest in the New Yorker cartoon contest.
method Comparison of 12 automatic methods for identifying funniest captions.
result Negative sentiment, human-centeredness, and lexical centrality are key predictors of funniest captions.

System tackles indeterminacies in automated audio captioning.

problem Word selection and sentence length indeterminacies in automated audio captioning.
method Solves caption generation and sub-indeterminacy problems through multi-task learning to estimate keywords and sentence length.
result Model achieved 20.7 SPIDEr score, significantly outperforming baseline.

Develops a variational autoencoder for image, label, and caption modeling.

problem Deep learning of images, labels, and captions.
method Uses a Deep Generative Deconvolutional Network (DGDN) and a Convolutional Neural Network (CNN) for image and latent feature encoding.
result Able to predict labels or captions for new images using latent code distributions.

Improves text-to-image translation by using GANs and captioning networks.

problem Generating images that accurately reflect the meaning of a sentence.
method Uses cycle consistent adversarial networks and captioning networks to improve image generation.
result Significantly improved image quality compared to existing methods.

JECL clusters images and captions by jointly learning representations and assignments.

problem Clustering image-caption pairs with limited structured training data.
method Parallel encoders trained with clustering and alignment objectives, minimizing KL divergence and maximizing Jensen-Shannon divergence, with regularizers.
result JECL outperforms single-view and multi-view methods on large image-caption datasets.

Paper develops a framework for generating coherent image captions using visual features and hierarchical topics.

problem Generating semantically coherent paragraphs to describe image content.
method Plug-and-play hierarchical-topic-guided image paragraph generation framework integrating visual extractor and deep topic model.
result Proposed models can distill interpretable multi-layer semantic topics and generate diverse and coherent captions.

Study improves image-caption retrieval by quantifying feature and posterior uncertainty.

problem Improving reliability in image-caption retrieval tasks with deep learning models.
method Quantified feature and posterior uncertainty for model averaging and reliability measure in image-caption retrieval.
result Consistent improvement in retrieval performance with different datasets and architectures.

Study finds CLIP's caption-based learning outperforms image-only methods under certain conditions.

problem Comparing CLIP's performance with traditional image-only methods in learning transferable representations.
method Controlled comparison of CLIP and image-only methods using a dataset with descriptive captions and specific criteria.
result CLIP's caption-based learning outperforms image-only methods when certain conditions are met, but can be detrimental in others.

Improved satellite image captions enhance descriptiveness without large models.

problem Extracting meaningful text from satellite imagery.
method Evaluated seven models on a large benchmark, extended vocabulary, and introduced a novel confusion matrix.
result Reduced model size by 100x without sacrificing accuracy, offering new deployment opportunities.

Enhances image captioning with novel context combination methods.

problem Improving machine learning for image captioning with structured learning and meaningful interpretation.
method Combines Feature Distribution Composition (FDC), Multiple Role Representation Crossover (MRRC) attention layers, and language decoder.
result Significantly improved image captioning performance (35.3%) and established new standards.

This paper improves image captioning by aligning visual attention with sentence generation.

problem Generating accurate and meaningful image captions from visual scenes.
method Region-based attention and scene factorization to align visual perception with sentence generation.
result Combining region-based attention and scene-specific contexts improves image captioning performance.

Improved video and movie description using multitask learning.

problem Lack of training data and poor generalization in video captioning.
method Multitask learning encoder-decoder framework for video sequences.
result Improved performance on multi-caption and single-caption datasets.

Generative Score Inference improves uncertainty quantification for multimodal data.

problem Accurate uncertainty quantification in multimodal learning tasks.
method Generative Score Inference (GSI) uses synthetic samples to approximate conditional score distributions.
result GSI achieves state-of-the-art performance in hallucination detection and image captioning uncertainty estimation.

Paper tackles model failures producing outputs of undesirable length.

problem Model failures producing outputs of undesirable length.
method Develops a differentiable proxy objective and a verification approach.
result Can produce outputs 50 times longer than input and prove output length below a certain size.

CMT efficiently manages memory by inserting and querying memories in logarithmic time.

problem Managing large memory stores efficiently for quick access and updates.
method Designing a Contextual Memory Tree (CMT) that inserts and retrieves memories in logarithmic time.
result CMT improves classification algorithms and image-captioning tasks, demonstrating better computational efficiency.

Seq-CVAE learns a latent space for each word position to capture sentence intention.

problem Capturing diversity in image captioning models.
method Seq-CVAE learns a sequential latent space for each word position, mimicking future sentence summaries.
result Significantly improves diversity metrics on MSCOCO dataset compared to baselines.

Unified framework for generating meteorological time series from text.

problem Lack of large-scale, physically grounded multimodal datasets and architectures ignoring spectral-temporal structure.
method Introduce MeteoCap-3B dataset and MTransformer model.
result State-of-the-art generation quality, accurate cross-modal alignment, strong semantic controllability.

The paper proposes a method to learn from neighbor annotations to improve predictions in ambiguous tasks.

problem Ambiguity in structured prediction problems with large output spaces makes exhaustive annotation impractical.
method Proposes an objective to transfer supervision from neighboring examples to improve predictions.
result Consistent improvements over standard maximum likelihood training and baselines in multi-label classification and image-grounded sequence modeling tasks.

A new technique called fraternal dropout improves RNN performance.

problem Optimizing recurrent neural networks (RNNs) is harder than feed-forward networks.
method Train two identical RNNs with different dropout masks to encourage robust representations.
result Achieves state-of-the-art results on sequence modeling tasks and improves image captioning and semi-supervised learning.