Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

69138206275 · Jun 202019922001200920172026
48 results for image description

Most existing text-to-image synthesis tasks are static single-turn generation, based on pre-defined textual descriptions of images. To explore more practical and interactive real-life applications, we introduce a new task - Interactive Image Editing, where users can guide an agent to edit images via multi-turn textual …

2018-12-20abs ↗pdf ↗

Generates garden paintings from text descriptions using deep learning.

problem Lack of firsthand material for traditional Chinese garden reconstruction.
method Deep learning model trained on text and paintings of Ming Dynasty gardens.
result Model generates garden paintings in Ming Dynasty style based on textual descriptions.

Improved satellite image captions enhance descriptiveness without large models.

problem Extracting meaningful text from satellite imagery.
method Evaluated seven models on a large benchmark, extended vocabulary, and introduced a novel confusion matrix.
result Reduced model size by 100x without sacrificing accuracy, offering new deployment opportunities.

Paper evaluates using app images for classification, improving accuracy.

problem Improving app classification accuracy when text descriptions are missing or inadequate.
method Used OCR, pic2vec, captionbot.ai, and object detection to convert images into text or vectors for classification.
result Improved classification accuracy of 96% for some app categories when images are added.

Paper proposes a deep learning architecture for generating long stories from images.

problem Maintaining context in long event sequences for visual storytelling.
method Hierarchical deep learning architecture with encoder-decoder networks and natural language descriptions.
result Our method outperforms state-of-the-art techniques on automatic evaluation metrics.

We introduce a new dataset of 293,008 high definition (1360 x 1360 pixels) fashion images paired with item descriptions provided by professional stylists. Each item is photographed from a variety of angles. We provide baseline results on 1) high-resolution image generation, and 2) image generation conditioned on the gi…

2018-06-21abs ↗pdf ↗

The recent success of Generative Adversarial Networks (GAN) is a result of their ability to generate high quality images from a latent vector space. An important application is the generation of images from a text description, where the text description is encoded and further used in the conditioning of the generated i…

2019-05-16abs ↗pdf ↗

Recent progress in using recurrent neural networks (RNNs) for image description has motivated the exploration of their application for video description. However, while images are static, working with videos requires modeling their dynamic temporal structure and then properly integrating that information into a natural…

2015-02-27abs ↗pdf ↗

Study improves neural network performance in sequential learning for image classification.

problem Improving neural network performance in sequential learning for image classification.
method Evaluation of approaches for computing prequential description lengths, proposing forward-calibration and replay-streams.
result Improved description lengths for image classification datasets, outperforming previous results.

Tensor networks reveal limitations for efficient text description but suggest potential for images.

problem Efficiently describing large text and image data sets using tensor networks.
method Investigation of mutual information scaling, introduction of mutual information estimators, and use of autoregressive and convolutional neural networks.
result Text data cannot be efficiently described by 1D tensor networks, while images may be better described by 2D tensor networks.

Generating a description of an image is called image captioning. Image captioning requires to recognize the important objects, their attributes and their relationships in an image. It also needs to generate syntactically and semantically correct sentences. Deep learning-based techniques are capable of handling the comp…

2018-10-06abs ↗pdf ↗

FCDD improves image anomaly detection without post-hoc explainers.

problem Image anomaly detection, especially pixel-wise.
method Fully Convolutional Data Description (FCDD) directly addresses anomaly detection without post-hoc methods.
result FCDD achieves state-of-the-art results on pixel-wise AD tasks.

An analytic approach and description are presented for the moduli cotangent sheaf for suitable stable curve families including noded fibers. For sections of the square of the relative dualizing sheaf, the residue map at a node gives rise to an exact sequence. The residue kernel defines the vanishing residue subsheaf. F…

2012-04-17abs ↗pdf ↗

Study subgroup of mapping class group related to handlebody, answering a question about Johnson homomorphism.

problem Understanding the image of the second Johnson homomorphism for a specific subgroup of mapping class groups.
method Introduced trace-like operators and used them to compute images of Johnson homomorphisms.
result Answered a question about algebraic description of the image of the second Johnson homomorphism.

Generating an image from its description is a challenging task worth solving because of its numerous practical applications ranging from image editing to virtual reality. All existing methods use one single caption to generate a plausible image. A single caption by itself, can be limited, and may not be able to capture…

2018-09-20abs ↗pdf ↗

The contraction of the image of the Johnson homomorphism is called the Chillingworth class. In this paper, we derive a combinatorial description of the Chillingworth class for Putman's subsurface Torelli groups. We also prove the naturality and uniqueness properties of the map whose image is the dual of the Chillingwor…

2019-03-09abs ↗pdf ↗

We propose JECL, a method for clustering image-caption pairs by training parallel encoders with regularized clustering and alignment objectives, simultaneously learning both representations and cluster assignments. These image-caption pairs arise frequently in high-value applications where structured training data is e…

2019-01-04abs ↗pdf ↗

Starting from the description of Segre forms as direct images of (powers of) the first Chern form of the (anti)tautological line bundle on the projectivized bundle of a holomorphic hermitian vector bundle, we derive a version of the pointwise Kobayashi-Lübke inequality.

2015-03-09abs ↗pdf ↗

In this study, we investigated multi-modal approaches using images, descriptions, and titles to categorize e-commerce products on Amazon. Specifically, we examined late fusion models, where the modalities are fused at the decision level. Products were each assigned multiple labels, and the hierarchy in the labels were …

2019-06-30abs ↗pdf ↗

Johnson has defined a surjective homomorphism from the Torelli subgroup of the mapping class group of the surface of genus gg with one boundary component to 3H\wedge^3 H, the third exterior product of the homology of the surface. Morita then extended Johnson's homomorphism to a homomorphism from the entire mapping cla…

2007-08-28abs ↗pdf ↗

Self-guidance controls image generation by extracting properties from diffusion model representations.

problem Generating images from text descriptions is challenging due to the complexity of visual details.
method Self-guidance uses internal representations of diffusion models to control image generation.
result Properties like object shape, location, and appearance can be extracted and used to steer image generation.

We discuss holomorphic isometric embeddings of the projective line into quadrics using a generalisation of the theorem of do Carmo--Wallach to provide a description of their moduli spaces up to image and gauge--equivalence. Moreover, we show rigidity of the real standard map from the projective line into quadrics.

2014-08-14abs ↗pdf ↗

In this work we propose a new computational framework, based on generative deep models, for synthesis of photo-realistic food meal images from textual descriptions of its ingredients. Previous works on synthesis of images from text typically rely on pre-trained text models to extract text features, followed by a genera…

2019-05-09abs ↗pdf ↗

The Burau representation is a fundamental bridge between the braid group and diverse other topics in mathematics. A 1974 question of Birman asks for a description of the image; in this paper we give a "strong approximation" to the answer. Since a 1984 paper of Squier it has been known that the Burau representation pres…

2019-03-27abs ↗pdf ↗

CDL index improves clustering validation for non-convex data.

problem Selecting clustering algorithms and hyperparameters without labeled data.
method CDL uses compactness, centers, and covariances to compute a probabilistic description length bound.
result CDL outperforms conventional CVIs on synthetic and image benchmarks.

New method evaluates text-to-image synthesis for realism, variety, and semantic accuracy.

problem Lack of metrics revealing semantic accuracy in text-to-image synthesis.
method Uses Inception network representations and t-SNE visualization for semantic evaluation.
result Classification accuracy of generated images to real images' visual concepts correlates with semantic accuracy.

Zero-shot learning (ZSL) is a framework to classify images belonging to unseen classes based on solely semantic information about these unseen classes. In this paper, we propose a new ZSL algorithm using coupled dictionary learning. The core idea is that the visual features and the semantic attributes of an image can s…

2019-06-10abs ↗pdf ↗

Study finds CLIP's caption-based learning outperforms image-only methods under certain conditions.

problem Comparing CLIP's performance with traditional image-only methods in learning transferable representations.
method Controlled comparison of CLIP and image-only methods using a dataset with descriptive captions and specific criteria.
result CLIP's caption-based learning outperforms image-only methods when certain conditions are met, but can be detrimental in others.

Many problems in image processing and computer vision (e.g. colorization, style transfer) can be posed as 'manipulating' an input image into a corresponding output image given a user-specified guiding signal. A holy-grail solution towards generic image manipulation should be able to efficiently alter an input image wit…

2017-03-21abs ↗pdf ↗

In this paper we explain how non-abelian Hodge theory allows one to compute the L2L^2 cohomology or middle perversity higher direct images of harmonic bundles and twistor D-modules in a purely algebraic manner. Our main result is a new algebraic description for the fiberwise L2L^2 cohomology of a tame harmonic bundle o…

2016-12-19abs ↗pdf ↗

We describe the νν-lines of curvature of an embedding of the double torus into R4\mathbb R^4, defined as the link of the real part of the Milnor fibration of a polynomial, where νν is its gradient. Through this analysis, we present a complete description of the foliation of lines of curvature of the embedding, define…

2019-02-05abs ↗pdf ↗