Most existing text-to-image synthesis tasks are static single-turn generation, based on pre-defined textual descriptions of images. To explore more practical and interactive real-life applications, we introduce a new task - Interactive Image Editing, where users can guide an agent to edit images via multi-turn textual …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Synthesizing high-quality images from text descriptions is a challenging problem in computer vision and has many practical applications. Samples generated by existing text-to-image approaches can roughly reflect the meaning of the given descriptions, but they fail to contain necessary details and vivid object parts. In…
Generates garden paintings from text descriptions using deep learning.
Improved satellite image captions enhance descriptiveness without large models.
Paper evaluates using app images for classification, improving accuracy.
Paper proposes a deep learning architecture for generating long stories from images.
We introduce a new dataset of 293,008 high definition (1360 x 1360 pixels) fashion images paired with item descriptions provided by professional stylists. Each item is photographed from a variety of angles. We provide baseline results on 1) high-resolution image generation, and 2) image generation conditioned on the gi…
Generates realistic faces from detailed textual descriptions.
The recent success of Generative Adversarial Networks (GAN) is a result of their ability to generate high quality images from a latent vector space. An important application is the generation of images from a text description, where the text description is encoded and further used in the conditioning of the generated i…
Recent progress in using recurrent neural networks (RNNs) for image description has motivated the exploration of their application for video description. However, while images are static, working with videos requires modeling their dynamic temporal structure and then properly integrating that information into a natural…
Study improves neural network performance in sequential learning for image classification.
Synthesizing images or texts automatically is a useful research area in the artificial intelligence nowadays. Generative adversarial networks (GANs), which are proposed by Goodfellow in 2014, make this task to be done more efficiently by using deep neural networks. We consider generating corresponding images from an in…
Although Generative Adversarial Networks (GANs) have shown remarkable success in various tasks, they still face challenges in generating high quality images. In this paper, we propose Stacked Generative Adversarial Networks (StackGAN) aiming at generating high-resolution photo-realistic images. First, we propose a two-…
Tensor networks reveal limitations for efficient text description but suggest potential for images.
Generating a description of an image is called image captioning. Image captioning requires to recognize the important objects, their attributes and their relationships in an image. It also needs to generate syntactically and semantically correct sentences. Deep learning-based techniques are capable of handling the comp…
FCDD improves image anomaly detection without post-hoc explainers.
An analytic approach and description are presented for the moduli cotangent sheaf for suitable stable curve families including noded fibers. For sections of the square of the relative dualizing sheaf, the residue map at a node gives rise to an exact sequence. The residue kernel defines the vanishing residue subsheaf. F…
Gradient flow on diffeomorphisms for image registration, with well-posedness proven.
Study subgroup of mapping class group related to handlebody, answering a question about Johnson homomorphism.
Study geodesics on finite-dimensional manifolds.
Generating an image from its description is a challenging task worth solving because of its numerous practical applications ranging from image editing to virtual reality. All existing methods use one single caption to generate a plausible image. A single caption by itself, can be limited, and may not be able to capture…
The contraction of the image of the Johnson homomorphism is called the Chillingworth class. In this paper, we derive a combinatorial description of the Chillingworth class for Putman's subsurface Torelli groups. We also prove the naturality and uniqueness properties of the map whose image is the dual of the Chillingwor…
We propose JECL, a method for clustering image-caption pairs by training parallel encoders with regularized clustering and alignment objectives, simultaneously learning both representations and cluster assignments. These image-caption pairs arise frequently in high-value applications where structured training data is e…
Starting from the description of Segre forms as direct images of (powers of) the first Chern form of the (anti)tautological line bundle on the projectivized bundle of a holomorphic hermitian vector bundle, we derive a version of the pointwise Kobayashi-Lübke inequality.
In this study, we investigated multi-modal approaches using images, descriptions, and titles to categorize e-commerce products on Amazon. Specifically, we examined late fusion models, where the modalities are fused at the decision level. Products were each assigned multiple labels, and the hierarchy in the labels were …
Johnson has defined a surjective homomorphism from the Torelli subgroup of the mapping class group of the surface of genus with one boundary component to , the third exterior product of the homology of the surface. Morita then extended Johnson's homomorphism to a homomorphism from the entire mapping cla…
Paper proposes method to generate images from text using GANs trained on uncaptioned images.
RNNs are crucial for text and speech tasks, explained in this overview.
We describe the range of a restricted spherical mean transform, which sends a function supported inside a closed ball in a hyperbolic space to its mean values on the geodesics spheres centered at the boundary of the ball. The description resembles that of the same transform on the Euclidean spaces obtained by Mark Agra…
Researchers describe the Gromov boundary of a graph related to surfaces.
This paper creates a tagging system for paintings using historical descriptions.
Self-guidance controls image generation by extracting properties from diffusion model representations.
We discuss holomorphic isometric embeddings of the projective line into quadrics using a generalisation of the theorem of do Carmo--Wallach to provide a description of their moduli spaces up to image and gauge--equivalence. Moreover, we show rigidity of the real standard map from the projective line into quadrics.
In this work we propose a new computational framework, based on generative deep models, for synthesis of photo-realistic food meal images from textual descriptions of its ingredients. Previous works on synthesis of images from text typically rely on pre-trained text models to extract text features, followed by a genera…
The Burau representation is a fundamental bridge between the braid group and diverse other topics in mathematics. A 1974 question of Birman asks for a description of the image; in this paper we give a "strong approximation" to the answer. Since a 1984 paper of Squier it has been known that the Burau representation pres…
Single-stage neural architecture improves text-to-image synthesis.
CDL index improves clustering validation for non-convex data.
New method evaluates text-to-image synthesis for realism, variety, and semantic accuracy.
Zero-shot learning (ZSL) is a framework to classify images belonging to unseen classes based on solely semantic information about these unseen classes. In this paper, we propose a new ZSL algorithm using coupled dictionary learning. The core idea is that the visual features and the semantic attributes of an image can s…
Study finds CLIP's caption-based learning outperforms image-only methods under certain conditions.
New method learns convolution-like structures from scratch.
Recent progress on automatic generation of image captions has shown that it is possible to describe the most salient information conveyed by images with accurate and meaningful sentences. In this paper, we propose an image caption system that exploits the parallel structures between images and sentences. In our model, …
Data-driven methods such as convolutional neural networks (CNNs) are known to deliver state-of-the-art performance on image recognition tasks when the training data are abundant. However, in some instances, such as change detection in remote sensing images, annotated data cannot be obtained in sufficient quantities. In…
Many problems in image processing and computer vision (e.g. colorization, style transfer) can be posed as 'manipulating' an input image into a corresponding output image given a user-specified guiding signal. A holy-grail solution towards generic image manipulation should be able to efficiently alter an input image wit…
In this paper we explain how non-abelian Hodge theory allows one to compute the cohomology or middle perversity higher direct images of harmonic bundles and twistor D-modules in a purely algebraic manner. Our main result is a new algebraic description for the fiberwise cohomology of a tame harmonic bundle o…
We continue to study twistor spaces on the connected sum of four complex projective planes, whose anticanonical map is of degree two over the image. In particular, we determine the defining equation of the branch divisor of the anticanonical map in an explicit form. Together with previous two articles (arXiv:1009.3153 …
We describe the -lines of curvature of an embedding of the double torus into , defined as the link of the real part of the Milnor fibration of a polynomial, where is its gradient. Through this analysis, we present a complete description of the foliation of lines of curvature of the embedding, define…
Learning social media data embedding by deep models has attracted extensive research interest as well as boomed a lot of applications, such as link prediction, classification, and cross-modal search. However, for social images which contain both link information and multimodal contents (e.g., text description, and visu…