Graph-RISE learns image embeddings for ultra-fine-grained semantics.
problem Learning image representations for fine-grained semantics.
method Graph-regularized neural graph learning framework.
result Graph-RISE outperforms state-of-the-art on image classification and triplet ranking.
Paper introduces LLISE for image structure learning using SSIM.
problem Image quality assessment using MSE or ℓ2 norm is not promising. method Locally Linear Image Structural Embedding (LLISE) using SSIM.
result LLISE captures image structure features and discriminates distortions.
We introduce a new multi-dimensional nonlinear embedding -- Piecewise Flat Embedding (PFE) -- for image segmentation. Based on the theory of sparse signal recovery, piecewise flat embedding with diverse channels attempts to recover a piecewise constant image representation with sparse region boundaries and sparse clust…
The purpose of this article is to investigate the relationship between suborbifolds and orbifold embeddings. In particular, we give natural definitions of the notion of suborbifold and orbifold embedding and provide many examples. Surprisingly, we show that there are (topologically embedded) smooth suborbifolds which d…
Enhances image classification by integrating semantic hierarchy into CNN models.
problem Limited use of external guidance in image classification.
method Integrates label-hierarchy knowledge into CNN-based classifiers and uses order-preserving embeddings.
result Boosts image classification performance through semantic hierarchy integration.
Learning social media data embedding by deep models has attracted extensive research interest as well as boomed a lot of applications, such as link prediction, classification, and cross-modal search. However, for social images which contain both link information and multimodal contents (e.g., text description, and visu…
Paper introduces privacy-preserving few-shot learning for images.
problem Privacy risk in few-shot learning systems.
method Discrete embedding vectors and one-way hash functions.
result Achieves computational pan privacy without storing embeddings.
Paper uses DL and image embedding to classify power grid disturbances.
problem Classifying transient disturbances in power grids.
method Transformed time series data into images using Gramian Angular Field, then applied CNN and RNN for classification.
result DL algorithms outperform traditional data mining methods in power grid disturbance classification.
Profiling cellular phenotypes from microscopic imaging can provide meaningful biological information resulting from various factors affecting the cells. One motivating application is drug development: morphological cell features can be captured from images, from which similarities between different drug compounds appli…
We explore the practicability of Nash's Embedding Theorem in vision and imaging sciences. In particular, we investigate the relevance of a result of Burago and Zalgaller regarding the existence of isometric embeddings of polyhedral surfaces in R3 and we show that their proof does not extended directly to hi…
New framework infers multiple classes per image for one-shot learning.
problem Inferring multiple classes per image in one-shot learning.
method Compositional embedding framework with joint training of embedding and composition/query functions.
result Compositional embedding models outperform existing methods on various datasets.
ECS evaluates synthetic CXR images' distributional fidelity.
problem Evaluating synthetic CXR images' distributional fidelity under privacy constraints.
method Characteristic function transforms of feature embeddings.
result ECS uncovers clinically relevant distributional discrepancies.
End-to-end deep metric learning tackles multi-label image classification.
problem Multi-label image classification problem.
method Two-way deep distance metric learning in a latent space with a reconstruction module.
result Our method outperforms state-of-the-arts on publicly available image datasets.
The thesis introduces methods to use semantic hierarchy in image classification.
problem Limited work in training image classifiers with non-conventional external guidance.
method Injects label hierarchy knowledge into arbitrary classifiers and uses order-preserving embeddings for image classification.
result Both embedding-based models and CNN-classifiers with hierarchical information outperform a hierarchy-agnostic model.
Paper improves image retrieval quality using nonlinear rank approximations.
problem Improving image retrieval quality in high-dimensional feature spaces.
method Computes normalized approximated ranks, converts to similarities, and uses them in a new loss function.
result Significant improvement in image retrieval quality on multiple datasets.
ADD embeds a 48-bit message into images, achieving high accuracy and speed.
problem Embedding high-fidelity messages into images to detect authenticity and source.
method Two-stage process: linear combination and addition of watermark to image, followed by decoding.
result ADD achieves 100% decoding accuracy for 48-bit watermarking, with minimal performance drop under various distortions.
Improved object detection for scientific document images.
problem Current object detectors fail to accurately localize regions in scientific document images.
method Revised R-CNN model with region embedding for fine-grained proposals.
result 17% mAP improvement over standard object detection models.
The paper analyzes and proposes an algorithm for multi-modal nonlinear embeddings with theoretical performance bounds.
problem Generalizability of multi-modal nonlinear embeddings to unseen data.
method Theoretical analysis and a multi-modal nonlinear representation learning algorithm motivated by performance bounds.
result The proposed algorithm yields promising performance in multi-modal image classification and cross-modal image-text retrieval applications.
The paper explores winding numbers of almost embeddings of a 4-vertex graph in the plane.
problem Understanding the winding numbers of almost embeddings of a 4-vertex graph in the plane.
method Constructing examples to show the only relation between the winding numbers of cycles in the graph.
result The sum of winding numbers is odd, and this is the only relation between them.
SketchEmbedNet learns image representations from sketches, useful for few-shot learning.
problem Learning image representations from sketches for few-shot learning.
method Training a model to produce sketches of images, focusing on informative embeddings.
result Model produces informative embeddings of novel images, classes, and datasets.
Method retrieves similar fashion items from images and text, enabling style refinement.
problem Lack of intuitive, interactive refinement in search engines for fashion items.
method Joint visual-textual embedding training, Mini-Batch Match Retrieval, attribute extraction.
result Improved performance in multimodal style search, demonstrated through benchmark.
NASES uses embedding space for efficient NAS in image classification tasks.
problem Difficulty in optimizing high-dimensional discrete architecture spaces.
method NASES employs architecture encoders and decoders to search in an embedding space using reinforcement learning.
result NASES discovers comparable final architectures to other NAS approaches in less time.
Deep-learning method estimates bone 3D structure from X-ray images.
problem Estimating bone 3D structure from X-ray images.
method Triplet loss-trained neural network selecting closest 3D bone shape from predefined set.
result Average RMS distance of 1.08 mm between predicted and true shapes.
DGE learns event representations from image sequences without manual annotations.
problem Data hunger and domain adaptation issues in self-supervised learning for temporal segmentation.
method Dynamic Graph Embedding (DGE) learns event representations by iteratively updating a graph and its embedding.
result DGE achieves robust temporal segmentation on benchmark datasets, outperforming state-of-the-art methods.
We construct a large class of pathological n-dimensional topological spheres in Rn+1 by showing that for any Cantor set C⊂Rn+1 there is a topological embedding f:Sn→Rn+1 of the Sobolev class W1,n whose image contains the Cantor set C.
We construct examples of complex algebraic surfaces not admitting normal embeddings (in the sense of semialgebraic or subanalytic sets) with image a complex algebraic surface.
JECL clusters images and captions by jointly learning representations and assignments.
problem Clustering image-caption pairs with limited structured training data.
method Parallel encoders trained with clustering and alignment objectives, minimizing KL divergence and maximizing Jensen-Shannon divergence, with regularizers.
result JECL outperforms single-view and multi-view methods on large image-caption datasets.
Low dimensional embeddings that capture the main variations of interest in collections of data are important for many applications. One way to construct these embeddings is to acquire estimates of similarity from the crowd. However, similarity is a multi-dimensional concept that varies from individual to individual. Ex…
Generalizes embedding complex Grassmannians into quadrics.
problem Holomorphic isometric embeddings of complex Grassmannians into quadrics.
method Generalization of do Carmo-Wallach theory for moduli spaces.
result Moduli spaces of embeddings discussed.
Study holomorphic isometric embeddings of a Grassmannian into quadrics.
problem Holomorphic isometric embeddings of complex Grassmannian into quadrics.
method Generalization of do Carmo--Wallach theory to study moduli space.
result Moduli space of embeddings up to equivalence discussed.
Selfie pretrains image embeddings without pixel-level predictions.
problem Training image embeddings efficiently with limited labeled data.
method Contrastive Predictive Coding loss for masked patches, using convolutional and attention mechanisms.
result Significant improvement in ResNet-50 accuracy on ImageNet 224 x 224 with 60 examples per class (5%).
In this article we study the Hamiltonian non-displaceability of Gauss images of isoparametric hypersurfaces in the spheres as Lagrangian submanifolds embedded in complex hyperquadrics.
Wassmap reduces image complexity while preserving key features.
problem Global nonlinear dimensionality reduction in imaging.
method Wassmap uses Wasserstein space and pairwise distances to create isometric embeddings.
result Wassmap can recover parameters of image manifolds like translations and dilations.
Study improves image-caption retrieval by quantifying feature and posterior uncertainty.
problem Improving reliability in image-caption retrieval tasks with deep learning models.
method Quantified feature and posterior uncertainty for model averaging and reliability measure in image-caption retrieval.
result Consistent improvement in retrieval performance with different datasets and architectures.
Super-AND improves unsupervised embedding learning with 89.2% accuracy on CIFAR-10.
problem Extracting good representations from data without labels.
method Super-AND uses unique losses to gather similar samples and maintain features.
result Super-AND achieves 89.2% accuracy on CIFAR-10 image classification.
The ability to characterize the color content of natural imagery is an important application of image processing. The pixel by pixel coloring of images may be viewed naturally as points in color space, and the inherent structure and distribution of these points affords a quantization, through clustering, of the color i…
For any n-dimensional compact Riemannian manifold (M,g), we construct a canonical t-family of isometric embeddings I_{t}: M->R^{q(t)}, with t>0 sufficiently small and q(t)>>t^{-n/2}. This is done by intrinsically perturbing the heat kernel embedding introduced in [BBG]. As t->0, asymptotic geometry of the embedded imag…
A fast binary embedding method preserves Euclidean distances in high-dimensional data.
problem Preserving Euclidean distances in high-dimensional datasets.
method Stable noise-shaping quantization of Ax with A a sparse Gaussian random matrix, followed by a linear transformation. result Euclidean distances are approximated by the ℓ1 norm on binary sequences, leading to accurate binary codes. HCRL learns hierarchical embeddings from deep embeddings of hierarchy components.
problem Flat clustering limits cohesive instance relations in hierarchical data.
method Simultaneously optimizes representation learning and hierarchical clustering in the embedding space.
result HCRL achieves best hierarchical clustering and data reconstruction.
Instance embeddings are an efficient and versatile image representation that facilitates applications like recognition, verification, retrieval, and clustering. Many metric learning methods represent the input as a single point in the embedding space. Often the distance between points is used as a proxy for match confi…
Injecting adversarial examples during training, known as adversarial training, can improve robustness against one-step attacks, but not for unknown iterative attacks. To address this challenge, we first show iteratively generated adversarial images easily transfer between networks trained with the same strategy. Inspir…
We present a method for training multi-label, massively multi-class image classification models, that is faster and more accurate than supervision via a sigmoid cross-entropy loss (logistic regression). Our method consists in embedding high-dimensional sparse labels onto a lower-dimensional dense sphere of unit-normed …
Proposes Group Loss for deep metric learning to improve clustering and image retrieval.
problem Improving deep metric learning for better clustering and image retrieval.
method Group Loss based on label-propagation method enforcing embedding similarity across all samples of a group.
result Shows state-of-the-art results on clustering and image retrieval on several datasets.
In this project we analysed how much semantic information images carry, and how much value image data can add to sentiment analysis of the text associated with the images. To better understand the contribution from images, we compared models which only made use of image data, models which only made use of text data, an…
Automatically evaluates image quality based on human judgment.
problem Difficulty in rigorously evaluating generated image quality.
method Generative model embeddings, human labels regression, and statistical matching.
result 66% accuracy in predicting human scores of image realism.
Study improves Poincaré-Sobolev inequalities for differential forms.
problem Improving Sobolev space embeddings for differential forms.
method Utilizes Lq,p-cohomology and bi-Lipschitz images to estimate embedding norms. result Estimates for embedding norms in Euclidean balls and their images.
Generative Adversarial Nets (GANs) and Variational Auto-Encoders (VAEs) provide impressive image generations from Gaussian white noise, but the underlying mathematics are not well understood. We compute deep convolutional network generators by inverting a fixed embedding operator. Therefore, they do not require to be o…
The paper introduces an inequality to detect when surfaces in 4-manifolds are not smoothly isotopic.
problem Detecting when surfaces in 4-manifolds are not smoothly isotopic.
method An adjunction inequality to distinguish non-isotopic surfaces.
result The minimal genus of a surface isotopic to its image under a diffeomorphism is generally larger.