Method recombines image content and style from different images.
problem Recombining image content and style from different images.
method Constructs content embedding, uses VAE with leakage filtering to ensure separation of style and content.
result Synthesizes novel images with state-of-the-art performance on few-shot learning tasks.
Content based image retrieval, a technique which uses visual contents of image to search images from large scale image databases according to users' interests. This paper provides a comprehensive survey on recent technology used in the area of content based face image retrieval. Nowadays digital devices and photo shari…
Efficiently transfers style to content without distorting the content structure.
problem Arbitrary style transfer in computer vision.
method Rigid alignment of style features to content features.
result High-quality stylized images with intact content structure.
DEMUD-VIS detects novel image content and explains it visually.
problem Detecting and explaining novel image content in large datasets.
method Uses CNN for feature extraction, reconstruction error for novelty detection, and up-convolutional networks for image reconstruction.
result Demonstrates visual explanations of novel image content on diverse datasets.
DCMIX learns channel importance for high content imaging.
problem Lack of channel importance information in deep learning-based image analysis.
method Image blending concepts with alpha compositing for arbitrary channels.
result DCMIX learns biologically relevant channel importance without sacrificing prediction performance.
A new method transcribes complex structured images like musical scores.
problem Transcribing content from images with complex internal structure.
method Hierarchical Spotlight Transcribing Network (STN) framework with two-stage approach.
result Demonstrated effectiveness through experiments on various structural image datasets.
A new method for image translation using disentangled style and content preservation.
problem Difficulty in maintaining original content during reverse diffusion in diffusion-based image translation.
method Disentangled style and content representation using intermediate keys from ViT model, CLIP loss, semantic divergence loss, and resampling strategy.
result Outperforms state-of-the-art models in text-guided and image-guided translation tasks.
DMAN embeds social images with deep multimodal attention networks.
problem Sub-optimal social image representation for social media data.
method Deep Multimodal Attention Networks (DMAN) that jointly embed multimodal contents and link information.
result DMAN achieves significant improvement in multi-label classification and cross-modal search compared to state-of-the-art image embeddings.
A mixture of CNNs improves adult content recognition.
problem Recognizing and restricting inappropriate images like pornography.
method A weighted sum of multiple CNN models trained using OLS.
result The proposed model outperforms single and average models.
Zero-shot contrastive loss improves text-guided image style transfer without extra training.
problem Stochastic nature of diffusion models leads to trade-offs between style transformation and content preservation.
method Proposes a zero-shot contrastive loss for diffusion models that doesn't require additional fine-tuning or auxiliary networks.
result Method outperforms existing methods while preserving content and requiring no additional training.
A new method for image translation without paired data.
problem Image-to-image translation between two domains with content preservation.
method Energy-based model in latent space of pretrained autoencoder.
result Improved translation quality and content preservation.
MIXGAN combines concepts from different domains for new image generation.
problem Generating new images with mixed content and style from different domains.
method MIXGAN is a mixture generative adversarial network that learns content and style from two domains and generates new images combining them.
result MIXGAN effectively generates new images with mixed content and style from different domains.
Proposes a VAE variant for ordinal content factors.
problem Isolating ordinal-valued content factors in deep latent variable models.
method Introduces a partially ordered set (poset) structure and a conditional Gaussian spacing prior model.
result Significant improvements in content-style separation over previous non-ordinal approaches.
Kernel Mean Matching enhances GANs for content-addressable generation.
problem Creating models that can generate images consistent with specified examples.
method Kernel Mean Matching applied to GANs.
result The method generates images consistent with specified input sets while maintaining original model quality.
Improved image translation using asymmetric gradient guidance.
problem Trade-off between style transformation and content preservation in diffusion models.
method Asymmetric gradient guidance to guide reverse diffusion sampling.
result Our method outperforms state-of-the-art models in image translation tasks.
New method uses SVD entropy to price artworks.
problem Lack of fine measurements in traditional art pricing models.
method SVD entropy of painting images for content measurement.
result SVD entropy positively affects sales price at 1% significance level.
RB-Modulation trains free diffusion models without external adapters.
problem Training-free personalization of diffusion models with style and content control.
method Stochastic optimal control with a style descriptor and cross-attention aggregation.
result Precise content and style extraction and control without external adapters.
System optimizes product images for e-commerce, enhancing customer engagement.
problem Optimizing product images for e-commerce to improve customer engagement.
method Machine learning, deep learning, and computer vision techniques applied to large e-commerce catalogs.
result System produces superior image sets tailored to customer preferences.
Proposes a general deep neural network method for digital watermarking.
problem Protecting intellectual content in a massive, IoT-acquired image dataset.
method Train a neural network on an image set and use it to protect distinct test images in bulk.
result Demonstrates the robustness and practicality of the proposed method.
Framework translates images between domains without supervision.
problem Challenges in unsupervised image-to-image translation, especially handling multimodality.
method Proposes a Multimodal Unsupervised Image-to-Image Translation (MUNIT) framework, decomposing images into content and style codes.
result Demonstrates improved generation of diverse outputs from a single source image.
Single auto-encoder learns cross-domain image translation.
problem Cross-domain image-to-image translation using a single encoder-decoder architecture.
method Single auto-encoder with independent domain and content encodings.
result Cross-domain mapping achieved without separate encoders.
Most content-based image retrieval systems consider either one single query, or multiple queries that include the same object or represent the same semantic information. In this paper we consider the content-based image retrieval problem for multiple query images corresponding to different image semantics. We propose a…
Method interprets GAN latent space via latent variable correlation analysis.
problem Understanding the inner workings of GANs.
method Analyzing correlation between latent variables and semantic contents in generated images.
result A method for controllable semantic content generation in GANs.
New method detects unusual images in large datasets.
problem Detecting new images in large image data sets.
method Combines novelty detection with CNN image features.
result Rapid discovery with interpretable explanations.
SLUG method detects bias and out-of-distribution content in generative models.
problem Generative models can underrepresent certain groups and fail on out-of-distribution data.
method SLUG: A new uncertainty quantification method for VAEs combining Laplace approximations and stochastic trace estimators.
result SLUG's UQ score correlates with bias and out-of-distribution content.
This paper improves image retrieval accuracy through novel relevance feedback methods.
problem Improving image retrieval accuracy in Content-Based Image Retrieval (CBIR).
method Novel addition to feature re-weighting and classification techniques, focusing on 0-th iteration improvement.
result Significantly improved retrieval accuracy from relevance feedback.
End-to-end framework for static image generation from dynamic content.
problem Generating static images from dynamic content with occluded backgrounds.
method Conditional GAN for static image generation, convolutional network for dynamic object detection.
result Generated static images are realistic and can be used for augmented reality and robot localization.
Proposes Gaussian optimal transport for image style transfer.
problem Image style transfer and mixing different artistic styles.
method Encoder/Decoder framework with optimal transport for Gaussian measures.
result Simple methodology for generating stylized content interpolating between many styles.
Paper classifies brain signals using eigenvalues for 2D and 3D educational content questions.
problem Classifying brain signals for 2D and 3D educational content questions.
method Eigenvalues of covariance matrix used as features; KNN and SVM classifiers applied.
result No significant difference in learning, memory retention, and recall between 2D and 3D educational content.
New model learns content and transformation separately from data.
problem Learning disentangled representations from data without explicit labels.
method Group-based variational autoencoders, assuming content and transformation groups.
result Model learns generalizable content representations from unseen data.
Paper introduces Latent-CLIP for efficient text-image comparison in latent space.
problem Efficiently compare text and images in latent space without costly decoding.
method Trains CLIP model in latent space, uses Latent-CLIP rewards for noise optimization, and guides generation away from harmful content.
result Latent-CLIP matches CLIP performance on text-image classification and harmful content detection.
The ability to characterize the color content of natural imagery is an important application of image processing. The pixel by pixel coloring of images may be viewed naturally as points in color space, and the inherent structure and distribution of these points affords a quantization, through clustering, of the color i…
This work improves disentanglement by preventing style variables from encoding content-related features.
problem Disentanglement of content and style in data representations using Variational Autoencoders.
method Adversarial training with mutual information minimization to prevent content information leakage in style representations.
result The method efficiently separates content and style related attributes and generalizes to unseen data.
Review of Neural Style Transfer algorithms and their applications.
problem Creating artistic imagery by separating and recombining image content and style.
method Taxonomy and evaluation of current Neural Style Transfer algorithms.
result Comparison and discussion of different NST algorithms.
Study reveals differences in medical image models' hidden representation refinement.
problem Understanding how intrinsic dimensionality changes in neural network hidden representations across different domains.
method Analysis of 11 natural and medical image datasets using 6 network architectures.
result Medical image models refine hidden representations earlier, suggesting differences in feature abstraction.
Paper generates personalized fonts from a few characters.
problem Creating personalized fonts from a limited set of characters.
method Designs a network framework to extract and recombine character content and style using various neural networks.
result Generated characters are structurally similar to real characters.
ITAL uses mutual information for active learning in image retrieval.
problem Acquiring meaningful user feedback for content-based image retrieval.
method Information-Theoretic Active Learning (ITAL) maximizing mutual information between predicted relevance and user feedback.
result ITAL achieves state-of-the-art performance across various datasets.
We present a Bayesian nonparametric framework for multilevel clustering which utilizes group-level context information to simultaneously discover low-dimensional structures of the group contents and partitions groups into clusters. Using the Dirichlet process as the building block, our model constructs a product base-m…
Self-supervised method improves CBIR of CT liver images.
problem Limited labeled data and lack of transparency in deep CBIR systems.
method Proposes a self-supervised learning framework with domain-knowledge integration.
result Improved performance and generalization across datasets.
System converts 3D lung nodule images into embeddings for retrieval.
problem Retrieving similar 3D lung nodule images for radiologist decision support.
method 3D deep learning, semantic representation, transfer learning, similarity score.
result System can measure similarity between nodule annotations and CBIR results.
Jointly trains images and videos using residual vectors.
problem Generating high-quality videos from images and vice versa.
method Simultaneously learns latent variables for images and videos using residual vectors.
result Improves sample quality and diversity in video generation and image generation.
Compression-based similarity measures are effectively employed in applications on diverse data types with a basically parameter-free approach. Nevertheless, there are problems in applying these techniques to medium-to-large datasets which have been seldom addressed. This paper proposes a similarity measure based on com…
Generative Adversarial Network purifies images from steganography without degrading quality.
problem Destruction of image steganography while maintaining visual quality.
method Generative Adversarial Network (GAN) optimized for steganography destruction.
result High rate of steganographic content destruction with minimal visual quality degradation.
End-to-end deep generative model for video compression.
problem Efficiently compressing video data with deep learning.
method Variational autoencoder (VAE) for sequential data, combined with neural image compression techniques.
result Our model achieves competitive rate-distortion results on diverse video content.
UMRL network tackles single image de-raining by learning rain content at different scales.
problem De-raining single images with varying rain streaks of different sizes, directions, and densities.
method UMRL network learns rain content at different scales and uses confidence measures to guide learning.
result UMRL achieves significant improvements over state-of-the-art methods.
TimbreTron transfers musical timbre using CQT and WaveNet.
problem Transfer musical timbre while preserving pitch, rhythm, and loudness.
method Apply image domain style transfer to CQT representation, then generate high-quality waveform with WaveNet.
result TimbreTron recognizably transfers timbre while preserving musical content.
Deep neural net improves low-light image denoising.
problem Noise in low-light images captured by mobile devices.
method Intelligent integration of multiple short noisy frames using a recurrent fully convolutional deep neural net (CNN).
result Achieves state-of-the-art denoising results on burst datasets.
DESS MRI and kernel learning enable precise myelin water quantification.
problem Quantifying myelin water content in the brain.
method Optimized DESS scans combined with kernel learning for precise estimation of myelin water fraction.
result DESS PERK ff estimates are quantitatively similar to conventional MESE MWF estimates.