Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

131262392523 · Jun 202019922001200920182026
48 results for Image representation

The paper proposes a model to learn disentangled representations using mutual information.

problem Learning disentangled representations from shared and exclusive attributes.
method Mutual information maximization for shared attributes and minimization for disentanglement.
result The proposed model outperforms state-of-the-art models in representation disentanglement.

A new image representation method using hypernetworks.

problem Representing images in a way that allows for continuous manipulation and analysis.
method Constructing a hypernetwork that maps pixel positions to colors, allowing for continuous image manipulation.
result Comparable image super-resolution results to existing methods using a single model.

BigBiGAN improves unsupervised representation learning using image generation quality.

problem Improving unsupervised representation learning methods.
method Extending BigGAN to include an encoder and modifying the discriminator for representation learning.
result BigBiGAN models achieve state-of-the-art performance in unsupervised representation learning and unconditional image generation.

Hierarchical autoregressive models improve image quality by learning abstract representations.

problem Local structure bias in autoregressive models leads to lack of large-scale coherence in generated images.
method Propose two methods to learn discrete representations of images that abstract away local detail, and train autoregressive priors on these representations.
result Hierarchical autoregressive models produce high-fidelity reconstructions and realistic images with large-scale coherence.

A new method, REC, compresses images by encoding their latent representations efficiently.

problem Efficiently compressing single images with latent representations.
method Relative Entropy Coding (REC) that directly encodes latent representations with codelength close to relative entropy.
result REC is more efficient for single image compression compared to previous methods and is competitive for lossy compression.

Study reveals differences in medical image models' hidden representation refinement.

problem Understanding how intrinsic dimensionality changes in neural network hidden representations across different domains.
method Analysis of 11 natural and medical image datasets using 6 network architectures.
result Medical image models refine hidden representations earlier, suggesting differences in feature abstraction.

Paper mitigates information leakage in image representations using maximum entropy.

problem Mitigating unintended leakage of user information from image representations.
method Formulates an adversarial non-zero sum game to find an embedding function that maximizes task-dependent discriminative information while minimizing entropy of sensitive attributes.
result Proposed approach learns image representations with high task performance and reduced leakage of sensitive information.

A method to prevent image representation collapse through data-dependent augmentation.

problem Representation collapse due to image augmentations that damage information.
method Formalizing a stochastic encoding process with a tug-of-war between corruption and preserved information, using infoMax objective.
result Learning a data-dependent distribution of augmentations to avoid representation collapse.

The method learns disentangled representations for localized image manipulations.

problem Image generating neural networks are viewed as black boxes with global effects.
method Localized ResNet Autoencoder with multiple loss functions.
result The network can transfer specific facial attributes like shape and color of eyes, hair, mouth, etc. between persons.

For any knot, the following are equivalent. (1) The infinite cyclic cover has uncountably many finite covers; (2) there exists a finite-image representation of the knot group for which the twisted Alexander polynomial vanishes; (3) the knot group admits a finite-image representation such that the image of the fundament…

2007-08-28abs ↗pdf ↗

A novel feature representation method for non-image based features.

problem Inability of Convolutional Neural Networks for non-image based features or features without spatial correlations.
method REFINED: Representation of Features as Images with Neighborhood Dependencies.
result Higher prediction accuracy compared to existing methodologies.

A new model decouples global and local image representations without supervision.

problem Learning decoupled global and local image representations without supervision.
method Variational auto-encoding framework with invertible generative flow.
result The model effectively learns decoupled representations of images.

SketchEmbedNet learns image representations from sketches, useful for few-shot learning.

problem Learning image representations from sketches for few-shot learning.
method Training a model to produce sketches of images, focusing on informative embeddings.
result Model produces informative embeddings of novel images, classes, and datasets.

Deep learning method improves gene ontology classification of neural images.

problem Classifying gene expression in neural in situ hybridization images.
method End-to-end deep learning using convolutional denoising autoencoders (CDAE).
result Significant improvement in classification accuracy (96% reduction in error rate).

Pix2Shape learns 3D scene representations from single images without supervision.

problem Learning 3D scene information from a single image without supervision.
method Pix2Shape uses an encoder, decoder, and critic network to generate 2.5D surfel-based reconstructions.
result Pix2Shape can generate complex 3D scenes from a single image, scaling with on-screen resolution.

This paper addresses the following questions pertaining to the intrinsic dimensionality of any given image representation: (i) estimate its intrinsic dimensionality, (ii) develop a deep neural network based non-linear mapping, dubbed DeepMDS, that transforms the ambient representation to the minimal intrinsic space, an…

2018-03-26abs ↗pdf ↗

This work proposes a model to disentangle image factors effectively and control their manipulation.

problem Controlling disentanglement during image editing while preserving object identity.
method Encoder-decoder architecture with decorrelation regularization and soft target representations.
result The model successfully disentangles image factors and manipulates them effectively.

A framework compares image representations based on local geometry.

problem Comparing image representations based on global structure overlooks local differences.
method Quantify local geometry using Fisher information matrix and optimize differentiation with principal distortions.
result Identifies differences in local sensitivities between models.

New method evaluates text-to-image synthesis for realism, variety, and semantic accuracy.

problem Lack of metrics revealing semantic accuracy in text-to-image synthesis.
method Uses Inception network representations and t-SNE visualization for semantic evaluation.
result Classification accuracy of generated images to real images' visual concepts correlates with semantic accuracy.

New text-to-image diffusion models improve scene understanding for AI agents.

problem Fine-grained scene understanding for AI agents from text and images.
method Pre-trained text-to-image diffusion models optimized for generating images from text prompts.
result Policies learned with Stable Control Representations outperform state-of-the-art approaches on various control tasks.

Solved a specific case of Salter's question on Burau representation.

problem Under what conditions are matrices in the image of the Burau representation of B3B_3.
method Algorithmically constructed a counterexample to Salter's specific question.
result The central quotient of the Burau image group is not the central quotient of a certain subgroup of the unitary group.

CLIP learns joint image-text representations for zero-shot learning.

problem Understanding and improving zero-shot transfer performance in CLIP.
method Formal study of transferrable representation learning and analysis of zero-shot transfer performance.
result Proposes a new CLIP-type approach that outperforms existing methods.

A simple method flags images as out-of-distribution based on their distance to nearest neighbors.

problem Detecting images not aligned with a trained model's in-distribution data.
method Flag images as OOD if their average distance to K nearest neighbors is large in the classifier's representation space.
result Simple methods can outperform more complex ones when considering learned representations.

Paper proposes a method to extract disentangled features for multi-task learning in medical images.

problem Indiscriminate mixing of image properties leads to poor generalization in deep learning.
method Uses deep neural networks and adversarial regularization to disentangle features.
result Demonstrates improved performance on images with new properties like artifacts.

Scattering networks improve image representation learning without deep learning.

problem Improving image representation learning without deep learning.
method Scattering networks as generic representations in scattering space.
result Scattering networks achieve competitive results in supervised and unsupervised learning.

A new image interpolation model using sparse representation and nonlocal linear regression.

problem Image interpolation without blurring and noise.
method Sparse representation, nonlocal self-similarity, nonlocal linear regression, adaptive sub-dictionary learning, weighted encoding.
result Our method outperforms state-of-the-art methods in quantitative measures and visual quality.

Self-guidance controls image generation by extracting properties from diffusion model representations.

problem Generating images from text descriptions is challenging due to the complexity of visual details.
method Self-guidance uses internal representations of diffusion models to control image generation.
result Properties like object shape, location, and appearance can be extracted and used to steer image generation.

Scattering representations simplify SBI for images without extra compression.

problem Efficiently performing simulation-based inference on images with limited data.
method Use scattering representations for compression and learning, combined with spatial averaging and expressive density estimators.
result Scattering representations provide more information than traditional methods, without requiring additional simulations.

CNNs improve medical image classification with few samples.

problem Classifying medical images with limited training data.
method Transfer learning using CNNs, representation extraction, and a novel metric for performance prediction.
result CNN-based transfer learning outperforms feature-based methods with high correlation to test set performance.

Maximizes mutual info across views for better image representations.

problem Improving image representation learning through multiple views.
method Maximizing mutual information between features from multiple views.
result ImageNet accuracy of 68.1% using linear evaluation, significantly outperforming prior methods.

Synthesizes images from audio and visual data using spike-based autoencoders.

problem Extracting meaningful information from spatio-temporal data for image synthesis.
method Spike-based autoencoders trained to learn spatio-temporal representations of audio and visual data.
result Synthesized images from audio samples with high fidelity, achieving competitive performance.

DGE learns event representations from image sequences without manual annotations.

problem Data hunger and domain adaptation issues in self-supervised learning for temporal segmentation.
method Dynamic Graph Embedding (DGE) learns event representations by iteratively updating a graph and its embedding.
result DGE achieves robust temporal segmentation on benchmark datasets, outperforming state-of-the-art methods.

Conventional SVM-based image coding methods are founded on independently restricting the distortion in every image coefficient at some particular image representation. Geometrically, this implies allowing arbitrary signal distortions in an nn-dimensional rectangle defined by the ε\varepsilon-insensitivity zone in eac…

2013-10-18abs ↗pdf ↗

This paper tackles disentanglement in image editing and reconstruction.

problem Learning disentangled image representations and balancing disentanglement strength and reconstruction quality.
method Distance covariance based decorrelation regularization for disentanglement, soft target representation for reconstruction, and collapsing AE decoder and GAN generator.
result The proposed model improves the disentanglement strength and perceptual quality of generated images.

Study finds CLIP's caption-based learning outperforms image-only methods under certain conditions.

problem Comparing CLIP's performance with traditional image-only methods in learning transferable representations.
method Controlled comparison of CLIP and image-only methods using a dataset with descriptive captions and specific criteria.
result CLIP's caption-based learning outperforms image-only methods when certain conditions are met, but can be detrimental in others.