ZegOT uses optimal transport to zero-shot segment images with text prompts.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Generative model downgrades coarse satellite images to fine resolution.
New unsupervised learning task improves RL performance.
LDA improves image classification accuracy with fewer features.
Method adapts frozen models for few-shot tasks without training.
S-GAI initializes MLPs using spectral geometry from data, improving performance.
New method finds compact strong lottery tickets in partially frozen networks.
Robust CLIP improves vision models' resistance to attacks.
Revises Bayesian model averaging for foundation models.
The study examines how language models learn to represent the world, identifying conditions for ecological veridicality.
Framework adds invariance to pretrained networks without fine-tuning.
This study rethinks the latent space in generative modeling, improving performance with less complex models.
CFA improves model's ability to generalize across unseen domain-class combinations.
Single auto-encoder learns cross-domain image translation.
Spatial Adapter adds structured spatial representation to frozen predictors.
Image denoising is always a challenging task in the field of computer vision and image processing. In this paper, we have proposed an encoder-decoder model with direct attention, which is capable of denoising and reconstruct highly corrupted images. Our model consists of an encoder and a decoder, where the encoder is a…
DistillKac generates images quickly using damped wave equations.
In "extreme" computational imaging that collects extremely undersampled or noisy measurements, obtaining an accurate image within a reasonable computing time is challenging. Incorporating image mapping convolutional neural networks (CNN) into iterative image recovery has great potential to resolve this issue. This pape…
For bidirectional joint image-text modeling, we develop variational hetero-encoder (VHE) randomized generative adversarial network (GAN), a versatile deep generative model that integrates a probabilistic text decoder, probabilistic image encoder, and GAN into a coherent end-to-end multi-modality learning framework. VHE…
We define systems of pre-extremals for the energy functional of regular rheonomic Lagrange manifolds and show how they induce well-defined Hamilton orthogonal nets. Such nets have applications in the modelling of e.g. wildfire spread under time- and space-dependent conditions. The time function inherited from such a Ha…
Generative adversarial networks (GANs) have demonstrated to be successful at generating realistic real-world images. In this paper we compare various GAN techniques, both supervised and unsupervised. The effects on training stability of different objective functions are compared. We add an encoder to the network, makin…
A novel approach stores encoded images as centroids and covariance matrices to improve classification accuracy with less memory.
Humans are able to imagine a person's voice from the person's appearance and imagine the person's appearance from his/her voice. In this paper, we make the first attempt to develop a method that can convert speech into a voice that matches an input face image and generate a face image that matches the voice of the inpu…
New method generates synthetic time series paths with more flexibility.
Bidirectional VAE reduces parameters and improves image tasks.
This paper explores a new framework for lossy image encryption and decryption using a simple shallow encoder neural network E for encryption, and a complex deep decoder neural network D for decryption. E is kept simple so that encoding can be done on low power and portable devices and can in principle be any nonlinear …
A new method, REC, compresses images by encoding their latent representations efficiently.
Reconstructing observed images from fMRI brain recordings is challenging. Unfortunately, acquiring sufficient "labeled" pairs of {Image, fMRI} (i.e., images with their corresponding fMRI responses) to span the huge space of natural images is prohibitive for many reasons. We present a novel approach which, in addition t…
Generative adversarial networks (GANs) are capable of producing high quality image samples. However, unlike variational autoencoders (VAEs), GANs lack encoders that provide the inverse mapping for the generators, i.e., encode images back to the latent space. In this work, we consider adversarially learned generative mo…
WEINCE improves contrastive learning by correcting softmax biases.
Unified principle LZN unifies generative modeling, representation learning, and classification.
This paper introduces a novel generative encoder (GE) model for generative imaging and image processing with applications in compressed sensing and imaging, image compression, denoising, inpainting, deblurring, and super-resolution. The GE model consists of a pre-training phase and a solving phase. In the pre-training …
For surfaces, we brush a reasonably sharp picture of the influence of the fundamental group upon the complexity of foliated-dynamics. A metaphor emerges with phase-changes through the solid-liquid-gaseous states. Groups of ranks are frozen with intransitivity reigning ubiquitously. When , th…
In this manuscript we propose two objective terms for neural image compression: a compression objective and a cycle loss. These terms are applied on the encoder output of an autoencoder and are used in combination with reconstruction losses. The compression objective encourages sparsity and low entropy in the activatio…
Medical imaging models may encode demographic attributes without violating fairness, depending on the approach.
We introduce a novel generative autoencoder network model that learns to encode and reconstruct images with high quality and resolution, and supports smooth random sampling from the latent space of the encoder. Generative adversarial networks (GANs) are known for their ability to simulate random high-quality images, bu…
TSSC images enhance chaotic signal classification using ConvNets.
Skull stripping is usually the first step for most brain analysisprocess in magnetic resonance images. A lot of deep learn-ing neural network based methods have been developed toachieve higher accuracy. Since the 3D deep learning modelssuffer from high computational cost and are subject to GPUmemory limit challenge, a …
In this work, we propose an end-to-end block-based auto-encoder system for image compression. We introduce novel contributions to neural-network based image compression, mainly in achieving binarization simulation, variable bit rates with multiple networks, entropy-friendly representations, inference-stage code optimiz…
Image compression is an essential approach for decreasing the size in bytes of the image without deteriorating the quality of it. Typically, classic algorithms are used but recently deep-learning has been successfully applied. In this work, is presented a deep super-resolution work-flow for image compression that maps …
Performance of neural networks can be significantly improved by encoding known invariance for particular tasks. Many image classification tasks, such as those related to cellular imaging, exhibit invariance to rotation. We present a novel scheme using the magnitude response of the 2D-discrete-Fourier transform (2D-DFT)…
CAFLOW uses auto-regressive flows to translate images efficiently.
In this paper, we present UNet++, a new, more powerful architecture for medical image segmentation. Our architecture is essentially a deeply-supervised encoder-decoder network where the encoder and decoder sub-networks are connected through a series of nested, dense skip pathways. The re-designed skip pathways aim at r…
This work proposes a novel autoencoder for fusing visible and infrared images.
End-to-end meta-learned system for image compression.
This study presents a new lossy image compression method that utilizes the multi-scale features of natural images. Our model consists of two networks: multi-scale lossy autoencoder and parallel multi-scale lossless coder. The multi-scale lossy autoencoder extracts the multi-scale image features to quantized variables a…
Deep learning transforms time series into images for anomaly detection in industrial assets.
Recently deep learning-based methods have been applied in image compression and achieved many promising results. In this paper, we propose an improved hybrid layered image compression framework by combining deep learning and the traditional image codecs. At the encoder, we first use a convolutional neural network (CNN)…