Benchmark improves object detection robustness in winter weather.
problem Assessing object detection models' performance under image corruptions.
method Developed three benchmark datasets with various image corruptions; used data augmentation to improve robustness.
result Simple data augmentation significantly enhances model robustness across different corruptions and datasets.
A dataset of 10 molecule types for machine learning studies.
problem Lack of suitable datasets for machine learning in molecular imaging.
method Generated 2D cross-sectional projections of 10 molecule types from Molecular Dynamics trajectories.
result Benchmark dataset for machine learning, deep learning, and image processing in scattering, imaging, and microscopy.
We present Fashion-MNIST, a new dataset comprising of 28x28 grayscale images of 70,000 fashion products from 10 categories, with 7,000 images per category. The training set has 60,000 images and the test set has 10,000 images. Fashion-MNIST is intended to serve as a direct drop-in replacement for the original MNIST dat…
FLAIR dataset for federated learning benchmarks.
problem Lack of suitable federated learning datasets.
method Curated large-scale annotated image dataset for multi-label classification.
result FLAIR captures real-world federated learning challenges.
This paper benchmarks OoDD methods for medical imaging.
problem Medical models trained for one domain may fail on images from a different domain.
method Defined 3 categories of OoD examples and benchmarked methods in 3 medical imaging domains.
result Simple binary classifier on feature representation yields best accuracy and AUPRC.
COOS-7 dataset benchmarks image classifier generalization.
problem Measuring generalization of image classifiers to out-of-sample datasets.
method Created COOS-7 dataset with varying covariate shifts; benchmarked multiple models.
result All classifiers failed to generalize to datasets with greater covariate shifts.
Deep learning model predicts tropical cyclone intensification using satellite images.
problem Accurately predicting rapid intensification of tropical cyclones.
method Attention-based deep learning model using satellite images.
result Deep learning models outperform traditional methods in RI prediction.
Paper benchmarks adversarial robustness methods on image classification.
problem Vulnerability of deep neural networks to adversarial examples.
method Established a comprehensive benchmark with robustness curves.
result Found important findings on adversarial attack and defense methods.
New benchmarks measure image generation models' ability to generalize beyond training data.
problem Trivially memorizing training data yields better scores than state-of-the-art models on current benchmarks.
method Developed neural network divergences (NNDs) as evaluation metrics requiring large samples.
result Implemented and validated a black-box metric that measures diversity, sample quality, and generalization.
Public dataset for benchmarking deep learning CT reconstruction methods.
problem Lack of a fair benchmark for comparing deep learning CT reconstruction methods.
method Processed and simulated over 40,000 CT scan slices from the LIDC/IDRI Database.
result First baseline results provided for comparison.
A DenseNet model classifies metastatic cancer in medical images.
problem Classifying metastatic cancer in medical images efficiently and accurately.
method Proposes a DenseNet-based model for metastatic cancer classification on medical images.
result The proposed model outperformed other classical methods like Resnet34, Vgg19.
Paper introduces FJD metric for cGAN benchmarking.
problem Quantitative evaluation of cGANs using multiple metrics.
method Frechet Joint Distance (FJD) metric.
result FJD provides a single metric for cGAN benchmarking and model selection.
DRSVM uses deep learning to rank relative attributes between image pairs.
problem Classifying relative attributes between image pairs.
method Deep Siamese network with rank SVM loss function.
result DRSVM outperforms state-of-the-art methods on multiple datasets.
This paper benchmarks privacy-preserving machine learning on medical images.
problem Ensuring privacy in medical image analysis while maintaining model accuracy.
method Comparing Local-DP and DP-SGD for differential privacy in medical imagery.
result Theoretical privacy guarantees do not fully align with real-world performance.
This study benchmarks deep learning for unsupervised near-duplicate image detection.
problem Detecting near-duplicates in large image datasets with high specificity.
method Binary classification using Receiver Operating Curve (ROC) for comparison of different descriptors.
result Fine-tuning deep convolutional networks generally outperforms off-the-shelf features, with best performance on MFND dataset.
A fundamental challenge in calcium imaging has been to infer the timing of action potentials from the measured noisy calcium fluorescence traces. We systematically evaluate a range of spike inference algorithms on a large benchmark dataset recorded from varying neural tissue (V1 and retina) using different calcium indi…
Study subjective perception of low light restored images and develop an unsupervised QA model.
problem Lack of subjective QA for low light restored images and challenges in collecting human opinion scores.
method Create a dataset, conduct subjective QA study, develop self-supervised contrastive learning technique to extract features.
result Unsupervised NR QA model achieves state-of-the-art performance for low light restored images.
Combines variational and evolutionary optimization for generative models.
problem Optimizing generative models with discrete latent variables.
method Truncated posteriors as variational distributions, evolutionary algorithms applied to variational parameters.
result Evolutionary algorithms effectively optimize variational bounds for generative models.
Facebook's ResNeXt WSL models show exceptional robustness against image corruptions and adversarial attacks.
problem Image recognition model robustness against corruptions and adversarial attacks.
method Training with 1B images from Instagram and fine-tuning on ImageNet.
result ResNeXt WSL models achieve state-of-the-art results on ImageNet-C, ImageNet-P, and ImageNet-A.
Paper introduces a real-world image dataset for federated learning.
problem Lack of high-quality real-world data for federated learning.
method Created a dataset from 26 street cameras and 7 categories, implemented YOLO and Faster R-CNN.
result Benchmarked model performance, efficiency, and communication in federated learning.
SPN and multidimensional upsampling generate high-fidelity images from small inputs.
problem Generating high-fidelity images from small inputs.
method Subscale Pixel Network (SPN) and multidimensional upsampling.
result Achieved state-of-the-art likelihood results and high-fidelity samples.
C3 compresses images and videos with low complexity and high performance.
problem High complexity and low performance in neural compression models.
method Overfits a small model to each image or video separately, improving RD performance with low complexity.
result Matches the RD performance of state-of-the-art neural and video codecs with significantly lower decoding complexity.
Unified model trained on images and videos using masked autoencoding.
problem Training a single model for multiple visual modalities.
method Masked autoencoding on a Vision Transformer.
result Unified model achieves comparable or better performance than single-modality models.
New benchmark evaluates BDL methods in medical retinopathy diagnosis.
problem Evaluate robustness and scalability of BDL methods in medical applications.
method Developed a new benchmark with real-world diabetic retinopathy tasks.
result Some BDL techniques overfit uncertainty to datasets, underperforming on new benchmark.
Pythae is a Python library for benchmarking VAE models.
problem Improving variational autoencoders for various tasks.
method Unified implementation and framework for 19 generative autoencoder models.
result Benchmarking 19 VAE models across multiple tasks.
Generative model learns to autoencode and generate sets of images.
problem Learning to represent and generate sets of images with unknown number of sets.
method Set Distribution Networks (SDNs) learn set encoder, discriminator, generator, and prior.
result SDNs can reconstruct and generate sets of images with preserved attributes.
Paper introduces Latent-CLIP for efficient text-image comparison in latent space.
problem Efficiently compare text and images in latent space without costly decoding.
method Trains CLIP model in latent space, uses Latent-CLIP rewards for noise optimization, and guides generation away from harmful content.
result Latent-CLIP matches CLIP performance on text-image classification and harmful content detection.
This study benchmarks algorithms for automatic segmentation of LGE-MRI images of the left atrium.
problem Challenging segmentation of LGE-MRI images due to low contrast.
method Organized a large-scale benchmarking challenge with 154 3D LGE-MRIs and 27 teams.
result Top method achieved 93.2% dice score and 0.7 mm mean surface to surface distance.
DGE learns event representations from image sequences without manual annotations.
problem Data hunger and domain adaptation issues in self-supervised learning for temporal segmentation.
method Dynamic Graph Embedding (DGE) learns event representations by iteratively updating a graph and its embedding.
result DGE achieves robust temporal segmentation on benchmark datasets, outperforming state-of-the-art methods.
Baseline for few-shot image classification outperforms state-of-the-art.
problem Few-shot image classification challenges.
method Fine-tuning deep networks trained with cross-entropy loss, transductively.
result Outperforms state-of-the-art on various datasets.
Multimodal bitransformer boosts image-text classification.
problem Combining text and image modalities for improved classification.
method Supervised multimodal bitransformer model integrating text and image encoders.
result State-of-the-art performance on multimodal classification benchmarks.
Atlas dataset categorizes clothing products with high accuracy.
problem Lack of real-world datasets for e-commerce clothing product categorization.
method Collected and labeled a dataset of 186,150 images, established a benchmark for image classification and sequence models.
result Benchmark model achieved a micro f-score of 0.92.
OpenDataVal benchmarks data valuation algorithms for diverse datasets.
problem Improving model performance and mitigating biases in training datasets.
method Unified benchmark framework for data valuation algorithms.
result No single algorithm performs uniformly best across all tasks.
Public data boosts MICCAI papers' citations, but sharing remains low.
problem Slow adoption of public benchmarks in medical image computing.
method Analysis of MICCAI papers from 2014-2018.
result Public data boosts citations, but sharing is low.
RobustBench aims to standardize adversarial robustness evaluation in image classification.
problem Lack of systematic understanding and error-prone robustness evaluations.
method Standardized benchmark with restricted models and adaptive attacks.
result Reflects current state of the art in adversarial robustness.
MaCow improves flow-based models for image density estimation.
problem Flow-based models struggle with density estimation compared to autoregressive models.
method Introduced masked convolutional generative flow (MaCow) using masked convolution.
result Significant improvements in density estimation on image benchmarks.
This paper benchmarks neural network robustness to corruptions and perturbations.
problem Establishing benchmarks for image classifier robustness to corruptions and perturbations.
method Developed ImageNet-C and ImageNet-P datasets to evaluate robustness to corruptions and perturbations, not adversarial attacks.
result There are negligible changes in relative corruption robustness from AlexNet to ResNet classifiers.
It has recently been observed that certain extremely simple feature encoding techniques are able to achieve state of the art performance on several standard image classification benchmarks including deep belief networks, convolutional nets, factored RBMs, mcRBMs, convolutional RBMs, sparse autoencoders and several othe…
Neural DEs improve single image super-resolution.
problem Challenging tasks in image super-resolution.
method Applied Neural Differential Equations to image super-resolution, using variational methods and backpropagation.
result Differential models match state-of-the-art performance.
Generative modeling over natural images is one of the most fundamental machine learning problems. However, few modern generative models, including Wasserstein Generative Adversarial Nets (WGANs), are studied on manifold-valued images that are frequently encountered in real-world applications. To fill the gap, this pape…
This paper benchmarks speech LVMs against deterministic models and adapts a video model to speech.
problem Speech generation models are inferior to deterministic models.
method Developed a speech benchmark of LVMs and compared them against deterministic models.
result The Clockwork VAE outperforms previous LVMs and reduces the gap to deterministic models.
Review of deep learning methods in medical image registration.
problem Improving accuracy and efficiency of medical image registration.
method Classification and detailed analysis of seven categories of DL-based registration methods.
result Comprehensive comparison of DL-based methods for lung and brain registration.
Survey of deep learning methods for fMRI natural image reconstruction.
problem Reconstructing natural images from fMRI brain activity.
method Survey of deep learning approaches, including architectural design, datasets, and evaluation metrics.
result Performance evaluation across standardized metrics.
This work improves medical image segmentation with limited annotations using contrastive learning.
problem Lack of labeled data for medical image segmentation.
method Contrastive learning framework for semi-supervised segmentation with domain-specific and problem-specific cues.
result Significant improvements in segmentation performance compared to other methods.
New image restoration method using localized patches and external databases.
problem Image restoration challenges.
method Localized structured prediction and non-linear multi-task learning for optimizing a penalized energy function.
result Strong statistical guarantees and practical effectiveness demonstrated on various image restoration problems.
Few-shot unsupervised image-to-image translation model learns from a few examples.
problem Current unsupervised image-to-image translation methods require many images at training time.
method Coupling adversarial training with a novel network design for few-shot learning.
result Model achieves effective few-shot image-to-image translation.
New method speeds up image denoising models without sacrificing performance.
problem Efficiently training models for image denoising tasks.
method Introduces superkernel techniques for fast training of dense prediction models.
result Demonstrates effectiveness on SIDD+ benchmark with 6-8 RTX2080 GPU hours.
Neural network estimates rigid motion in stroke imaging to improve image quality.
problem Rigid patient motion during C-arm CBCT imaging reduces image quality.
method Neural network trained to regress reprojection error based on image information.
result Neural network outperforms entropy-based method in motion estimation.