Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

10192938 · Jun 202019922001200920182026
48 results for white-box defenses

New research shows many recent defenses against adversarial examples are ineffective against black-box attacks.

problem The robustness of recent defenses against adversarial examples is insufficient, especially against black-box attacks.
method Evaluation of nine defenses on two black-box adversarial models and six attacks on CIFAR-10 and Fashion-MNIST datasets.
result Most recent defenses provide only marginal improvements in security (<25%<25\%) compared to undefended networks.

New defense method inspired by encryption improves visual classification accuracy.

problem Conventional defenses reduce accuracy and are defeated by obfuscated gradients.
method Block-wise pixel shuffling with secret key for training and test images.
result Achieves high accuracy (91.55%) on clean images and (89.66%) on adversarial examples.

Defends classifiers from adversarial attacks using self-supervised data estimation.

problem Protecting classifiers from adversarial attacks with full attacker access.
method RIDE, a self-supervised learning algorithm for individual data estimation.
result Significant improvement in adversarial defense performance (98%, 76%, 43% test accuracy on MNIST, CIFAR-10, and ImageNet datasets respectively).

MulDef defends neural networks against adversarial examples by combining multiple models.

problem Vulnerability of neural networks to adversarial examples.
method A general defense framework based on multiple models with robustness diversity.
result Substantially improved accuracy on adversarial examples (22-74%) while maintaining similar accuracy on legitimate examples.

Evaluates SHIELD's effectiveness against adaptive adversaries in various threat models.

problem Evaluating SHIELD's efficacy against adaptive adversaries in different threat models.
method Empirical analysis of SHIELD's robustness against adaptive attacks using Projected Gradient Descent (PGD) attacks in various threat models (white-box, gray-box).
result The targeted PGD attack success rate drops from 64.3% to 48.9% when models are trained from scratch instead of retrained.

ATHENA builds flexible defenses against adversarial attacks.

problem Extensive research on adversarial attacks is domain-specific and cannot be easily extended.
method Designing an extensible framework based on diverse weak defenses.
result Comprehensive empirical study demonstrates the effectiveness of ATHENA.

Study various image defenses against adversarial attacks using basis functions.

problem Defending against adversarial images in deep networks.
method Experiments with low-pass filtering, PCA, JPEG compression, wavelet approximation, and soft-thresholding.
result JPEG compression and soft-thresholding outperform other defenses in most settings, with specific cases showing different performance.

Defensive distillation fails against targeted adversarial attacks, forcing a tradeoff between learning and security.

problem Defensive distillation's limitations in blocking targeted adversarial attacks.
method Systematic exploration of defensive distillation's effectiveness and limitations.
result Defensive distillation is effective against non-targeted attacks but fails against targeted attacks, necessitating a tradeoff between learning and security.

S2SNets defend against adversarial attacks by interpreting perturbations as signal.

problem Fragility of deep neural networks to adversarial attacks.
method Two-stage training of S2SNets: unsupervised first, fine-tuning second, using classifier gradients.
result S2SNets achieve comparable resilience in white-box attacks and robustness in gray-box attacks.

Paper introduces a new method to create adversarial examples against gradient-obfuscating defenses.

problem Crafting adversarial examples to fool gradient-obfuscating defenses.
method Stochastic Substitute Training (SST), a gray-box approach.
result Adversaries can create adversarial examples without knowledge or limited information about the defense.

ME-Net defends neural nets against adversarial attacks by reconstructing images.

problem Adversarial attacks on deep neural networks.
method ME-Net uses matrix estimation to preprocess images, destroying adversarial noise while preserving global structure.
result ME-Net consistently outperforms state-of-the-art defenses on various benchmarks.

Research evaluates adversarial attacks and defenses on 3D point cloud classifiers.

problem Robustness of 3D object classifiers against adversarial attacks.
method Extending 2D adversarial attacks to 3D point clouds and proposing new defenses.
result 3D point cloud classifiers are weak to adversarial attacks but more defensible.

Paper proposes HRS to improve neural network robustness without significant accuracy loss.

problem Vulnerability of neural networks to adversarial attacks and performance degradation.
method Hierarchical Random Switching (HRS) for robustness without sacrificing accuracy.
result HRS significantly improves adversarial robustness with minimal accuracy loss.

Tricks adversarial attacks to target specific classes, improving classifier accuracy.

problem Recent adversarial defense approaches have failed to protect classifiers from untargeted attacks.
method Target Training defense tricks untargeted attacks into targeted attacks on designated classes, then derives the real class.
result 86.2% accuracy for CW-L2 (confidence=0) in CIFAR10, outperforming unsecured classifiers.

This paper reviews defenses against adversarial learning attacks on statistical classifiers.

problem Adversarial attacks on machine learning systems, particularly statistical classifiers.
method Survey of recent work on test-time evasion, data poisoning, and reverse engineering attacks and defenses.
result Novel insights that challenge conventional AL wisdom and target unresolved issues.

GanDef uses GANs to defend against adversarial examples in neural networks.

problem Defending against adversarial examples in neural networks.
method GAN-based adversarial training defense using a competition game to regulate feature selection.
result GanDef trains a classifier to defend against adversarial examples with high accuracy.

RTFE provides adversarial robustness to multiple models.

problem Adversarial examples can transfer to other models, compromising robustness.
method Proposes RTFE, a deep learning-based pre-processing mechanism.
result RTFE provides adversarial robustness to multiple independently trained classifiers.

Denoised smoothing defends pretrained classifiers against adversarial attacks.

problem Adversarial attacks on pretrained classifiers.
method Prepending a denoiser to any off-the-shelf classifier using randomized smoothing.
result Guaranteed p\ell_p-robustness to adversarial examples without modifying the pretrained classifier.

This paper explores RL for cyber defense in SDN, resisting poisoning attacks.

problem Adversaries exploit ML adaptability to poison training and evade classification.
method Investigates RL algorithms for autonomous cyber defense in SDN, studying various attack types.
result RL agents can effectively react to poisoning attacks in SDN.

Detects adversarial examples in deep speech recognition.

problem Vulnerability of deep speech recognition systems to adversarial attacks.
method Formulated as a classification problem, generated adversarial and normal datasets, trained CNN.
result Accurately distinguishes between adversarial and normal examples for known attacks.

Proposes RPG-RT for red-teaming T2I models without internal access.

problem Evaluating T2I models' security through red-teaming is challenging due to their closed-source nature and unknown defense mechanisms.
method Integrates LLM and rule-based preference modeling to dynamically adapt to unknown defense mechanisms.
result Demonstrates superior and practical approach for red-teaming T2I models.

The paper proposes logit regularization for improving neural network robustness against adversarial attacks.

problem Vulnerability of neural networks to adversarial examples, especially in safety-critical applications.
method Logit regularization techniques combined with other methods to enhance adversarial robustness.
result Logit regularization can improve the effectiveness of existing adversarial defenses and create stronger black-box attacks.

Paper proposes a new black-box attack method for deep neural networks.

problem Understanding and defending against adversarial attacks on deep neural networks.
method The algorithm finds a probability distribution over a region around the input, making a sample likely an adversarial example.
result The method outperforms state-of-the-art black-box or white-box attack methods for most test cases.

A new activation function k-WTA improves neural network defenses against adversarial attacks.

problem Improving neural network robustness against gradient-based adversarial attacks.
method Proposes k-Winners-Take-All activation function and analyzes its effectiveness.
result k-WTA activation significantly enhances neural network robustness against adversarial attacks.

SONet stabilizes ODE networks for robustness without adversarial training.

problem Improving adversarial robustness of neural networks without sacrificing natural accuracy.
method SONet uses skew-symmetric ODE blocks and DOPRI5 solver for robustness.
result SONet achieves comparable robustness to adversarial defense methods without trade-off.

FineFool attacks deep models by focusing on object contours, improving attack performance.

problem Adversarial attacks on deep learning models, especially those that focus on perturbation size and success rate.
method FineFool uses attention to focus on object contours, producing more efficient and imperceptible perturbations.
result FineFool achieves better attack performance compared to state-of-the-art attacks, including higher success rate and smaller perturbations.

BlurNet defends against adversarial attacks by filtering feature maps.

problem Adversarial attacks on deep neural networks, especially for image classification.
method BlurNet introduces a depthwise convolution layer with standard blur kernels after the first layer to filter high frequency noise.
result The defense reduces the success rate of adversarial attacks from 90% to 20% with total variation regularization.

The paper investigates how Gaussian Processes handle adversarial examples and their uncertainty.

problem Adversarial examples' impact on uncertainty in Gaussian Process models.
method Investigates Gaussian Processes in the context of Bayesian inference to study adversarial examples.
result Gaussian Processes show varying levels of uncertainty that reflect adversarial perturbations.

SmoothFool efficiently computes smooth adversarial perturbations for deep networks.

problem Vulnerability of deep neural networks to adversarial attacks with specific statistical properties.
method SmoothFool: a general and computationally efficient framework for computing smooth adversarial perturbations.
result Smoothness significantly enhances robustness against adversarial attacks and improves transferability.

Colored noise improves neural network robustness against adversarial attacks.

problem Vulnerability of neural networks to adversarial perturbations.
method Injection of colored noise into network weights and activations during adversarial training.
result Our approach outperforms previous methods in terms of adversarial accuracy on CIFAR-10 and CIFAR-100 datasets.

Deep Latent Defence combines adversarial training with a detection system to protect neural networks.

problem Vulnerability of neural networks to adversarial attacks, especially those that cause misclassification.
method Adversarial training combined with a kk-NN classifier in a latent space.
result Deep Latent Defence effectively detects and mitigates adversarial attacks, even under strong attack models.

Improved neural network robustness to adversarial attacks through smoothed inference.

problem Vulnerability of deep neural networks to adversarial attacks.
method Randomized smoothing applied to adversarial training, improving both robustness and performance.
result Significant improvement in accuracy on adversarial attacks (e.g., 60.4% on CIFAR-10 with ResNet-20, outperforming previous methods by 11.7%).