ScieNet improves deep learning resilience to input perturbations.
problem Deep learning's poor resilience to input perturbations in real-world scenarios.
method Hybrid architecture combining SNN for contextual info extraction and DNN for classification.
result Significant improvement in accuracy on noisy and rainy images without prior training.
We propose a novel framework for the differentially private ERM, input perturbation. Existing differentially private ERM implicitly assumed that the data contributors submit their private data to a database expecting that the database invokes a differentially private mechanism for publication of the learned model. In i…
Classifiers such as deep neural networks have been shown to be vulnerable against adversarial perturbations on problems with high-dimensional input space. While adversarial training improves the robustness of image classifiers against such adversarial perturbations, it leaves them sensitive to perturbations on a non-ne…
While deep learning is remarkably successful on perceptual tasks, it was also shown to be vulnerable to adversarial perturbations of the input. These perturbations denote noise added to the input that was generated specifically to fool the system while being quasi-imperceptible for humans. More severely, there even exi…
This research investigates reliable local explanations for machine listening models.
problem Generating reliable local explanations for machine listening models.
method Investigates the sensitivity of SoundLIME explanations to input perturbations and proposes a novel method for identifying suitable content types.
result SoundLIME explanations are sensitive to the content in occluded input regions, and the average magnitude of input mel-spectrogram bins is the most suitable content type for temporal explanations.
Following great success in the image processing field, the idea of adversarial training has been applied to tasks in the natural language processing (NLP) field. One promising approach directly applies adversarial training developed in the image processing field to the input word embedding space instead of the discrete…
It is well-known that the robustness of artificial neural networks (ANNs) is important for their wide ranges of applications. In this paper, we focus on the robustness of the classification ability of a spiking neural network which receives perturbed inputs. Actually, the perturbation is allowed to be arbitrary styles.…
Neural networks are known to be vulnerable to adversarial examples, inputs that have been intentionally perturbed to remain visually similar to the source input, but cause a misclassification. It was recently shown that given a dataset and classifier, there exists so called universal adversarial perturbations, a single…
Jacobian regularization boosts neural network robustness without degrading generalization.
problem Ensuring robustness of machine learning models against input perturbations.
method Developed a computationally efficient Jacobian regularization technique.
result Significant improvements in robustness measured against random and adversarial perturbations.
DANCE improves saliency maps by adding subtle input variations.
problem Poor performance of saliency methods in saturated gradients, adversarial perturbations, and inter-feature dependence.
method Two-step procedure: 1) Perturbation mechanism, 2) Aggregation of saliency maps.
result DANCE saliency method outperforms existing methods qualitatively and quantitatively.
A new efficient PGD method generates smaller perturbation adversarial examples.
problem Adversarial examples in deep neural networks.
method Modified Project Gradient Descent (PGD) method for ensemble models.
result Generates smaller perturbation adversarial examples than PGD method.
This work generates diverse adversarial attacks for different domains using latent variable perturbation.
problem Adversarial attacks on deep neural networks are limited to a single perturbation.
method Frame adversarial attacks as learning a distribution of perturbations, enabling generation of diverse attacks.
result Framework generates competitive or superior adversarial attacks across diverse domains (images, text, graphs).
Adversarial perturbations fool deepfake detectors with high accuracy.
problem Improving deepfake detection accuracy against adversarial attacks.
method Used adversarial perturbations and two defenses: Lipschitz regularization and Deep Image Prior (DIP).
result Deepfake detectors achieved 27% accuracy on perturbed images, compared to 95% on unperturbed.
Adaptive algorithm generates unrestricted adversarial inputs, defeating robust classifiers.
problem Vulnerability of neural networks to unrestricted adversarial inputs.
method Adaptive algorithm for generating unrestricted adversarial inputs.
result Adversarial inputs defeat robust classifiers.
New research evaluates various perturbation methods for improving neural network robustness.
problem Understanding and improving robustness of Convolutional Neural Networks (CNNs) against adversarial attacks.
method Detailed evaluation of five main perturbation-based defenses, comparing random and deterministic approaches.
result Perturbation-based defenses are equivalent in efficacy, and attacks transfer between them.
State-of-the-art machine learning models frequently misclassify inputs that have been perturbed in an adversarial manner. Adversarial perturbations generated for a given input and a specific classifier often seem to be effective on other inputs and even different classifiers. In other words, adversarial perturbations s…
Paper proposes a method to generate adversarial perturbations for black-box attacks without accessing inner states.
problem Generating adversarial perturbations for black-box attacks without accessing inner states of a DNN.
method Matrix-free generation method that requires fewer query trials.
result The proposed method successfully deceives a DNN for semantic segmentation more effectively than random noise.
New adversarial training methods generate multiplicative perturbations for robust DNN training.
problem Training Deep Neural Networks with adversarial examples to improve robustness.
method Proposes xAT and xVAT, generating multiplicative perturbations for robust training.
result xAT and xVAT match or outperform state-of-the-art classification accuracies and are faster.
New method for robustly interpreting ML models using quantile constraints and Wasserstein projections.
problem Assessing robustness of black-box models to input misspecification.
method Quantile-constrained Wasserstein projections for robust interpretability.
result Analytical solution for perturbation problem and smooth perturbations.
DBPA assesses LLM perturbations using frequentist hypothesis testing.
problem Quantifying input perturbation impacts on LLM outputs.
method DBPA reformulates perturbation analysis as frequentist hypothesis testing, using Monte Carlo sampling for empirical null and alternative distributions.
result DBPA provides interpretable p-values and scalar effect sizes for LLM perturbations.
Adversarial weight perturbations can inject backdoors into trained neural models.
problem Security risk of using publicly available trained models due to backdoors.
method Extended adversarial perturbations to model weights, using a composite loss and projected gradient descent.
result Adversarial weight perturbations can be successfully injected with very small changes, exposing security risks across various tasks.
We consider a neural network architecture with randomized features, a sign-splitter, followed by rectified linear units (ReLU). We prove that our architecture exhibits robustness to the input perturbation: the output feature of the neural network exhibits a Lipschitz continuity in terms of the input perturbation. We fu…
AWP improves robustness by flattening weight loss landscape.
problem Improving robustness of deep neural networks against adversarial examples.
method Explicitly regularizes the flatness of weight loss landscape through adversarial weight perturbation.
result AWP forms a double-perturbation mechanism in adversarial training, leading to flatter weight loss landscape.
New neural network units resist adversarial attacks effectively.
problem Adversarial attacks on machine learning models.
method Introduced MWD units, developed training techniques, and computed robustness.
result MWD networks are significantly more robust to adversarial attacks.
The paper introduces extremal perturbations for better attribution analysis in deep networks.
problem Identifying input parts responsible for model outputs.
method Extremal perturbations, smooth masks, and technical innovations for computation.
result Demonstrates excellent sensitivity to spatial properties of deep neural networks.
Lower class selectivity makes networks more robust to natural perturbations but more vulnerable to adversarial attacks.
problem Understanding how class selectivity affects robustness to different types of perturbations in neural networks.
method Investigated the relationship between class selectivity and robustness to natural and adversarial perturbations in neural networks.
result Lower class selectivity increases robustness to natural perturbations but decreases robustness to adversarial attacks.
Study of deep neural networks using finite-time Lyapunov exponents.
problem Understanding the geometric structures in input space formed by deep neural networks.
method Analogy with dynamical systems, computing finite-time Lyapunov exponents.
result Ridges of large positive exponents divide input space into regions associated with different classes.
We propose a new input perturbation mechanism for publishing a covariance matrix to achieve (ε,0)-differential privacy. Our mechanism uses a Wishart distribution to generate matrix noise. In particular, We apply this mechanism to principal component analysis. Our mechanism is able to keep the positive semi-definitene…
Localized uncertainty attacks target uncertain regions to create imperceptible adversarial examples.
problem Adversarial examples that are imperceptible to humans and strong under deterministic classifiers.
method Localized uncertainty attacks by perturbing uncertain regions, using predictive uncertainty or surrogate models.
result Localized uncertainty attacks produce strong adversarial examples that retain input similarity.
BPN defends against adversarial attacks by generating beneficial perturbations.
problem Adversarial attacks cause deep neural networks to misclassify clean inputs.
method BPN generates beneficial perturbations during training to neutralize future adversarial attacks.
result BPN is robust to adversarial examples and more efficient than classical adversarial training.
Unified analysis of removal-based feature attributions robustness.
problem Robustness of removal-based feature attributions is not well understood.
method Theoretical analysis and upper bounds derivation for removal-based feature attributions under input and model perturbations.
result Upper bounds for the difference between intact and perturbed attributions derived under various perturbation settings.
Paper introduces input perturbation for privacy in machine learning models.
problem Protecting both training data and model parameters while maintaining privacy.
method Add noise to training data and train with perturbed data for differential privacy.
result Achieves (ε,δ)-differential privacy on the final model with privacy on original data.
We analyze the adversarial examples problem in terms of a model's fault tolerance with respect to its input. Whereas previous work focuses on arbitrarily strict threat models, i.e., ε-perturbations, we consider arbitrary valid inputs and propose an information-based characteristic for evaluating tolerance to diverse …
Discretizing input space improves DLN robustness against adversarial attacks.
problem Improving machine learning models' resistance to adversarial attacks.
method Input discretization and Binary Neural Networks (BNNs).
result 2-bit input discretization significantly enhances adversarial robustness with minimal accuracy loss.
SmoothLLM defends LLMs from jailbreaking attacks by randomly perturbing inputs.
problem Adversaries can fool large language models into generating objectionable content.
method SmoothLLM randomly perturbs multiple copies of a prompt and aggregates predictions to detect adversarial inputs.
result SmoothLLM sets the state-of-the-art for robustness against various jailbreak attacks.
Residual networks analyzed using linearization for stability under perturbations.
problem Understanding the behavior of residual networks under small input perturbations.
method Linearization of residual units and network stages, using singular value decomposition for stability analysis.
result Most singular values of residual units are 1, but scaling and weights significantly affect them.
Mixup inference improves adversarial robustness by mixing inputs with clean samples.
problem Adversarial examples can fool deep networks due to local non-linearity.
method Develops mixup inference, which mixes inputs with clean samples to shrink adversarial perturbations.
result Mixup inference enhances adversarial robustness for mixup-trained models.
Deep Convolutional Networks (DCNs) have been shown to be vulnerable to adversarial examples---perturbed inputs specifically designed to produce intentional errors in the learning algorithms at test time. Existing input-agnostic adversarial perturbations exhibit interesting visual patterns that are currently unexplained…
This work examines how adversarial vulnerability changes with the dimensionality of the subspace of perturbations.
problem Understanding adversarial vulnerability in constrained input spaces.
method Investigates adversarial vulnerability in subspace V of the input space X with varying dimensions, using PGD attacks and analyzing the dependence on ε and dim(V)/dim(X). result Adversarial success of PGD attacks is a monotonically increasing function of $ε(rac{dim(V)}{dim(X)})^{rac{1}{q}}$.
New geometric interpretation explains over-parameterized models and adversarial perturbations.
problem Geometric understanding of over-parameterized regression and adversarial perturbations.
method Alternative geometric interpretation of regression in feature space.
result Adversarial perturbations are a natural feature of biased models due to underlying geometry.
Paper proposes a method to improve deep learning models' robustness to real-world variations.
problem Deep learning models fail to generalize to small variations of the input.
method Adversarial mixing with disentangled representations to enforce robustness to real-world transformations.
result Improves generalization and reduces spurious correlations, as shown by experiments.
Semantify-NN verifies neural network robustness against semantic perturbations.
problem Verifying robustness of neural networks against semantic adversarial attacks.
method Inserting semantic perturbation layers (SP-layers) into neural networks to verify robustness.
result Semantify-NN significantly improves robustness verification performance over ℓp-norm-based methods. Batch normalization makes neural networks more vulnerable to small adversarial perturbations.
problem Adversarial vulnerability of neural networks trained with batch normalization.
method Investigated the impact of batch normalization on adversarial robustness and compared it to weight decay.
result Substituting weight decay for batch norm nullifies adversarial vulnerability.
The characteristics (or numerical patterns) of a feature vector in the transform domain of a perturbation model differ significantly from those of its corresponding feature vector in the input domain. These differences - caused by the perturbation techniques used for the transformation of feature patterns - degrade the…
LSDAT reduces query efficiency for decision-based adversarial attacks.
problem Improving query efficiency for decision-based adversarial attacks.
method Low-rank and sparse decomposition (LSD) to craft perturbations.
result LSDAT achieves superior fooling rates with fewer queries.
The paper addresses fairness in machine learning by adjusting input distributions.
problem Reducing disparate impact in machine learning models over different groups.
method The approach involves learning a counterfactual distribution to adjust input variables for disadvantaged groups.
result The method can reduce disparate impact without training a new model.
MARGINATTACK improves zero-confidence adversarial attacks' accuracy and efficiency.
problem Improving zero-confidence adversarial attacks' accuracy and efficiency.
method Proposes MARGINATTACK, a zero-confidence attack framework that computes margin with improved accuracy and efficiency.
result MARGINATTACK computes a smaller margin than state-of-the-art zero-confidence attacks and matches state-of-the-art fix-perturbation attacks.
New method enhances neural network robustness against adversarial attacks.
problem Enhancing neural network robustness against adversarial attacks.
method Variational framework with per-sample noise level selector.
result Enhanced empirical robustness and certified robustness.