Interval attacks find more adversarial examples than existing methods.
problem Evaluating robustness of adversarially trained neural networks against unknown attacks.
method Symbolic interval propagation for bound over-approximation and gradient-guided attacks.
result Interval attacks find on average 47% more violations than state-of-the-art methods.
Proves the existence of accurate, certifiably robust neural networks.
problem Training neural networks to be robust against adversarial attacks.
method Proves the existence of networks that approximate continuous functions and ensure robustness through interval-bound propagation.
result Proves the existence of accurate, interval-certified ReLU networks.
Efficiently trains robust models with faster and more stable interval bounds.
problem Training verifiably robust models against adversarial attacks.
method Integrates interval arithmetic and an additional cost function term to keep hidden layer bounds small.
result Comparable or better results achieved with fewer training iterations and more stability.
This work improves neural network robustness to symbol substitutions using formal verification.
problem Neural networks' vulnerability to adversarial attacks, especially under discrete text perturbations.
method Formal verification using Interval Bound Propagation on a simplex model of input perturbations.
result Models show improved verified accuracy under perturbations with formal guarantees.
PRISM-FCP improves federated prediction robustness against Byzantine attacks.
problem Byzantine attacks in federated learning.
method Partial model sharing and distance-based maliciousness scores.
result Maintains nominal coverage guarantees under Byzantine attacks.
Study probabilistic safety of BNNs under adversarial attacks.
problem Evaluate vulnerability of BNNs to adversarial attacks.
method Relaxation techniques from non-convex optimization to compute probabilistic safety bounds.
result Certify probabilistic safety of BNNs with millions of parameters.
Improves deep learning robustness by considering task and model.
problem Adversarial attacks on deep learning systems.
method Binary and interval label encoding strategy to redefine classification tasks and design corresponding loss functions.
result Our method enhances robustness without sacrificing accuracy.
IBP-R improves verified adversarial robustness with simple, effective interval bound propagation.
problem Improving verifiability of adversarially trained networks.
method Coupling adversarial attacks with interval bound propagation for minimized verification gap.
result State-of-the-art verified robustness-accuracy trade-offs for small perturbations on CIFAR-10.
New method defends against patch attacks with high-certainty guarantees.
problem Patch attacks on images, especially physical adversarial attacks.
method Randomized smoothing, exploiting patch constraints.
result Meaningfully large robustness certificates against patch attacks.
The paper tackles individual fairness in ML models, developing statistical methods to detect bias.
problem Detecting and measuring violations of individual fairness in machine learning models.
method Formalizing the problem as adversarial attack, developing inference tools for the adversarial cost function.
result Statistical methods to assess and test hypotheses of model fairness with non-coverage error rate control.
RL agents learn to detect and mitigate adversarial attacks.
problem Adversarial attacks on Deep RL algorithms deployed in real-life applications.
method Meta-Learned Advantage Hierarchy (MLAH) agent using meta-learning for online robustness.
result The MLAH agent exhibits hierarchical coping behaviors and maintains higher reward distributions over time.
New MCMC method estimates differential privacy from multiple MIAs without worst-case assumptions.
problem Bayesian estimation of differential privacy from membership inference attacks.
method Bayesian estimation via MCMC algorithm (MCMC-DP-Est).
result More cautious privacy analysis with joint estimation of MIA strengths and privacy parameter.
Adversarial attacks add perturbations to the input features with the intent of changing the classification produced by a machine learning system. Small perturbations can yield adversarial examples which are misclassified despite being virtually indistinguishable from the unperturbed input. Classifiers trained with stan…
Making neural networks robust against adversarial inputs has resulted in an arms race between new defenses and attacks. The most promising defenses, adversarially robust training and verifiably robust training, have limitations that restrict their practical applications. The adversarially robust training only makes the…
Training Deep Neural Networks that are robust to norm bounded adversarial attacks remains an elusive problem. While exact and inexact verification-based methods are generally too expensive to train large networks, it was demonstrated that bounded input intervals can be inexpensively propagated from a layer to another t…
"Feint Attack", as a new type of APT attack, has become the focus of attention. It adopts a multi-stage attacks mode which can be concluded as a combination of virtual attacks and real attacks. Under the cover of virtual attacks, real attacks can achieve the real purpose of the attacker, as a result, it often caused hu…
This paper studies adversarial attacks on Gaussian process bandits.
problem Adversarial attacks on Gaussian process bandits to manipulate optimal function regions.
method Proposes various adversarial attack methods on GP bandits, including white-box and black-box attacks.
result Adversarial attacks can force GP bandits to optima in target regions even with low attack budgets.
Subpopulation attacks poison data to misclassify naturally distributed points.
problem Improving accuracy of machine learning predictions through adversarial data modification.
method Introducing a novel subpopulation attack framework, using influence functions and gradient optimization.
result Subpopulation attacks are effective and stealthy, making them difficult to defend against.
Efficient attacks on DRL models without model access and low computation.
problem Vulnerabilities of DRL models to adversarial attacks.
method Adapting black-box attacks, introducing efficient online sequential attacks, exploring perturbations in environment dynamics, and generating robust physical perturbations.
result Demonstrated the effectiveness of proposed attacks on real-world robots.
Reward-poisoning attacks can force RL agents to learn bad policies, and we categorize and quantify their feasibility.
problem Reward-poisoning attacks can manipulate RL agents to learn undesirable policies.
method Categorize attacks by infinity-norm constraint, provide thresholds for feasibility, and develop adaptive attack strategies.
result Adaptive reward-poisoning attacks can achieve the nefarious policy in polynomial steps, while non-adaptive attacks require exponential steps.
Spanning attack improves black-box attacks with unlabeled data.
problem Query inefficiency in black-box attacks due to high input space dimensionality.
method Proposes spanning attack by constraining adversarial perturbations in a low-dimensional subspace via an auxiliary unlabeled dataset.
result Significantly improves query efficiency of black-box attacks.
Headless attacks bypass classification heads to fool transfer learning models.
problem Adversarial attacks against transfer learning models without access to the classification head.
method Label-blind adversarial attacks that do not require class-label information.
result Transfer attack lowers ResNet18 accuracy on CIFAR10 by over 40%.
New attack manipulates UCB algorithm, new defense algorithm reduces pseudo-regret.
problem Adversarial attacks on stochastic bandit algorithms.
method Introducing action-manipulation attacks and proposing a robust defense algorithm.
result Proposed defense algorithm reduces pseudo-regret to O(max{log T, A}).
This paper explores evasion attacks against Bayesian models.
problem Bayesian predictive models are vulnerable to evasion attacks.
method Developed gradient-based attacks for specific point predictions and entire posterior distributions.
result Optimal evasion attacks can be designed against Bayesian models.
Adversarial attacks pose a threat to deep neural networks, especially in safety-critical applications.
problem Adversarial attacks can misclassify deep neural networks, leading to safety issues.
method Adversarial attacks are categorized into white-box and black-box attacks based on the attacker's knowledge. They can be targeted or non-targeted.
result Adversarial attacks are effective and can transfer between different models and real-world scenarios.
Adversarial attacks hide cyber-physical attacks in ICS.
problem Hiding cyber-physical attacks in industrial control systems.
method Modeling an attacker compromising sensors, manipulating data, and evaluating attacks on both continuous and mixed data.
result Successfully hides cyber-physical attacks with 2.87 out of 12 sensors compromised on average.
New approach deflects adversarial attacks by causing them to resemble target classes.
problem Ongoing cycle of stronger defenses being broken by more advanced attacks.
method Combines three detection mechanisms in Capsule Networks to achieve state-of-the-art performance on both standard and defense-aware attacks. Uses human study to show attacks can no longer be called adversarial.
result Attack images can no longer be called adversarial because they are classified the same way as humans do.
This paper optimizes attacks on reinforcement learning policies, reducing their effectiveness.
problem Optimizing adversarial attacks on reinforcement learning policies to minimize rewards.
method Designing optimal attacks for both white-box and black-box scenarios using Markov Decision Processes and Reinforcement Learning.
result Optimal attacks can reduce the effectiveness of reinforcement learning policies, especially for smooth policies.
New defense method against physical attacks on image classification models.
problem Defending against physically realizable attacks on image classification models.
method Proposed a new abstract adversarial model, rectangular occlusion attacks, and developed two approaches for efficiently computing adversarial examples.
result Adversarial training using the new attack yields robust image classification models against physical attacks.
A structured approach to generating adversarial attacks for ML systems.
problem Vulnerability of ML systems to adversarial perturbations.
method Developed an 'attack generator' to systematically create adversarial attacks.
result Summarized and extended existing adversarial perturbation taxonomies.
RayS attack improves hard-label adversarial attacks by reducing query complexity and identifying false robust models.
problem Challenges in hard-label adversarial attacks, especially in terms of effectiveness and efficiency.
method Reformulates continuous problem into discrete problem without gradient estimation and uses a fast check step to eliminate unnecessary searches.
result Significantly reduces the number of queries needed for hard-label attacks and identifies false robust models.
New attacks on RL agents' action space improve understanding of cyber-physical systems vulnerabilities.
problem Understanding and improving the robustness of RL agents in CPS against action space attacks.
method Proposed white-box MAS and LAS attack algorithms to optimize and temporally couple attack budgets.
result LAS attacks cause significantly more performance degradation than MAS attacks.
Depending on how much information an adversary can access to, adversarial attacks can be classified as white-box attack and black-box attack. For white-box attack, optimization-based attack algorithms such as projected gradient descent (PGD) can achieve relatively high attack success rates within moderate iterates. How…
New method handles uncertainty in adversarial attacks using ensemble noise simulation.
problem Uncertainty in adversarial attacks on neural networks.
method Simulates attacker's noisy perturbation using various gradient-based attack algorithms and a pre-processing Denoising Autoencoder (DAE) defense.
result Significant improvements in post-attack accuracy with the proposed ensemble-trained defense.
Paper creates universal adversarial attacks.
problem Creating universal, transferable, and targeted adversarial attacks.
method Learn a universal mapping to map sources to adversarial examples.
result Examples can fool networks into classifying all into one targeted class and have strong transferability.
New attacks reveal membership in label-only ML models.
problem Vulnerability of ML models to membership inference attacks.
method Developed decision-based membership inference attacks.
result Label-only exposures are vulnerable to membership leakage.
Defends against ML inference attacks using adversarial examples.
problem Automated inference attacks using ML classifiers pose privacy and security threats.
method Turns ML classifier vulnerabilities into defenses by adding adversarial noise to public data.
result Adversarial examples can mislead ML classifiers and protect private data.
This paper proposes multi-view attack strategies for deep models.
problem Vulnerability of multi-view deep models to adversarial attacks.
method Two multi-view attack strategies: two-stage attack (TSA) and end-to-end attack (ETEA).
result Proposed multi-view attack strategies are effective on multi-view deep models.
Adversarial attacks for image classification are small perturbations to images that are designed to cause misclassification by a model. Adversarial attacks formally correspond to an optimization problem: find a minimum norm image perturbation, constrained to cause misclassification. A number of effective attacks have b…
New attacks can infer model training membership using only label predictions, not confidence.
problem Inferring whether a data point was used to train a machine learning model.
method Evaluate model's predicted labels under perturbations to infer membership.
result Label-only attacks perform as well as confidence-based attacks and break defenses that rely on confidence masking.
Efficiently attacks large-scale graphs without using the whole graph.
problem Vulnerability of graph neural networks to adversarial attacks.
method Simplified Gradient-based Attack (SGA) method for large-scale graphs.
result SGA achieves significant time and memory efficiency improvements.
Pixle attacks images by rearranging pixels, bypassing neural networks.
problem Vulnerability of neural networks to black-box adversarial attacks.
method A novel attack that rearranges a small number of pixels in images.
result Successfully attacks a high percentage of samples on various datasets and models.
Adversarial training makes models more vulnerable to privacy attacks.
problem Privacy attacks on robust models become feasible.
method Demonstrated through model inversion attacks on robustly trained models.
result Privacy attacks on robust models are now feasible.
We introduce two tactics to attack agents trained by deep reinforcement learning algorithms using adversarial examples, namely the strategically-timed attack and the enchanting attack. In the strategically-timed attack, the adversary aims at minimizing the agent's reward by only attacking the agent at a small subset of…
A new method uses adversarial attacks to detect other adversarial attacks.
problem Detecting iterative adversarial attacks on deep neural networks.
method Using the Carlini-Wagner (CW) attack as a detector itself, under certain assumptions.
result The method provides asymptotically optimal separation of original and attacked images.
Paper proposes a black-box adversarial attack method for graph embedding models.
problem Robustness of graph embedding models against adversarial attacks.
method GF-Attack constructs a generalized adversarial attacker by the graph filter and feature matrix, performing the attack on the graph filter in a black-box fashion.
result GF-Attack can consistently make strong attacks on different graph embedding models even with small perturbations.
Study on reward poisoning attacks on CMAB, revealing attackability depends on adversary's knowledge.
problem Reward poisoning attacks on Combinatorial Multi-Armed Bandits (CMAB).
method Provided a sufficient and necessary condition for attackability, devised an attack algorithm.
result Attackability of CMAB depends on adversary's knowledge of the bandit instance.
Many machine learning algorithms are vulnerable to almost imperceptible perturbations of their inputs. So far it was unclear how much risk adversarial perturbations carry for the safety of real-world machine learning applications because most methods used to generate such perturbations rely either on detailed model inf…