New attacks found in adversarial training make defenses vulnerable.
problem Vulnerability of adversarial training to new attacks.
method Analysis of adversarial training effectiveness and blind-spot attacks.
result Adversarial training is susceptible to blind-spot attacks.
This paper explores how to fool ECG diagnosis systems with adversarial ECGs.
problem Vulnerability of DNN-powered ECG diagnosis systems to adversarial attacks.
method Analyzed ECG properties to design effective adversarial attacks under two models.
result Demonstrates weaknesses in DNN-powered ECG diagnosis systems under adversarial attacks.
Agents trained in simulation may make errors in the real world due to mismatches between training and execution environments. These mistakes can be dangerous and difficult to discover because the agent cannot predict them a priori. We propose using oracle feedback to learn a predictive model of these blind spots to red…
CNNs can develop blind spots due to uneven padding in feature maps.
problem Spatial bias in convolutional networks leads to blind spots in certain tasks.
method Identified and analyzed the role of padding in convolutional networks, proposing solutions to mitigate bias.
result Mitigating spatial bias improves model accuracy, especially in tasks like small object detection.
A framework to quantify deployment risk in ML systems, especially for rare states.
problem Under-supported rare states in ML models lead to unreliable performance in unseen data.
method Blind-Spot Mass (B_n(tau)) using Good-Turing unseen-species estimation.
result Identifies and quantifies the risk of under-supported states in ML models.
Explores security challenges of machine learning in real-world systems.
problem Vulnerabilities in machine learning models deployed in safety-critical systems.
method Broadens systems security view of ML vulnerabilities, identifies novel challenges, proposes mitigation suggestions.
result Highlights novel challenges and proposes mitigation strategies for securing ML systems.
New method trains image denoising models without clean reference images.
problem Training deep image denoising models without clean reference images.
method Employing networks with a 'blind spot' in the receptive field.
result Image quality is on par with state-of-the-art neural network denoisers.
The study reveals flaws in pruning criteria and proposes a new assumption for better filter selection.
problem Flaws in existing pruning criteria for CNNs.
method Empirical experiments and Convolutional Weight Distribution Assumption.
result The Convolutional Weight Distribution Assumption improves filter selection in pruning.
Decoding strategies often exclude human-like tokens, creating a detectable gap in generated text.
problem Decoding strategies exclude contextually appropriate but statistically rare tokens, creating a detectable gap in generated text.
method Analysis of 1.8 million texts across 8 language models, 5 decoding strategies, and 53 hyperparameter configurations.
result 8-18% of human-selected tokens fall outside typical truncation boundaries, indicating a detectable gap.
Theoretical model for iterative user discovery in recommender systems.
problem Iterative feedback loops in recommender systems and their biases.
method Theoretical framework to model system evolution and convergence properties.
result Theoretical bounds and convergence properties on user discovery and blind spots.
This paper explains tax policy for crypto assets in a rapidly evolving tech landscape.
problem Rapid technological changes in crypto assets create regulatory and tax policy blind spots.
method Explains principles of crypto assets, their technology, and tax issues.
result Tax policies are lagging behind innovation in blockchain and crypto.
Serial problems can't be efficiently parallelized, affecting machine learning models.
problem Inefficiency of parallelization in inherently serial problems.
method Formalized distinction in complexity theory, demonstrated with diffusion models.
result Diffusion models cannot solve inherently serial problems.
Christoffel function characterizes the corruption a bounded-degree certificate cannot remove in robust halfspace learning.
problem Robust halfspace learning under malicious noise
method Sum-of-Squares degree of outlier-removal certificate
result Christoffel function bounds the corruption a bounded-degree certificate cannot remove
This paper detects multi-stage Feint Attacks using Bi-RNN and few-shot learning.
problem Detecting multi-stage Feint Attacks due to lack of professional datasets and semantic relationships.
method Fuzzy clustering for attack chain mining, few-shot deep learning, Bi-RNN for feature extraction.
result Accurately detected Feint Attacks using Bi-RNN and few-shot learning.
We introduce MosAIc, an interactive web app that allows users to find pairs of semantically related artworks that span different cultures, media, and millennia. To create this application, we introduce Conditional Image Retrieval (CIR) which combines visual similarity search with user supplied filters or "conditions". …
This paper studies adversarial attacks on Gaussian process bandits.
problem Adversarial attacks on Gaussian process bandits to manipulate optimal function regions.
method Proposes various adversarial attack methods on GP bandits, including white-box and black-box attacks.
result Adversarial attacks can force GP bandits to optima in target regions even with low attack budgets.
Subpopulation attacks poison data to misclassify naturally distributed points.
problem Improving accuracy of machine learning predictions through adversarial data modification.
method Introducing a novel subpopulation attack framework, using influence functions and gradient optimization.
result Subpopulation attacks are effective and stealthy, making them difficult to defend against.
Efficient attacks on DRL models without model access and low computation.
problem Vulnerabilities of DRL models to adversarial attacks.
method Adapting black-box attacks, introducing efficient online sequential attacks, exploring perturbations in environment dynamics, and generating robust physical perturbations.
result Demonstrated the effectiveness of proposed attacks on real-world robots.
LemonadeBench evaluates LLMs' economic intuition through a simulated lemonade stand.
problem Evaluating LLMs' economic understanding and decision-making in simple markets.
method Simulated lemonade stand business to test LLMs' long-term planning and profit maximization.
result Models achieve profitability but exhibit local rather than global optimization.
Reward-poisoning attacks can force RL agents to learn bad policies, and we categorize and quantify their feasibility.
problem Reward-poisoning attacks can manipulate RL agents to learn undesirable policies.
method Categorize attacks by infinity-norm constraint, provide thresholds for feasibility, and develop adaptive attack strategies.
result Adaptive reward-poisoning attacks can achieve the nefarious policy in polynomial steps, while non-adaptive attacks require exponential steps.
Spanning attack improves black-box attacks with unlabeled data.
problem Query inefficiency in black-box attacks due to high input space dimensionality.
method Proposes spanning attack by constraining adversarial perturbations in a low-dimensional subspace via an auxiliary unlabeled dataset.
result Significantly improves query efficiency of black-box attacks.
Headless attacks bypass classification heads to fool transfer learning models.
problem Adversarial attacks against transfer learning models without access to the classification head.
method Label-blind adversarial attacks that do not require class-label information.
result Transfer attack lowers ResNet18 accuracy on CIFAR10 by over 40%.
New attack manipulates UCB algorithm, new defense algorithm reduces pseudo-regret.
problem Adversarial attacks on stochastic bandit algorithms.
method Introducing action-manipulation attacks and proposing a robust defense algorithm.
result Proposed defense algorithm reduces pseudo-regret to O(max{log T, A}).
This paper explores evasion attacks against Bayesian models.
problem Bayesian predictive models are vulnerable to evasion attacks.
method Developed gradient-based attacks for specific point predictions and entire posterior distributions.
result Optimal evasion attacks can be designed against Bayesian models.
Adversarial attacks pose a threat to deep neural networks, especially in safety-critical applications.
problem Adversarial attacks can misclassify deep neural networks, leading to safety issues.
method Adversarial attacks are categorized into white-box and black-box attacks based on the attacker's knowledge. They can be targeted or non-targeted.
result Adversarial attacks are effective and can transfer between different models and real-world scenarios.
A new Frank-Wolfe framework improves efficiency and effectiveness of adversarial attacks.
problem Develop efficient and effective optimization-based adversarial attack algorithms.
method Proposes a Frank-Wolfe algorithm variant for both white-box and black-box adversarial attacks.
result Demonstrates improved efficiency and effectiveness compared to existing methods.
A new adversarial attack improves model perturbation efficiency.
problem Improving adversarial attacks to better perturb images.
method LogBarrier method for solving constrained minimization problem.
result LogBarrier attack performs better on challenging images.
Adversarial attacks hide cyber-physical attacks in ICS.
problem Hiding cyber-physical attacks in industrial control systems.
method Modeling an attacker compromising sensors, manipulating data, and evaluating attacks on both continuous and mixed data.
result Successfully hides cyber-physical attacks with 2.87 out of 12 sensors compromised on average.
New approach deflects adversarial attacks by causing them to resemble target classes.
problem Ongoing cycle of stronger defenses being broken by more advanced attacks.
method Combines three detection mechanisms in Capsule Networks to achieve state-of-the-art performance on both standard and defense-aware attacks. Uses human study to show attacks can no longer be called adversarial.
result Attack images can no longer be called adversarial because they are classified the same way as humans do.
This paper optimizes attacks on reinforcement learning policies, reducing their effectiveness.
problem Optimizing adversarial attacks on reinforcement learning policies to minimize rewards.
method Designing optimal attacks for both white-box and black-box scenarios using Markov Decision Processes and Reinforcement Learning.
result Optimal attacks can reduce the effectiveness of reinforcement learning policies, especially for smooth policies.
New defense method against physical attacks on image classification models.
problem Defending against physically realizable attacks on image classification models.
method Proposed a new abstract adversarial model, rectangular occlusion attacks, and developed two approaches for efficiently computing adversarial examples.
result Adversarial training using the new attack yields robust image classification models against physical attacks.
A structured approach to generating adversarial attacks for ML systems.
problem Vulnerability of ML systems to adversarial perturbations.
method Developed an 'attack generator' to systematically create adversarial attacks.
result Summarized and extended existing adversarial perturbation taxonomies.
RayS attack improves hard-label adversarial attacks by reducing query complexity and identifying false robust models.
problem Challenges in hard-label adversarial attacks, especially in terms of effectiveness and efficiency.
method Reformulates continuous problem into discrete problem without gradient estimation and uses a fast check step to eliminate unnecessary searches.
result Significantly reduces the number of queries needed for hard-label attacks and identifies false robust models.
New attacks on RL agents' action space improve understanding of cyber-physical systems vulnerabilities.
problem Understanding and improving the robustness of RL agents in CPS against action space attacks.
method Proposed white-box MAS and LAS attack algorithms to optimize and temporally couple attack budgets.
result LAS attacks cause significantly more performance degradation than MAS attacks.
New method handles uncertainty in adversarial attacks using ensemble noise simulation.
problem Uncertainty in adversarial attacks on neural networks.
method Simulates attacker's noisy perturbation using various gradient-based attack algorithms and a pre-processing Denoising Autoencoder (DAE) defense.
result Significant improvements in post-attack accuracy with the proposed ensemble-trained defense.
This paper investigates why adversarial attacks can be effective against different models.
problem Understanding why adversarial attacks transfer between models.
method Developed a unifying optimization framework and formal definition for attack transferability.
result Identified two main factors: target model's vulnerability and surrogate model complexity.
New attacks reveal membership in label-only ML models.
problem Vulnerability of ML models to membership inference attacks.
method Developed decision-based membership inference attacks.
result Label-only exposures are vulnerable to membership leakage.
Paper creates universal adversarial attacks.
problem Creating universal, transferable, and targeted adversarial attacks.
method Learn a universal mapping to map sources to adversarial examples.
result Examples can fool networks into classifying all into one targeted class and have strong transferability.
Defends against ML inference attacks using adversarial examples.
problem Automated inference attacks using ML classifiers pose privacy and security threats.
method Turns ML classifier vulnerabilities into defenses by adding adversarial noise to public data.
result Adversarial examples can mislead ML classifiers and protect private data.
This paper proposes multi-view attack strategies for deep models.
problem Vulnerability of multi-view deep models to adversarial attacks.
method Two multi-view attack strategies: two-stage attack (TSA) and end-to-end attack (ETEA).
result Proposed multi-view attack strategies are effective on multi-view deep models.
New attacks can infer model training membership using only label predictions, not confidence.
problem Inferring whether a data point was used to train a machine learning model.
method Evaluate model's predicted labels under perturbations to infer membership.
result Label-only attacks perform as well as confidence-based attacks and break defenses that rely on confidence masking.
Paper studies attacks on bandit algorithms and shows how attackers can manipulate data to hijack behavior.
problem Potential attacks on bandit algorithms can cause catastrophic loss in real-world applications.
method Proposes a framework of offline and online attacks on bandit algorithms using convex optimization and adaptive strategies.
result Attackers can force bandit algorithms to pull target arms with high probability by manipulating data.
Efficiently attacks large-scale graphs without using the whole graph.
problem Vulnerability of graph neural networks to adversarial attacks.
method Simplified Gradient-based Attack (SGA) method for large-scale graphs.
result SGA achieves significant time and memory efficiency improvements.
Pixle attacks images by rearranging pixels, bypassing neural networks.
problem Vulnerability of neural networks to black-box adversarial attacks.
method A novel attack that rearranges a small number of pixels in images.
result Successfully attacks a high percentage of samples on various datasets and models.
Adversarial training makes models more vulnerable to privacy attacks.
problem Privacy attacks on robust models become feasible.
method Demonstrated through model inversion attacks on robustly trained models.
result Privacy attacks on robust models are now feasible.
We introduce two tactics to attack agents trained by deep reinforcement learning algorithms using adversarial examples, namely the strategically-timed attack and the enchanting attack. In the strategically-timed attack, the adversary aims at minimizing the agent's reward by only attacking the agent at a small subset of…
A new method uses adversarial attacks to detect other adversarial attacks.
problem Detecting iterative adversarial attacks on deep neural networks.
method Using the Carlini-Wagner (CW) attack as a detector itself, under certain assumptions.
result The method provides asymptotically optimal separation of original and attacked images.
Interval attacks find more adversarial examples than existing methods.
problem Evaluating robustness of adversarially trained neural networks against unknown attacks.
method Symbolic interval propagation for bound over-approximation and gradient-guided attacks.
result Interval attacks find on average 47% more violations than state-of-the-art methods.