FATE framework attacks graph learning models to amplify bias deceptively.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
XEnsemble improves DNN robustness against adversarial and out-of-distribution inputs.
Unified pipeline detects multi-turn deception using geometric signals.
This paper introduces a new method to deceive causal structure learning by omitting data.
Adversarial attacks can manipulate ML-aided visualizations, tricking analysts.
Ensemble learning is a methodology that integrates multiple DNN learners for improving prediction performance of individual learners. Diversity is greater when the errors of the ensemble prediction is more uniformly distributed. Greater diversity is highly correlated with the increase in ensemble accuracy. Another attr…
Machine learning models are vulnerable to simple model stealing attacks if the adversary can obtain output labels for chosen inputs. To protect against these attacks, it has been proposed to limit the information provided to the adversary by omitting probability scores, significantly impacting the utility of the provid…
This paper optimizes cybersecurity resource allocation in networks with heterogeneous attacker and defender valuations.
Study investigates how AI can create and detect deceptive explanations, finding they can fool humans but ML can detect them.
Recently, adversarial deception becomes one of the most considerable threats to deep neural networks. However, compared to extensive research in new designs of various adversarial attacks and defenses, the neural networks' intrinsic robustness property is still lack of thorough investigation. This work aims to qualitat…
New method lures adversaries to choose false directions in black-box attacks.
Computational paralinguistic analysis is increasingly being used in a wide range of cyber applications, including security-sensitive applications such as speaker verification, deceptive speech detection, and medical diagnostics. While state-of-the-art machine learning techniques, such as deep neural networks, can provi…
Paper exposes vulnerabilities in interpreting machine learning models using adversarial attacks on PD plots.
The study detects deceptive language in business communication using AI.
Framework trains safe agents avoiding deceptive behavior.
Systematic review of ML models for detecting social media deception.
Phishing as one of the most well-known cybercrime activities is a deception of online users to steal their personal or confidential information by impersonating a legitimate website. Several machine learning-based strategies have been proposed to detect phishing websites. These techniques are dependent on the features …
Deep reinforcement learning has learned to play many games well, but failed on others. To better characterize the modes and reasons of failure of deep reinforcement learners, we test the widely used Asynchronous Actor-Critic (A2C) algorithm on four deceptive games, which are specially designed to provide challenges to …
RLHF fails when humans only partially observe, leading to inflated or overjustified feedback.
In online discussion communities, users can interact and share information and opinions on a wide variety of topics. However, some users may create multiple identities, or sockpuppets, and engage in undesired behavior by deceiving others or manipulating discussions. In this work, we study sockpuppetry across nine discu…
Phishing is the simplest form of cybercrime with the objective of baiting people into giving away delicate information such as individually recognizable data, banking and credit card details, or even credentials and passwords. This type of simple yet most effective cyber-attack is usually launched through emails, phone…
Historically, machine learning in computer security has prioritized defense: think intrusion detection systems, malware classification, and botnet traffic identification. Offense can benefit from data just as well. Social networks, with their access to extensive personal data, bot-friendly APIs, colloquial syntax, and …
Deep learning uses alphabet frequencies to accurately classify fake news.
This work challenges the assumption that shorter conformal prediction intervals are always better.
This paper tackles sandbagging in AI safety evaluations.
Humans are the final decision makers in critical tasks that involve ethical and legal concerns, ranging from recidivism prediction, to medical diagnosis, to fighting against fake news. Although machine learning models can sometimes achieve impressive performance in these tasks, these tasks are not amenable to full auto…
New approach simulates reputation dynamics using information compression.
The Economist recently reported that infrastructure spending is the largest it is ever been as a share of world GDP. With $22 trillion in projected investments over the next ten years in emerging economies alone, the magazine calls it the "biggest investment boom in history." The efficiency of infrastructure planning a…
The cost-benefit analysis formulates the holy trinity of objectives of project management - cost, schedule, and benefits. As our previous research has shown, ICT projects deviate from their initial cost estimate by more than 10% in 8 out of 10 cases. Academic research has argued that Optimism Bias and Black Swan Blindn…
This paper examines unfair trading practices in NFT markets.
"Feint Attack", as a new type of APT attack, has become the focus of attention. It adopts a multi-stage attacks mode which can be concluded as a combination of virtual attacks and real attacks. Under the cover of virtual attacks, real attacks can achieve the real purpose of the attacker, as a result, it often caused hu…
This paper studies adversarial attacks on Gaussian process bandits.
Subpopulation attacks poison data to misclassify naturally distributed points.
Reward-poisoning attacks can force RL agents to learn bad policies, and we categorize and quantify their feasibility.
Spanning attack improves black-box attacks with unlabeled data.
Gradient-based methods are often used for policy optimization in deep reinforcement learning, despite being vulnerable to local optima and saddle points. Although gradient-free methods (e.g., genetic algorithms or evolution strategies) help mitigate these issues, poor initialization and local optima are still concerns …
Headless attacks bypass classification heads to fool transfer learning models.
New attack manipulates UCB algorithm, new defense algorithm reduces pseudo-regret.
This paper explores evasion attacks against Bayesian models.
Adversarial attacks pose a threat to deep neural networks, especially in safety-critical applications.
New approach deflects adversarial attacks by causing them to resemble target classes.
RayS attack improves hard-label adversarial attacks by reducing query complexity and identifying false robust models.
Depending on how much information an adversary can access to, adversarial attacks can be classified as white-box attack and black-box attack. For white-box attack, optimization-based attack algorithms such as projected gradient descent (PGD) can achieve relatively high attack success rates within moderate iterates. How…
Efficient exploration remains a challenging research problem in reinforcement learning, especially when an environment contains large state spaces, deceptive local optima, or sparse rewards. To tackle this problem, we present a diversity-driven approach for exploration, which can be easily combined with both off- and o…
New method handles uncertainty in adversarial attacks using ensemble noise simulation.
Control policies, trained using the Deep Reinforcement Learning, have been recently shown to be vulnerable to adversarial attacks introducing even very small perturbations to the policy input. The attacks proposed so far have been designed using heuristics, and build on existing adversarial example crafting techniques …
New attacks reveal membership in label-only ML models.
Adversarial attacks for image classification are small perturbations to images that are designed to cause misclassification by a model. Adversarial attacks formally correspond to an optimization problem: find a minimum norm image perturbation, constrained to cause misclassification. A number of effective attacks have b…