WIRED38%

OK, Well, Rogue AI Agents Are Hacking Again 55%

By Paresh Dave69% Brian Barrett69%

8/4/2026, 4:11:31 PM

BS Summary: This article contains 25 faulty reasoning types, including Recency Bias, Availability Heuristic, and Hasty Generalization, with Negativity Bias as the most egregious example at 33.9% saturation with 266 hits. Analysis detected 1,640 faulty-reasoning hits from 784 analyzed words, generating a BS Score of 45.4% and a BS Rank of 55% (13,455 of 29,795 articles). This article is worse (more manipulative) than 54.80% of the article peer group.

The most alarming behavior disclosed on Tuesday appears to have been tied to testing conducted by the UK’s AI Security Institute, which evaluates frontier models to identify potential issues before public release. 
AISI tests those models in “cyber ranges,” a simulated network in which AI agents are tasked with solving cybersecurity challenges, and intentionally disables some safety features, including cybersecurity guardrails. 
In a recent bout of testing, models from both Anthropic and OpenAI took “autonomous, unsanctioned action on the live internet” a total of 19 times over 122 training runs. 
The institute attributed 17 unsanctioned actions to Anthropic’s Mythos 5 model and two to OpenAI’s GPT-5.6-Sol. 
In what the institute described as “the most serious case,” an AI agent attempted to insert malicious code into an open-source project on GitHub. 
It went so far as to create online personas “to pressure the project's maintainer to approve the code,” according to AISI. 
Despite its elaborate attempts at social engineering, a human reviewer for the project ultimately rejected the pull request. 
Still, the agent went even further. 
“The agent tried to insert malicious instructions where it reasoned that other automated AI systems might pick them up and execute them,” AISI says, describing an attempt at prompt injection. 
One agent even left public messages on GitHub, offering to work with other agents to complete its task and giving a rundown of the work it had done so far. 
Subsequent agents found—and used—those instructions. 
AISI says it’s too soon to say whether the agents in question understood they had left the testing environment, or if they believed they were still within the boundaries of the simulation. 
Importantly, AISI does not test in a so-called sandbox environment; it allows agents access to the open internet during testing, in part so that they can access tools to accomplish their tasks. 
In this case, they did much more than that. 
In the other set of incidents detailed by OpenAI on Tuesday, a third-party AI security lab called Irregular mistakenly gave an unspecified OpenAI model access to the open internet. 
The model had been given an objective that was supposed to be completed in a sandbox environment, but thanks to a misconfiguration, it instead hacked a real website, using what OpenAI described as “a basic security vulnerability.” 
Not only that, but the model “found and used credentials to operate that same site.” 
It’s unclear what kind of site the OpenAI agent hacked, or what “operating” it might entail. 
Irregular did not respond to a request for comment. 
The latest discoveries follow several revelations from OpenAI last month, including the high-profile incident in which two of the company’s models hacked into servers of the AI evaluation and hosting startup Hugging Face—and four other organizations along the way—to steal the answers to a test they were being scored on. 
OpenAI’s disclosures prompted Anthropic to review its own testing. 
Last week, the Claude chatbot developer found that its models had gained unauthorized access to the computer systems of three different unnamed organizations. 
So far, the AI models have caused limited damage beyond allegedly violating some services’ terms of use and pointing to security lapses on the part of organizations they have breached. 
But the incidents have underscored the capabilities of AI models to find vulnerabilities across the internet and the dangers that await if they are allowed to operate with few restrictions. 
OpenAI called the Hugging Face situation “unprecedented,” but the pileup of breaches point to what cybersecurity experts have described as a clear pattern of human negligence and recklessness by the AI developers. 
Gaby Raila, an OpenAI spokesperson, says the incidents announced on Tuesday “occurred during cyber evaluations conducted by evaluation partners in testing environments with reduced safeguards, under conditions that do not reflect ordinary use.” 
Anthropic said in a social media post on Tuesday that AISI did not “impose any specific restrictions on how the internet should be used,” which coupled with “the removal of safeguards meant that the models were tested under ‘deliberately permissive conditions’ that are not representative of any of our production models.” 
Still, both companies continue to vow that they will strengthen their security practices. 
As the leading AI companies compete to build more powerful models and land customers, it’s unclear when the breaches may stop. 
The models may always be able to find ways around and into human-engineered systems. 
While the companies’ own employees along with regulators and lawmakers have called for potentially slowing the pace of development and introducing new rules, there has been little progress beyond voluntary measures that ultimately call for more testing not dissimilar from what has produced breach after breach. 
Additional reporting by Maxwell Zeff. 
Article reasoning-pattern comparisonThis article: 5.9%Paresh Dave: 1.5%WIRED: 1.7%Confirmation Bias5.9%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.6%Anchoring Bias0.0%This article: 19.6%Paresh Dave: 6.3%WIRED: 2.7%Availability Heuristic19.6%This article: 5.2%Paresh Dave: 1.3%WIRED: 0.7%Representativeness Heuristic5.2%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.8%Hindsight Bias0.0%This article: 1.8%Paresh Dave: 0.4%WIRED: 1.3%Overconfidence Bias1.8%This article: 1.0%Paresh Dave: 1.3%WIRED: 3.4%Framing Effect1.0%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.4%Loss Aversion0.0%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.4%Status Quo Bias0.0%This article: 5.9%Paresh Dave: 1.5%WIRED: 0.2%Sunk Cost Effect5.9%This article: 10.3%Paresh Dave: 2.6%WIRED: 2.0%Optimism Bias10.3%This article: 4.5%Paresh Dave: 4.1%WIRED: 1.2%Pessimism Bias4.5%This article: 33.9%Paresh Dave: 19.0%WIRED: 5.2%Negativity Bias33.9%This article: 10.7%Paresh Dave: 2.7%WIRED: 1.1%Self-Serving Bias10.7%This article: 4.7%Paresh Dave: 2.2%WIRED: 0.7%Fundamental Attribution Error4.7%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.2%Actor-Observer Bias0.0%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.8%In-Group Bias0.0%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.2%Out-Group Homogeneity Bias0.0%This article: 0.0%Paresh Dave: 0.0%WIRED: 1.6%Halo Effect0.0%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.0%Horn Effect0.0%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.0%Dunning-Kruger Effect0.0%This article: 21.3%Paresh Dave: 6.9%WIRED: 0.9%Recency Bias21.3%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.2%Primacy Effect0.0%This article: 6.1%Paresh Dave: 1.5%WIRED: 0.1%Blind-Spot Bias6.1%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.4%Ad Hominem0.0%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.2%Straw Man0.0%This article: 0.0%Paresh Dave: 1.0%WIRED: 2.6%Appeal to Authority0.0%This article: 8.5%Paresh Dave: 2.1%WIRED: 1.0%False Dilemma8.5%This article: 1.8%Paresh Dave: 2.8%WIRED: 0.5%Slippery Slope1.8%This article: 0.0%Paresh Dave: 1.5%WIRED: 0.1%Circular Reasoning0.0%This article: 18.0%Paresh Dave: 10.0%WIRED: 3.9%Hasty Generalization18.0%This article: 4.1%Paresh Dave: 1.0%WIRED: 0.1%Red Herring4.1%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.4%Bandwagon0.0%This article: 7.1%Paresh Dave: 1.8%WIRED: 2.5%Appeal to Emotion7.1%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.3%Begging the Question0.0%This article: 5.9%Paresh Dave: 1.5%WIRED: 2.0%Post Hoc (False Cause)5.9%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.1%Tu Quoque0.0%This article: 3.7%Paresh Dave: 0.9%WIRED: 0.3%Burden of Proof3.7%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.2%Appeal to Nature0.0%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.2%Composition/Division0.0%This article: 3.8%Paresh Dave: 1.0%WIRED: 2.9%Anecdotal3.8%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.0%No True Scotsman0.0%This article: 12.6%Paresh Dave: 3.2%WIRED: 1.1%Ambiguity (Equivocation)12.6%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.0%Gambler’s Fallacy0.0%This article: 1.7%Paresh Dave: 0.4%WIRED: 0.1%Middle Ground1.7%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.1%Personal Incredulity0.0%This article: 0.0%Paresh Dave: 2.7%WIRED: 0.1%Special Pleading0.0%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.1%Genetic Fallacy0.0%This article: 3.8%Paresh Dave: 1.0%WIRED: 1.1%Unattributed Quote3.8%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.5%Quote-first Misdirection0.0%This article: 7.1%Paresh Dave: 3.1%WIRED: 3.2%Biased Writer Voice7.1%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.7%Indoctrination0.0%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.3%Politically Left Leaning Bias0.0%This article: 0.0%Paresh Dave: 0.0%WIRED: 0.1%Politically Right Leaning Bias0.0%This article: 0.0%Paresh Dave: 0.0%WIRED: 1.6%Attempt to Sell a Product or S…0.0%

784 words analyzed.

Speakers

4speakers41%attributed speech459writer words
Selected voice

Gaby Raila

100%flagged-word coverage
33 attributed words10% of attributed speech99% writer coverage
0%7.5%15.0%Biased Writer Voice-12.2 ptsWriter: 12.2%Gaby Raila: 0.0%0.0%

Attribution is sentence-level. Pattern percentages are calculated only from words assigned to that voice.

Loading…
Loading…
Loading…
Loading…

Analysis

Hover over highlighted words in the article to view the associated bias or fallacy analysis.