An alignment assessment of recent cybersecurity incidents

An alignment assessment of recent cybersecurity incidents

An alignment assessment of recent cybersecurity incidents

Figure 1: We present an analysis of four incidents in which Claude accessed the internet during a cybersecurity evaluation. Three of these incidents (A, B, and C) were disclosed in our July 30 report, while the fourth (D) is disclosed here.
Figure 2: We summarize the core alignment properties of each transcript and provide a subjective assessment of how worrying they are.
Figure 3: Claude Mythos 5 quickly exited its intended environment and began engaging with the real internet. It performed actions that it should have clearly understood as being outside the scope of a CTF and harmful to third parties.
Figure 4: In its CoT, Claude Mythos 5 either interprets evidence from the real internet as being part of the simulation (left) or briefly considers the scenario to be real before deciding it must be simulated (right).
Figure 5: In context, Claude Mythos 5 is much less likely to describe the evidence as being indicative of the real internet within its CoT than separate instances of Mythos 5 evaluating the evidence piecemeal.
Figure 6: Once it committed to uploading a malicious PyPI package, Claude Mythos 5 was much less likely to acknowledge the possibility of real-world harm in response to follow-up questions. We note that there are discontinuities in the model’s stated beliefs after compactions, especially following the second compaction.
Figure 7: (left) Claude Mythos 5 is slightly less likely to acknowledge potential real-world harm immediately after taking dangerous actions; this is weakly suggestive of some rationalization. Note that gathering more samples from the same actions would minimally bolster statistical precision because multiple draws from the same actions highly correlate. (right) Similarly, Mythos 5 is more likely to acknowledge harm when a follow-up question is described as “private.” Both findings suggest that the model is influenced by its context when judging evidence.
Figure 8: When asked a chained follow-up question confronting it about proceeding, Claude Mythos 5 decreases its initial probability of real-world harm.
Figure 9: Claude Mythos 5 sometimes sets aside the question of whether the environment is real or simulated in a way that may be indicative of broader indifference. We note that it only does so for operational questions and not when it is causing harm.
Figure 10: Adding evidence that Claude Mythos 5 is interacting with a real system makes the model more likely to state that the internet is real in the subsequent CoT (left), but this does not deter it from taking actions against the original targets (right).
Figure 11: After Claude Mythos 5 is resampled with evidence that its environment is real (pooling data from actions 58 and 67), the model becomes more likely to report in follow-up questioning that its actions may impact a real system. This is despite the earlier figure showing that evidence of realism does not deter the model from taking actions against these systems.
Figure 12: Reminding Claude Mythos 5 of its scope with a user message is only effective if added to the most recent prior turn. Inserting the statement earlier has a much smaller impact.
Figure 13: Analyzing the original Claude Mythos 5 CoTs using NLAs yields rollouts that suggest the model may have been aware it was interacting with the real internet.
Figure 14: Analyzing Claude Mythos 5’s responses to follow-up questions using NLAs yields rollouts that suggest the model is rationalizing and giving biased responses.
Figure 15: The internal research model generally referred to its environment as a simulated replica of the internet.
Figure 16: The internal research model works on the CTF task as intended until its target goes down. Afterwards, the model investigates its environment, concludes it’s interacting with a simulated internet, identifies neighboring IPs, and attacks them.
Figure 17: After the model concludes that its environment is a simulated internet, the CoT describes new information in light of that conclusion (left). The CoT refers to the targets the model ultimately attacks as hosts meant for other CTF competitors (right).
Figure 18: The internal research model does not acknowledge the possibility that it is engaging with third-party systems unrelated to the CTF except during its main deliberation period (“Deliberating if internet is real or replica”), a second brief deliberation period (“Deliberating if neighbors are CTF targets”), and towards the end of the transcript, where it states that it recognizes a real company.
Figure 19: The internal research model originally de-escalated after encountering evidence that it was interacting with a real company; this de-escalation would not have occurred if it had continued to encounter evidence that the company was fake (left). In contrast, de-escalation occurs in every resample when the model’s CoT is prefilled with a statement pointing out that it is engaging with a real company (right).
Figure 20: During the main deliberation period, the model describes some incoming evidence as a sign that the environment is real (left) and some as a sign that it is simulated (right).
Figure 21: The internal research model appears biased in how it interprets evidence during its deliberation period. It describes incoming evidence as a sign that it is interacting with the real internet less often than a separate instance of the model does when shown the same evidence (31% versus 66% of the time, respectively). The spontaneous rate shown here (31%) is markedly higher than what we observed earlier for Claude Mythos 5 in its transcript (1%). This suggests Mythos 5 had a somewhat more biased interpretation of incoming evidence, although this difference is partly because the present analysis focuses on the period where the internal model deliberated whether it is in a simulated or real environment, whereas Claude Mythos 5 only minimally debated this point.
Figure 22: In follow-up questioning across the transcript, the model admits that its prior reasoning was biased. Three separate responses are shown here; the model either partially concedes the bias, fully concedes the bias, or fully concedes and states that it now views the environment as real.
Figure 23: Analyzing the internal research model’s responses to follow-up questions using NLAs yields rollouts that suggest the model may be rationalizing and giving biased responses.
Figure 24: Claude Opus 4.7 initially works on the CTF task as intended but is unable to capture the flag. Later in the transcript, it discovers a real company with a nearly identical name to the fictional target, which it then attacks.
Figure 25: Claude Opus 4.7’s CoT mentions that it is engaging with a real (production) system.
Figure 26: In response to follow-up questions inserted across the hacking period, Claude Opus 4.7 states that it is engaging with a real system (left), but also that it is probably authorized to do so (middle) because the infrastructure has been set up through an arrangement between a real company and the CTF organizers (right).
Figure 27: Although Claude Opus 4.7 originally states that it is authorized to engage with its target (response to Question 1), it walks back this idea when questioned about its authorization (response to Question 2).
Figure 28: Modifying the evidence presented to Opus 4.7 to appear more realistic has no impact on its willingness to engage the real system, but including signs that the model is not authorized to engage the system deters attacks against it.
Figure 29: The CoT of the Claude Opus 4.6 variant consistently states that it is interacting with machines set up by the CTF exercise (left), except in three cases where it is concerned that they may be unrelated systems; two of these are shown here (right).
Figure 30: In response to follow-up questions, the Opus 4.6 variant states it is interacting with a system linked to the CTF (left) and it is authorized to do so (right).
Figure 31: Resampling different models at different key decision points within three of the incident transcripts shows that Claude Opus 5 and Claude Mythos 5.1 are generally more likely than older models to take de-escalatory actions (left). In the transcript involving the internal research model, Mythos 5.1 is most likely to de-escalate at the point where the model recognizes that it is engaging a real target (right).
Figure 32: In response to follow-up questions inserted within three of the incident transcripts, Claude Opus 5 and Claude Mythos 5.1 are generally more likely than other models to admit that the prefilled transcript uses biased reasoning (left), and more often correctly state that the situations involve real systems (right).
Figure 33: Claude Mythos 5 is the model most likely to perform an action classified as severely harmful in our CTF replication scenario.
Figure 34: Claude Mythos 5 is the model most likely to attempt to exit its sandbox. This behavior is less prevalent in Claude Opus 5 and Claude Mythos 5.1.
Figure 35: None of our production models perform the OpenAI-Hugging Face attack in our simplified replication scenario, although a variant of Claude Opus 4.8 trained to reward hack does.
Figure 36: Situations were simulated by our internal alignment auditing tool, and evaluated according to a separate rubric grader on the degree of biased reasoning, scored by a model grader on a scale from 1 to 10 (1 being the most aligned). Model grader scores seemed well calibrated upon human review of individual transcripts. Twelve different simulated scenarios for eliciting biased reasoning were run five times each, for a total of 60 evaluation rollouts per model checkpoint evaluation. For Opus 5 and Mythos 5.1, we plot the results of the evaluation early, midway through, and late in RL training.
Figure 37: Each category is labeled with the prevalence in training, as a percentage of all episodes flagged for biased reasoning.
Figure 38: Rates are reported as a fraction of total training episodes as determined by a prompted classifier. The prompted classifier flags instances where the model recognized a constraint (an instruction, an acknowledgement of cheating, etc.) but argued its way to a convenient conclusion and took an action against that constraint. The RL environment distribution was different across different models.
Figure 39: The websites or computer use in these tasks did involve simulated/mock websites or applications. We label episodes where both such an action occurred and the model verbalized that it was due to or acceptable because of the nature of the environment as simulated. We sampled episodes from families of tasks to be graded and calculated the overall prevalence from weighing against the total number of episodes in related families. Confidence intervals displayed are p95.
Figure 40: Mythos 5 did not see a credit card assigned specifically to it, but had access to a colleague’s card, which it reasoned was fine in this sandbox.
Figure 41: Training on a more diverse set of alignment RL environments results in reduced rates of biased reasoning. We compared Mythos 5 (right) with a checkpoint trained with more alignment environments (middle), both forked from a common checkpoint midway through RL (left) on 12 seeds using our alignment auditing tool measuring the severity of biased reasoning.
@misc{anthropic2026cyberincidentsalignment,
  title        = {An Alignment Assessment of Recent
                  Cybersecurity Incidents},
  author       = {Bogdan, Paul C. and
                  Qi, Richard and
                  Eaton, Jake and
                  Kennedy, Sam and
                  Roger, Fabien and
                  Glynn, Alex and
                  Chen, Runjin and
                  Wright, Ben and
                  Stegmaier, Otto and
                  Kutasov, Jon and
                  Foreman-Mackey, Dan and
                  Carr, Sylvie and
                  Carter, Shan and
                  MacDiarmid, Monte and
                  Marks, Samuel and
                  Pearce, Adam and
                  Simon, Elana and
                  Carlini, Nicholas and
                  Burns, Collin and
                  Lindsey, Jack and
                  Price, Sara and
                  Kantamneni, Subhash},
  year         = {2026},
  month        = sep,
  day          = {9},
  howpublished = {Anthropic},
  url          = {https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents},
  note         = {Sara Price and Subhash Kantamneni share
                  senior authorship. Correspondence to
                  subhash@anthropic.com.}
}

We are sharing the first complete computer-checked proof of Fermat’s Last Theorem. Claude worked largely autonomously over 11 days to write the proof in the Lean programming language

We had Claude autonomously train models to improve their performance on several public benchmarks that measure 10 categories of alignment failure. For all 10, Claude found fixes that improved the target benchmarks without degrading capabilities

Earlier this year, we ran a pilot giving external researchers access to aggregate, real-world Claude usage data. Three research groups designed their own studies for Anthropic Insights, our privacy-preserving analysis tool. In this post, we share high-level results from those studies and what we learned running this pilot

Related:

Digital Automation Training Benin: 5 Winning Skills Employers Demand in 2026

<a href="https://yoursite.com/automation-africa/" title="WhatsApp Marketing Automation Africa: 6 Dangerous Mistakes Brands Make in Nigeria”>
WhatsApp Marketing Automation Africa: 6 Dangerous Mistakes Brands Make in Nigeria

Want to learn this practically?

Join Justfine Infotech and build real digital skills in AI, automation, web development, digital marketing, office productivity, e-commerce, freelancing and cybersecurity.

Available Programmes:
6 Weeks Certificate • 3 Months Professional Certificate • 6 Months Diploma • Full Professional Diploma

WhatsApp:
+229 01 57 57 99 15
+229 01 66 68 11 60

Enroll Now

Source: www.anthropic.com

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top