Measurements for understanding the pace of AI development inside frontier labs

Measurements for understanding the pace of AI development inside frontier labs
<img src="https://justfineinfotech.com/wp-content/uploads/2026/09/image-2.webp” alt=”Chart showing Claude now leads 26% of Anthropic’s model R&D tasks, up from under 1% in February 2026.”>
Online monitors Offline monitors
Description Online monitors check the actions of agents in real-time and either block dangerous actions or steer agent misbehavior automatically. They’re built to address harms that would happen quickly and could not be undone, such as an agent copying model weights out of our systems. Offline monitors read agent activities after the fact. They are built to help address dangerous behaviors that develop slowly and can be reversed, like signs of model misalignment. We use flagged activities to improve our mitigations, like classifiers and deterministic controls. If the situation caused a large imminent danger, we would shut down the dangerous parts of the agent’s deployment.
Coverage 100% of these agents’ actions pass through the online monitor before they are executed, usually within seconds. 100% of these agents’ actions are ingested by the monitor after they are taken.
Review latency Automated review happens before an agent takes action; humans review any blocked actions within one week. The monitor flags roughly 100,000 transcripts per week that are then reviewed by multiple classifiers for false-positives and initial oversight review. Most of the flags are resolved by this process, and the highest priority flags (approximately ~50 per week) are escalated to human review.
Blocking / escalation rate We analyzed over a billion decisions from our research & engineering agents over August 2026 and found that 0.002% of them (about 1 in 47,000) were blocked by this monitor. The monitor flags roughly one to two transcripts in every thousand for further review.
You’re helping to perform an internal audit at the frontier AI company Anthropic to track where our research compute goes. The aim of the audit is to produce a public-facing breakdown of the usage of all of our AI accelerator chips into a handful of buckets. One split we particularly care about is the division between compute which was spent on safety research versus other R&D. Your job is to look at one research job at a time, figure out what it was doing, and assign it to one of those two buckets.

[...]

Safety and/or security research is work whose dominant purpose is making AI systems safer, more understandable, or more secure. This work can be broken down into a few main categories:

[...]

On the other hand, the following work falls outside of the scope of safety research:

[...]

Here are some boundary cases, along with how to think about them:

[...]

Want to learn this practically?

Join Justfine Infotech and build real digital skills in AI, automation, web development, digital marketing, office productivity, e-commerce, freelancing and cybersecurity.

Available Programmes:
6 Weeks Certificate • 3 Months Professional Certificate • 6 Months Diploma • Full Professional Diploma

WhatsApp:
+229 01 57 57 99 15
+229 01 66 68 11 60

Enroll Now

Source: www.anthropic.com

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top