Anthropic published something this week that no leading AI lab has published before: a measurement of AI building AI, quantifying how much of its own research is now done by its own AI. According to the Associated Press, the company announced on Thursday that Claude now leads 26 percent of its model research and development, up from less than 1 percent in February. The company introduced an R&D Automation Index to track the shift and urged other labs to publish similar numbers using a common methodology. The Associated Press has the full breakdown of the announcement.

Six months ago, AI directing AI research was a premise for research papers. This week it appeared in a staffing metric. AI building AI is now a line item in how a frontier lab runs itself. Anthropic also said that about 90 percent of the company's research and development now happens in at least "collaboration" with Claude, and that the model is not yet operating fully autonomously in any of the categories it measured. The number is less important than who checked it.

The student graded its own exam

Here is how Anthropic built the index. Each week of July, the company randomly sampled 20 percent of the staff in its model R&D departments and pulled their work records, including Slack messages and internal documents. That produced roughly fifteen thousand tasks, which Claude sorted into 542 categories. Then a separate Claude judge assigned each category an automation level. A detailed breakdown of the methodology shows how the self-measurement was constructed. The scale comes from Epoch AI and runs from AL0, no AI participation, to AL5, fully autonomous with no human in the loop. AL4, the "leads" tier, means the model completes most of a task end to end from a high-level prompt while a human supervises and reviews. This is the first public attempt to quantify AI building AI inside a frontier lab, and the yardstick was built by the lab itself.

The company disclosed all of this openly and acknowledged that errors and differences in interpretation were possible, since Claude performed the analysis and the assessment. Credit where it is due: most labs publish nothing at all. Disclosure is not verification, though, and as of the announcement no outside party has independently checked the numbers. The tool being measured did the measuring. The categories were drawn by the model, and the grades were assigned by the model. A self-report can be honest and still be the weakest kind of evidence. That is how audits work: the auditor and the audited cannot be the same party. Anthropic wrote that it has "been doing this kind of monitoring for months, not years."

The fine print nobody is quoting

Anthropic also broke out compute spending for one week in July. About 6 percent of AI research and development compute was classified as safety work, and among the compute used by AI agents in model development, the safety share was 12 percent. The company called both estimates conservative, noting that mixed-purpose work was counted as capabilities research.

The company also reported that around 30,000 AI agents run simultaneously on its main internal research platform. In August, its monitoring systems checked over a billion of their actions and blocked 0.002 percent. About 50 of the most critical alerts were forwarded to humans for review each week. These agent-monitoring figures were reported alongside the index. Every one of these figures comes from the company's own reporting, which is precisely the point.

Anthropic warned that models accelerating their own development could make it "more challenging for humans to understand or control these systems." In a blog post, the company wrote that society should minimize the gap between what frontier labs know and what the public knows. AI systems already shape consequential decisions well beyond the lab, from health care to the apps that decide what billions of people watch, and the public conversation keeps filling the vacuum with speculation instead of verified numbers. The index is a genuine attempt to narrow that gap, published by the lab with the most to prove.

Why the audit matters more than the number

The announcement landed as leading figures in AI, including Anthropic chief executive Dario Amodei, have been calling for a slowdown in development over safety concerns, according to the Associated Press. So the same week a lab says the industry should pump the brakes, it also says its AI now leads a quarter of its research. That tension is the story. You can hold both facts at once without resolving them.

Here is the hot take, stated plainly. The 26 percent figure deserves less attention than the audit trail behind it: an industry where the loudest voice for transparency measures the most consequential shift in research history with a ruler it made itself, and nobody else has held the ruler. The AI building AI number will keep climbing. From February to August it went from under 1 percent to 26, and whatever it reads next year, the oversight question will be harder, not easier.

Regulators, journalists, and researchers should treat the R&D Automation Index as the start of a standard, not a result. Anthropic did the field a service by publishing first. Now someone without a stake in the number needs to check it.