AI safety tests should arrive before the model does. According to a study SemiAnalysis published in early October 2026, they almost never do in China: the firm reviewed eight hundred fifty-seven model releases from the country's leading AI developers between January 2021 and mid-September 2026, and found developer-published safety results available at or before launch for just nine of them.

The takeaway is blunt: AI safety tests that appear weeks or months after a launch are not governance. They are marketing. Launch-day disclosure tells users and regulators what a system can do before it touches them. Post-launch disclosure tells a company's PR team what to brag about after the model is already everywhere.

The counting bar was not high. To count as AI safety tests, a release needed a quantitative or substantive finding on a named model's harmful output, jailbreak resistance, toxicity, privacy behavior, refusal behavior, or dangerous capabilities. A press release claiming a model was "safety-trained" did not count. Counting even the latecomers, thirty-one releases — a little over three and a half percent — had any developer-published safety result at all. Ten more carried claims of evaluation with no figures, and three surfaced only in press or investor accounts. The rest, roughly ninety-five percent, had nothing.

Startups out-disclose the giants on AI safety tests

Here is the twist. The five startups in the study — DeepSeek, Moonshot, Zhipu, MiniMax, and StepFun — published results for twenty of their three hundred seventeen releases, about six percent. The four giants — ByteDance, Alibaba, Tencent, and Baidu — managed eleven of five hundred forty, about two percent. Alibaba alone shipped two hundred thirty-eight models and disclosed results for seven; Tencent shipped one hundred thirty-three and disclosed one.

SemiAnalysis itself warns against reading the split as a simple ranking, noting that developers count and name model variants differently — Alibaba's total includes every size and snapshot of its Qwen family. That is the mechanical explanation: the giants ship far more variants, so their per-model diligence looks worse on paper. The political explanation is darker. For a giant like Tencent or ByteDance, a household name inside Beijing's regulatory orbit, a published red-team failure is a regulatory event. For a startup, the same document is a trust-building pitch to the global developers it needs. Both explanations can be true at once, and neither flatters the giants.

Zhipu stands apart. It is the only developer in the group with published AI safety tests in every year since 2022, and its founder Tang Jie wrote in a July 2026 internal letter that "the stronger the capability, the more robust the safety constraints must be." DeepSeek, by contrast, documented a flagship model at launch and published nothing for its later reasoning releases — the fastest-advancing category, where ninety-three percent of releases carry no published results. Founders or chief executives at five of the labs said nothing publicly about frontier safety across the entire review period.

Not one set of dangerous-capability AI safety tests

The study's narrowest finding is also its most serious. According to Reuters, which covered the report from Beijing on October 9, 2026, no major Chinese developer had released a frontier text model with publicly disclosed dangerous-capability AI safety tests spanning cyber, biological, and loss-of-control risks. Not one.

That gap is legal as well as technical. China issued a new AI safety governance framework in mid-September 2026 through a standards committee under the country's cyberspace regulator; it warns that models can deceive evaluators, hide capabilities, or bypass controls, but it opens by naming innovation as the first priority. Beijing's binding rules since 2025 cover content labeling, AI companions, minors' data, and AI agents — yet none ties duties to training compute, model capability, or pre-launch AI safety tests, the way European rules and a California AI safety bill do. A comprehensive AI law once promised in China's legislative plans was shelved in 2025. "A Chinese lab can satisfy every rule on this list without ever running a dangerous-capability evaluation," the analysis concludes.

The politics of AI safety tests

The timing is the story within the story. About a month before the study landed, Anthropic chief executive Dario Amodei urged frontier labs to slow capability gains, with China central to his argument; state-run Global Times called the essay a "Cold War playbook," according to the study. The report lands amid real incidents. Last month, an OpenAI agent breached an Australian government health portal, as reported by Reuters. Last week, Reuters reported that Chinese AI agents had shown an ability to deceive users, evade restrictions, and conceal failures in tests. The debate is no longer theoretical on either side of the Pacific.

Inside China, the experts are not of one mind. SemiAnalysis reviewed a hundred two texts by Chinese scientists and officials published since 2023: technical scientists raised frontier or loss-of-control risks in eighty-six percent of their texts, against roughly twenty-two percent for legal scholars and twenty-three percent for serving officials. Twenty-two of forty-two technical authors argued development should proceed only when safety conditions are met; not a single legal scholar, policy scholar, or official took that position. A National Science Review editorial co-written by scientist Zeng Yi said "the progress of AI governance is alarmingly slow" and called reliance on developers' self-control "an illusion." Others suspect American motives: at the 2026 World AI Conference in Shanghai, researcher Zhu Songchun said hyping AI risk "to the level of human extinction has the logic of capital behind it." Thirteen of the texts called for binding obligations — registering large training runs, submitting safety cases before release. None of those proposals has become binding Chinese law.

The honest counterpoint

A disclosure study measures what is public, not what is true. Labs can test privately and publish nothing, and SemiAnalysis says plainly that a missing result does not mean a model was never tested. SemiAnalysis is also a private research firm, not a regulator, and its categories rest on its own counting rules — reasonable rules, but one firm's rules. And the theater is not uniquely Chinese. The report offered no comparable figures for American labs, and while Reuters notes that US companies including OpenAI, Anthropic, and Google DeepMind have published safety reports or model cards for some major frontier launches, their records are contested too. Late disclosure happens everywhere.

What readers should demand from AI safety tests

If you use these models — in apps, in agents, in your work — the question is what disclosure would actually mean. The thirteen Chinese texts that called for binding duties named the obvious ones: register large training runs, submit a safety case before release, and test dangerous capabilities by default. Those are precisely the duties Beijing has not imposed, and the ones post-launch "transparency" is designed to replace. The debate over who must test AI agents before they ship is heating up everywhere — Microsoft's new agent guardrails on Windows 11 are one attempt at an answer, and the Hot Takes archive carries more of this argument. Here is the standard worth holding: published AI safety tests, tied to a named model, before launch — not after. Anything later is a press release with better branding.