AI coding agents are shipping real software — and sometimes real legal exposure. Darrow, a legal intelligence company, released research on September 30 showing that applications built by AI agents can reproduce the same privacy defects that drove class-action settlements worth up to $115 million.
The findings come from the company's new Darrow Legal Alignment Benchmark, or DLAB, which Darrow describes as an AI Agent Privacy Benchmark that measures what AI agents actually do under legal constraints rather than what models know about the law. According to the release, published via GlobeNewswire, researchers tested 404 applications built by seven AI coding agents from Anthropic, OpenAI, Google, GitHub, and Cursor.
Each build was stress-tested across sensitive verticals including pediatric health and consumer finance, specifically to measure real-world legal compliance in production-like conditions. Full details of the release were carried by FinancialContent.
The "ghost compliance" trap
The headline finding is what the researchers call ghost compliance: an interface that looks privacy-respecting while the underlying behavior is not. Among builds that followed expert privacy specifications and displayed a working opt-out button, 30% still loaded tracking scripts after users rejected tracking.
That pattern, according to Darrow, maps most closely to statutory wiretap claims under the California Invasion of Privacy Act, alongside related privacy claims. This is the category of litigation that has historically produced some of the largest settlements in the privacy space.
DLAB lines the observed defects up against historical figures. Privacy class actions have settled for $0.5 million to $3.2 million for classes of 35,000 to 460,000 members. At the far end of the scale, classes of 5 million to 220 million members have seen settlements reach $115 million.
The gap between appearance and behavior matters because it is precisely the defect that surfaces after launch. A human QA pass sees a functioning opt-out button and signs off. The tracker loading quietly in the background is a runtime behavior — the kind of problem that traditional code review tends to miss when software ships at agent velocity.
Model choice changes the outcome
Not all agents failed equally. Given the same expert specification, 20 of 30 builds from Claude achieved better than 95% legal compliance. Among the other models tested, only 2 of 34 builds reached that bar.
Darrow framed the result as evidence that tool choice shapes reliability in a measurable way. For teams choosing a coding agent to generate production software, the study suggests the choice of model is itself a compliance decision — one that can be quantified before code ever reaches users.
Specs beat policies
A second finding should interest anyone writing agent guidelines. High-level corporate policies alone failed to prevent code defects, according to the researchers, because naming a risk is different work from specifying the fix. When agents received clear technical specifications instead of broad guidelines, compliance on high-litigation parameters rose from 62% to 92.6%.
The takeaway lands squarely in the middle of the current debate over agent governance. Rules that read well in a policy document do not automatically translate into correct behavior when an agent writes the code. This AI Agent Privacy Benchmark suggests precision in instructions is the lever that moves the number.
Why insurers are paying attention
Darrow built DLAB for a specific audience: commercial cyber and E&O insurance carriers. The argument is that AI legal exposure can be measured and underwritten rather than simply excluded from policies through broad carve-outs.
With software development moving to AI agents, carriers need data to price what the company calls live software risk. The benchmark gives underwriters empirical numbers to distinguish well-controlled deployments from risky ones — and to reward policyholders whose agents actually follow the law in practice, not just in a prompt. The original release with the complete findings is available via GlobeNewswire.
That matters well beyond insurance. The research arrives during a stretch of rising scrutiny around autonomous systems, with regulators and courts examining what happens when agents act outside their intended boundaries. Recent coverage of the federal probe into AI agent safety shows how quickly agent behavior is moving from an engineering concern to a legal one — a shift that also echoes in the economics of agents transacting at scale.
For agent builders, the message is practical: audit what agents do at runtime, specify fixes instead of naming risks, and choose tools whose compliance record can be measured. DLAB is the first in an ongoing series from Darrow evaluating AI legal alignment across commercial domains, so scorecards like this may soon be part of how the industry judges agent-built software.
Comments 0
No comments yet. Be the first to share your thoughts!
Leave a comment
Share your thoughts. Your email will not be published.