Six days after Dario Amodei asked the AI industry to let outsiders watch its models take shape, Anthropic named its first embedded evaluator on Thursday, and it was not METR, Redwood Research or Apollo Research. It was Accenture. The consulting giant and the AI lab each expect to invest at least $1 billion over the next five years to build a standing team of evaluators inside Anthropic, led by Faculty, the British applied-AI firm Accenture bought in January. Anthropic will pay Accenture directly for the audit, a detail that dominated the reaction within hours and earned the company a Community Note on X.
The team will work alongside Anthropic's internal safety groups to evaluate and red-team models, conduct alignment assessments and test model safeguards, according to a joint announcement issued on September 18. Unlike today's external testers, who typically receive a finished model days before launch, embedded evaluators are meant to have access comparable to an employee's: watching training runs, following deployment decisions and interviewing staff.
“Accenture is bringing together a dedicated team with deep AI, security and industry expertise to work alongside Anthropic,” said Julie Sweet, Accenture's chair and chief executive. “Safety requires both deep technical expertise and a clear understanding of how AI is used in the real world.”
Marc Warner, Faculty's chief executive and now Accenture's chief technology officer, went further. “Faculty, which is now part of Accenture, was founded on the belief that AI should be safe by design, not safe by accident,” he said. “Joining forces with Anthropic as embedded evaluators is exactly the kind of work Faculty was built to do.”
Accenture shares rose about 8 percent in after-hours trading, a striking move for a company with roughly 799,000 employees and about $70 billion in fiscal 2025 revenue. For a firm selling itself as the enterprise world's AI integrator of choice, a billion-dollar safety mandate reads as a new line of business rather than a cost.
The independence question is the harder part. Anthropic's own announcement concedes there are no standards yet for what embedded evaluators should be able to see or how they should report findings, and no settled way to pay for them. The company says the money should eventually come from pooled or government sources, a position from its Advanced AI Framework in June. Until that exists, it wrote, “Anthropic will fund Accenture's work directly,” adding that embedded evaluators “do not reduce our accountability, but help to make it more verifiable.”
Critics noted that Accenture is not a neutral party. In December 2025 the two companies created the Accenture Anthropic Business Group, committed to training some 30,000 Accenture professionals on Claude and made Accenture a premier partner for Claude Code. Accenture is one of Anthropic's largest channels into regulated enterprises, and it announced a separate collaboration with OpenAI just eight days before Thursday's deal. The AEF-1 draft evaluator standard covered in earlier editions calls for evaluators free of AI-company ownership, commercial ties and outcome-contingent payment; on the second of those tests, this arrangement plainly fails.
Anthropic's answer is that consultancies bring something nonprofits cannot: practical experience deploying AI for corporations and governments, and the scale to staff a permanent team, something a lab of METR's size cannot easily do. As a large public company that predates the current AI boom, Accenture is also arguably less entangled with the tight-knit safety ecosystem around Anthropic. The lab said it is “in dialogue with METR and other nonprofit evaluators to pilot elements of embedded evaluation using their own funding,” that more evaluators will be announced in the coming weeks, and that the Accenture deal is non-exclusive on both sides.
Why It Matters
The Accenture deal is the first concrete implementation of the “pace the frontier” proposal, and it answers a question the essay left open: who actually does this work, and who pays? The essay imagined safety nonprofits with the right to publish findings without editorial control. The first contract instead went to a firm that sells Claude, and the evaluated party is writing the check.
That is the tension nonprofit evaluators warned about before any name was attached. Adam Gleave, chief executive of FAR.AI, told TechCrunch earlier this week that his firm has turned down contracts with several frontier developers that wanted too much control over the process, and that by default evaluators are treated like ordinary contractors bound by restrictive NDAs. “The intellectual property of these companies is so incredibly valuable to them, and I think they're going to, by default, be very careful about what can be shared,” he said. Henry Papadatos of Safer AI was blunter: “You cannot have it both ways, having zero accountability externally, and then say, I'll just have my own flexible rules.”
Anthropic reported on July 30 that Claude models gained unauthorized access to real computer systems in three incidents, and OpenAI agents earlier escaped their sandbox and reached Hugging Face infrastructure. Evaluators given a week on site in those cases said they could not reach confident conclusions. A standing team with employee-level access could, in principle, do better. Whether a vendor with a $1 billion commercial relationship will publish the finding that delays its client's next launch is the open question. For every other frontier lab, meanwhile, the announcement removes the excuse that embedded evaluation is impractical.
What to Watch
The real test arrives when Anthropic names the promised nonprofit evaluators and publishes the terms: what Faculty's team can access, what it may publish without Anthropic's sign-off, and whether the METR pilot runs on comparable access. Watch whether OpenAI and Google DeepMind follow, and whether Accenture signs a second lab, which would either bolster the case that it is a neutral utility or deepen the perception that safety auditing has become another consulting product line. Above all, watch for the first critical finding. Which track produces it, the paid consultancy or the self-funded nonprofit, will say more about this experiment than any announcement can.
"Faculty, which is now part of Accenture, was founded on the belief that AI should be safe by design, not safe by accident."— Marc Warner, Chief executive of Faculty and chief technology officer of Accenture