← All posts

Artificial intelligence

Anthropic Named Its First Independent Evaluator. It Is Also the Firm Running Anthropic's Largest Deployment

Anthropic will let Accenture staff sit inside the company with employee-level access to evaluate its models, with each side committing at least $1 billion over five years. The access is the most substantive transparency commitment a frontier lab has made — and the evaluator is the same firm Anthropic calls its largest ever deployment partner.

MAI
The header image Anthropic published on its own site alongside the announcement of its embedded-evaluation partnership with Accenture.

Anthropic said on 18 September that Accenture will supply the first team of embedded evaluators to work inside the company: outside staff with employee-level access who will red-team models, run alignment assessments, test safeguards and report what they find. Each company expects to invest at least $1 billion over five years. It is the first concrete piece of the three-step plan Dario Amodei published a week earlier in "We Must Pace the Frontier," and the only one of those steps a single lab can carry out alone.

It is also, on the same facts, a contract between Anthropic and the firm running its largest commercial deployment.

What was actually committed

TermWhat the announcement says
Access"comparable to an employee's"; the essay specifies "desks in our offices, access badges, and company laptops"
Workmodel evaluation, red-teaming, alignment assessments, safeguard testing
Publicationevaluators "can also report incidents and give the public a more informed account of benefits and risks"
Moneyat least $1bn from each company over five years
Who pays"Anthropic will fund Accenture's work directly"
Exclusivitynon-exclusive both ways; further evaluators "to be announced in the coming weeks"

The people come from Faculty, the British applied-AI company Accenture agreed to buy in January for more than £600 million, and whose chief executive Marc Warner is now Accenture's chief technology officer. Faculty's background is model testing for government, defence, healthcare and infrastructure customers — closer to assurance work than to alignment research.

The access is the substantive part

Outside evaluation of frontier models today mostly means a short pre-deployment window: an external organisation gets a finished checkpoint under a non-disclosure agreement, probes it for a few weeks, and publishes what the lab permits. Embedded access is a different instrument. Evaluators who sit inside the building during training can see the decisions, not only the artefact, and Amodei's essay says the right to publish should survive the lab's discomfort — "we can't redact findings just because they are unfavourable."

If that holds, it is the most substantive transparency commitment any frontier developer has made. Anthropic's own announcement is careful not to oversell it, noting that "there are, as yet, no standards for what information embedded evaluators should have access to, or how they should report what they find." The arrangement is being designed and staffed at the same time.

The independence problem is not incidental

On 9 December 2025, Accenture and Anthropic launched the Accenture Anthropic Business Group: roughly 30,000 Accenture professionals trained on Claude, tens of thousands of developers given Claude Code, and a stated aim of moving enterprise customers from pilots into production. Amodei described it at the time as "our largest ever deployment." Nine months later the same counterparty becomes the independent check on the same models, paid directly by the company it is checking. Accenture's shares rose about 8 per cent after hours on the news, per TechCrunch.

Anthropic addresses the objection rather than avoiding it. Its case is that Accenture is a large public company that existed long before the AI boom and does not depend on any one lab, that the partnership is non-exclusive in both directions, that it is "in dialogue with METR and other nonprofit evaluators to pilot elements of embedded evaluation using their own funding," and that direct funding is a starting arrangement — long term, "funding should come from pooled or government sources." Its summary line is the honest version of the claim:

Independent embedded evaluators do not reduce our accountability, but help to make it more verifiable.

The choice still surprised the people who had been arguing for embedded evaluation. That case was built around METR, Redwood Research and Apollo Research — small organisations with alignment-research staff and no revenue at stake. The counter-argument in Accenture's favour is a practical one the nonprofits cannot easily answer: a permanent team sitting inside a frontier lab for five years is a staffing problem, and consultancies are built to solve staffing problems while research nonprofits are not.

The test that has not been run

A week ago the open question was whether embedded evaluators would arrive with names, scope and published findings, or remain a principle. Two of the three are now answered. The third is the one that decides whether this is oversight or procurement, and it cannot be settled by any document signed this week.

It will be settled the first time Accenture's embedded team finds something Anthropic would rather not have published, and publishes it anyway — over the objection of a client that funds the work and sells its software through the same firm. Until that happens, what exists is a well-capitalised arrangement between two companies that already do a great deal of business together, on terms nobody outside them has agreed to. Anthropic says more evaluators are coming within weeks. Whether any of them is a party that can afford to lose the account is the thing worth watching.

Sources: Anthropic: Partnering with Accenture on embedded evaluation · Accenture newsroom: Accenture and Anthropic partner to build team of embedded evaluators at Anthropic · Dario Amodei, "We Must Pace the Frontier" · TechCrunch: Anthropic's first embedded evaluator is … Accenture? · CNBC: Anthropic selects Accenture as first embedded evaluator · Accenture newsroom: Accenture and Anthropic launch multi-year partnership (9 December 2025)

Keep reading