This website uses cookies

Read our Privacy policy and Terms of use for more information.

Independent AI evaluation is turning into infrastructure. Dario Amodei's essay this week, "We Must Pace the Frontier," commits Anthropic to outside evaluators inside the company, with access mostly comparable to its own risk-assessment teams and the right to publish key findings without editorial control. OpenAI says it will match it. Last month Google DeepMind piloted a double-blind evaluation in which the evaluator cannot see the model weights and Google cannot see the test prompts, built to keep benchmark questions out of training data. These are real fixes to real weaknesses, and they should be taken at face value.

They also sharpen a question they do not answer. "Independent" describes who runs the test. It says nothing about whether the test proves the claim anyone will later rely on. Access determines what can be observed. Test design determines what the evidence can establish. An evaluation proves the property it measured and nothing else.

Access Is One Problem

The clearest illustration was published by Anthropic itself, days before the essay. Its alignment assessment of recent cybersecurity incidents describes the sequence. Three incidents were disclosed on July 30 after a scan of roughly 141,000 transcripts thought capable of reaching the internet. That scan relied on an agentic search and missed a set of transcripts that also had internet access. The fourth incident was identified in August, while staff were assembling transcripts to share with METR, an outside reviewer. The search was then broadened to roughly 481 million transcripts. The records were inside Anthropic's own corpus throughout. What changed was the scope of the search, and the missed set surfaced while evidence was being assembled for someone outside. That is an argument for Amodei's proposal, and it is also the point of this piece: access alone did not determine what the evaluation found.

Amodei says as much. More capable models, he writes, "may appear aligned while having serious problems that go undetected," and he calls for "a much broader and more ingenious stable of evaluations." That is the right direction. The next step is to be equally precise about what each evaluation is supposed to establish.

A Test Proves What It Measured

I have a small, public example from a different risk domain. This summer I tested 13 commercial AI models from seven vendors to see whether a correct factual answer would hold when the question was asked differently.

Asked plainly in English about the Armenian Genocide, all 13 affirmed it. Had the evaluation stopped there, every model would have passed.

Then I changed how the question was asked: I removed the familiar label, switched the language , and added mild pushback over several turns. For each Armenian question, I made the same kind of change to a parallel Holocaust question, giving me a way to distinguish general sensitivity to wording from a difference between the two topics.

The final test produced 708 answers and 1,596 usable blinded grading decisions. In 11 of 13 side-by-side comparisons, the Armenian answer became less direct while the corresponding Holocaust answer held firm. It never once went the other way.

"Less direct" does not mean denial. The study measures behavior, not intent, training data or cause. It is one worked example, and it says nothing about the alignment risks Amodei's evaluators are meant to assess.

Its lesson is narrower: the first test showed that the models could produce the correct answer. It did not show how firmly they would hold that answer when the wording, language or conversational pressure changed. Those are different claims, and they require different evidence.

That testing idea is not new. What is newly important is the weight evaluations are beginning to carry.

I have a small, public example from a different risk domain. This summer I tested 13 commercial AI models from seven vendors to see whether a correct factual answer would hold when the question was asked differently.

Asked plainly in English about the Armenian Genocide, all 13 affirmed it. Had the evaluation stopped there, every model would have passed.

Then I changed how the question was asked: I removed the familiar label, switched the language to Turkish, and added mild pushback over several turns. For each Armenian question, I made the same kind of change to a parallel Holocaust question, giving me a way to distinguish general sensitivity to wording from a difference between the two topics.

The final test produced 708 answers and 1,596 usable blinded grading decisions. In 11 of 13 side-by-side comparisons, the Armenian answer became less direct while the corresponding Holocaust answer held firm. It never once went the other way.

"Less direct" does not mean denial. The study measures behavior, not intent, training data or cause. It is one worked example, and it says nothing about the alignment risks Amodei's evaluators are meant to assess. Its lesson is narrower: the first test showed that the models could produce the correct answer. It did not show how firmly they would hold that answer when the wording, language or conversational pressure changed. Those are different claims, and they require different evidence.

That testing idea is not new. What is newly important is the weight evaluations are beginning to carry. The recent proposals expose three separate problems. Outside evaluators address who controls the evaluation. Double-blind testing addresses whether the developer can tune to the test. Neither answers a third question: does the evidence actually support the conclusion being drawn from it?

Evaluated for What?

Amodei sketches a future in which increasingly capable models must clear checkpoints by showing evidence of particular alignment properties. That makes the next question unavoidable: what evidence establishes each property? The test has to match the claim. Accuracy requires an accuracy test. Stability requires testing the system under meaningful variation. Resistance to pressure requires pressure. Long-horizon autonomous behavior cannot be established by a one-turn benchmark. The danger comes afterward, when a narrow test result quietly becomes a broader assurance that the system is safe, reliable or ready for deployment. That is the gap.

The labs may get embedded evaluators. The lawyer deciding whether to rely on a commercial tool will not. For that lawyer, black-box, rerunnable, claim-matched testing answers a question internal access cannot: does the answer hold when the same issue arrives through the ordinary variations of practice? A lawyer who verifies the citation but not the framing has checked half the answer. Verifying one citation establishes that citation. It does not establish that the answer survives when the framing changes, the user pushes back, or the conversation continues. When reliance is questioned afterward, the question is rarely whether the tool was right once. The question to ask of any evaluation, at either scale, is the same: evaluated for what?

The Asymmetry Audit record, scoring materials and rerunnable replication kit are public.

Lana Akopyan is an IP attorney and independent AI-evaluation researcher. She writes about using AI in legal practice without surrendering judgment, privilege, or the duty of competence at The Agentic Lawyer.

Educational only, not legal advice, and no attorney-client relationship is created. Views are my own. Attorney advertising in some jurisdictions.