This website uses cookies

Read our Privacy policy and Terms of use for more information.

A lawyer files a declaration after fabricated citations turn up in a brief, and the declaration says a human reviewed the output. A firm circulates a policy for a new research tool, and the policy says a human reviews the output. A vendor is asked how its product handles privileged material, and the answer is that a human reviews the output.

The same sentence answers three unrelated questions. Usually it means someone inspected a particular output and did not find a visible problem. That has real value. A careful reviewer catches a nonexistent case, a misstated holding, a wrong date, a factual claim the source does not support. No serious approach to AI abandons that review.

A firm then asks the review to carry a larger claim: the memo was reviewed, therefore the tool is reliable; the output was checked, therefore the workflow is sound; a lawyer approved the answer, therefore the firm exercised reasonable care in selecting and deploying the system. None of those conclusions follows from the review.

A review tells you something about the document placed in front of the reviewer. It cannot tell you what the system failed to include, how the same system would respond to different wording or a different client’s facts or a later model version, or whether the system is reliable for a category of legal work. Someone inspected one artifact, and the firm is treating that inspection as an audit of the system that produced it.

I call that evidence theater: confidence that reaches further than the procedure offered to support it.

The review itself is not theater. The overclaim is.

What the reviewer cannot see

Suppose an AI system drafts a freedom-to-operate memorandum. A lawyer reads it closely. Every patent it discusses is real. Every claim construction is defensible. Every expiration date checks out. The lawyer signs off, and the sign-off is honest work.

The memorandum says nothing about the continuation still pending in one of those families. A continuation is a later application in the same family, with claims still under prosecution at the USPTO, and those claims can issue in a form that reaches the client’s product. Pending claims in a family you are already clearing are exactly the risk a freedom-to-operate analysis is supposed to flag. The memorandum is silent on it.

No sentence in the memorandum is wrong. The reviewer has nothing to catch. An omitted issue does not leave a false sentence behind. It leaves nothing.

A reviewer can examine what reached the page, and cannot inspect the analysis that never reached it. More care by the reviewer does not change that, because the limit belongs to the procedure rather than to the person performing it.

In most of these cases nobody is being deceptive at all. A policy says every output receives human review. A security presentation says a lawyer remains in the loop. A certification is signed. Each statement is true. Over time the organization treats them together as evidence that the system has been validated, which is a different proposition that nobody tested.

There are cases of real concealment. They are rarer than the ordinary version, where an actual check is performed and then someone describes it in broader terms than it earned.

Tested for what?

ABA Formal Opinion 512 is advisory rather than binding, and it does not say that every AI output must receive the same form of review. It says the level of review “will necessarily depend on the GAI tool and the specific task that it performs.”

That is the right instinct, and it moves the question one step back. To calibrate review to the tool, a firm has to know something about the tool. What did the testing establish?

A tool that performs well on a demonstration prompt has been tested on that prompt. A tool that produces accurate citations has been tested for citation accuracy. Neither result establishes that the tool reliably identifies every material issue in an unfamiliar agreement, finds controlling authority in a new jurisdiction, or holds its performance after the provider changes the underlying model.

If a firm intends to rely on a system to surface material issues, the testing has to measure issue coverage. Fluency does not measure it, citation accuracy does not measure it, and a vendor benchmark does not measure it unless the benchmark was built for that property under conditions resembling the firm’s actual work. Testing that measures something else is still useful. It just does not support the claim the firm makes when it says the system was tested.

Five questions before relying on the assurance

When someone offers a review, audit, certification, or human-in-the-loop process as evidence of reliability, ask:

  1. What was actually inspected?

  2. What conclusion is the inspection being used to support?

  3. Could the procedure reasonably establish that conclusion?

  4. When was the procedure performed, and what has changed since?

  5. What would the procedure fail to detect even if it were performed perfectly?

None of this asks a firm to build a validation program. Answering all five takes a conversation, and a firm that cannot answer them has learned something at low cost. The review stays in place. The only thing that changes is how broadly anyone describes it.

Lawyers already run this analysis on other people’s evidence. When an expert offers an opinion, we do not ask only whether the expert looked at the file. We ask what method was used, whether the method can support the opinion offered, and whether it was applied to facts the method actually fits. A qualified person having examined the material is where that inquiry starts. We apply that standard to opposing experts as a matter of course, and applying it to the systems we deploy ourselves asks nothing new of us.

So when someone says a human reviewed it, ask what was reviewed, what the review is being offered to prove, and what it had no chance of seeing.

If you have run into a policy, certification, or human-review provision that claims more than its procedure can establish, reply and tell me about it. I am gathering specimens. The category gets easier to recognize once you have seen a few.

The language from ABA Formal Opinion 512 quoted here was verified against the published opinion on July 24, 2026.

I write about using AI in legal practice without surrendering judgment, privilege, or the duty of competence at The Agentic Lawyer.

Educational only, not legal advice, and no attorney-client relationship is created. Views are my own. Attorney advertising in some jurisdictions.

Keep Reading