This website uses cookies

Read our Privacy policy and Terms of use for more information.

One of the most durable claims in legal AI began with a bar exam score.

You have almost certainly heard it: GPT-4 performed at roughly the 90th percentile on the Uniform Bar Examination.

The number originated in a specific technical document: OpenAI's March 2023 GPT-4 Technical Report. From there it traveled outward into conference presentations, innovation memos, news coverage, and vendor decks, where it became a convenient way to communicate a much larger idea: this system can perform at the level of a very capable lawyer.

Then a peer-reviewed reanalysis, published in the journal Artificial Intelligence and Law, checked the comparison population.

The 90th-percentile estimate was based on a February examination population that included a large proportion of repeat test takers. Repeat takers generally perform worse than first-time candidates.

Compared against different populations, the same score meant something else entirely. Eric Martinez, the author of that reanalysis, estimated the model ranked around the 62nd percentile among first-time test takers, and around the 48th percentile among those who actually passed.

The written-component numbers are more revealing, but only up to a point. Against first-time takers, the model came in around the 42nd percentile. Against those who passed, around the 15th.

Compared against

Overall percentile

Written components

Repeat-inclusive February population (the original basis)

~90th

-

First-time test takers

~62nd

~42nd

Those who passed

~48th

~15th

One GPT-4 score. Different comparison populations, different percentiles.

The essay and performance-test components resemble some parts of legal practice more closely than multiple-choice questions do. They still do not measure client judgment, factual investigation, strategic choice, or responsibility for a live matter.

So the 15th percentile is not the "real" lawyer number either. The point is that the 90th percentile was never entitled to carry the conclusion placed on it.

The underlying score had not changed.

The exam had not changed.

The denominator had.

That is the important part of the story. Nothing had to be fabricated for the claim to become misleading. "90th percentile" could be accurate within the population used to calculate it and still create an entirely wrong impression when offered as evidence of performance relative to newly qualified or practicing lawyers.

The number moved from one context into another and left its conditions behind.

I call that a traveling number: a measurement that is accurate where it was produced but misleading where it is later used.

A traveling number is one form of evidence theater: the claim expands while the test stays the same.

Benchmarks do not travel well

A benchmark measures performance under defined conditions:

  • a particular task set;

  • a particular model version;

  • a particular scoring method;

  • a particular comparison population;

  • and a particular date.

Buyers rarely hear any of that.

They hear "90th percentile," "95% accurate," or "trained on the complete patent corpus." Those phrases are easy to remember and easy to repeat. The conditions disappear and the number keeps moving.

By the time the claim reaches a procurement meeting, it is usually supporting a proposition that was never tested.

A model performed well on a standardized exam, therefore it can perform legal work.

A contract tool reached a particular accuracy rate on a benchmark dataset, therefore it will identify the material risks in our agreements.

A search product indexed millions of patents, therefore it will surface the reference that matters to this claim set.

Each statement contains a jump. The benchmark describes the test condition. The buyer supplies the real-world conclusion.

Patent lawyers should recognize the problem

Patent search and drafting tools often make claims about the size of the corpus they use, the quantity of documents they index, or the breadth of their training data.

Those facts may matter. They do not answer the question the patent lawyer actually cares about.

Indexing the full corpus of issued patents is not the same as reliably retrieving the reference that anticipates claim 3. One describes coverage of a database. The other describes performance on a retrieval task.

And when a tool does report a retrieval number, read its denominator the same way. Recall at k on a vendor's prior-art benchmark measures retrieval against that benchmark's known relevance set: how often the tool surfaced an answer someone had already identified, inside a curated collection. It does not establish that the tool will surface the reference that reads on a new claim set. So ask what the retrieval was actually keyed to - semantic similarity, lexical overlap, classification, citation structure, or mapping to the elements of the claim as construed - and whether the benchmark ever tested the last of those. A strong benchmark score and a missed 102 reference are entirely compatible outcomes. Your claim was never in the benchmark.

A system may also generate an application that resembles a professionally drafted one without reliably identifying the inventive contribution, preserving commercially meaningful claim scope, or anticipating the art most likely to matter during prosecution. Resembling the work product is not the same as doing the work.

In freedom-to-operate and validity work, the gap stops being an embarrassment and becomes a wrong opinion. "The search was 95 percent accurate" turns into "we cleared the product to launch," and a missed reference is now something a client relied on to spend money, to file, or to sue. The number's move from the benchmark slide to the opinion letter is the exact point where its conditions fall away and someone else assumes the risk.

The question is never whether the benchmark is impressive. It is whether the benchmark measured the property you intend to rely on.

Five questions for the next benchmark slide

Start with the denominator, and ask it plainly. The answer, or the inability to give one, is itself information about how much weight the number can carry.

  1. Who or what was the system compared against?

  2. What tasks were included, and how closely do they resemble the work we intend to perform?

  3. Which model, product version, and date did the result cover?

  4. What property was measured, who scored it, and against what standard?

  5. What material failure could the benchmark miss even if it were administered perfectly?

These are not hostile questions. A strong vendor should be able to answer them, or say precisely what has not been tested.

The answers may show the benchmark is highly relevant to you. They may show the result is real but narrow. Or they may show the number has traveled so far from its original conditions that it can no longer support the claim being made in front of you.

None of this means benchmarks are fraudulent. Most are not.

The ordinary path is duller than fraud, and more common. A flattering number gets repeated because it is flattering. Each person shortens the explanation slightly. The conditions fall away. Eventually the number stands in for a conclusion nobody measured. No bad actor is required.

Which is why the lawyer's ordinary habits are the whole defense here. We already know how to examine an expert opinion: identify the question, inspect the method, test the comparison population, and ask whether the conclusion reaches past the data. A benchmark deserves exactly the same treatment.

Before you rely on the number, make it show its passport.

If a vendor has put a performance number in front of your team, send me the claim and its source. I am selecting three for a written Benchmark Teardown, at no charge, while I develop the format. Anything I publish will be anonymized, and I will not use material you are not free to share. The specimens are teaching me more than the decks are.

Percentile figures are from Eric Martinez, "Re-evaluating GPT-4's Bar Exam Performance," Artificial Intelligence and Law (2024), verified against the published article on July 25, 2026.

I write about using AI in legal practice without surrendering judgment, privilege, or the duty of competence at The Agentic Lawyer.

Educational only, not legal advice, and no attorney-client relationship is created. Views are my own. Attorney advertising in some jurisdictions.

Warmly Ran GTM With No Sales Team. Here's How.

That's what Warmly proved. They defined ICP, scored buying intent, and surfaced the right accounts before a human ever touched a lead. HubSpot noticed.

On August 12, Max and Keegan are rebuilding it live in HubSpot — and showing you how to replicate it this week. HubSpot Credits included when you join HubSpot for Startups.

Stop typing what you could say in 10 seconds.

Wispr Flow turns your voice into clean, professional text inside any app. Emails, Slack, client updates — speak once, send without editing. 4x faster than typing.

Keep Reading