Abstract. A high benchmark score does not prove that an artificial intelligence system is dependable for every legal problem. It tells us how a particular version of a system performed on particular tasks, under particular rules and conditions. In law, the number of correct answers is only part of the evidence. The type and consequence of an error, the quality of the sources, the answer rate, the quality of abstention and the way the reference answers were chosen matter as well. A fair evaluation must therefore examine both the individual stages of a system and the complete path to an answer. It must define critical errors in advance and bind every result to a version, data set, date, evaluators and an explicit limit on what may be claimed. OpenLegalCore can develop this principle as a chain of bounded evidence, not as one grand score for legal reliability.
Imagine that a public institution is testing an AI system intended to assist with legal work. It prepares 100 questions. The system answers 99 of them correctly.
On the hundredth question, it misses a transitional provision. It applies a rule that was not yet applicable on the date of the event and calculates a deadline incorrectly. A person who acted on that answer could lose the opportunity to seek a legal remedy.
Yet the presentation slide still lights up with one large figure: 99%.
Did the system pass the test? Probably. Is its score high? Undoubtedly. Does the number prove that the system is dependable for the legal problem faced by the person in the hundredth case? No.
The example is fictional. It does not describe any provider or real benchmark. It illustrates a real problem: an average can conceal the very error that made legal evaluation necessary in the first place.
A high score is not, by itself, false. It becomes false when we turn it into a promise it cannot support. “Ninety-nine correct outcomes in 100 selected test cases” is not the same claim as “a legal AI system that is 99% reliable”. The first sentence reports a measurement. The second silently extends it to future cases, other fields of law, other documents, other periods and real-world consequences.
The question, then, is not whether legal AI should be measured. It must be. The question is what was measured, which failures the number concealed and how far the result is entitled to travel.
A percentage without a question is an unfinished sentence
A benchmark is a predefined collection of tasks, data, expected outcomes and scoring rules. It can help us compare two systems or determine whether a new version of the same system performs better than the previous one.
That makes it a valuable measurement tool. It does not make it a neutral window onto some general faculty called “legal intelligence”. Someone must choose the questions, the law, the language, the relevant period, the permitted sources and the scoring method. Someone must also decide what counts as a correct answer.
The statement “92%” is therefore an unfinished sentence. It needs at least the following qualifications:
- 92% of what, and out of how many cases;
- on which legal task;
- in which jurisdiction, language and period;
- with which documents and tools;
- using which system version and settings;
- evaluated by whom and under which rules;
- how many questions the system declined or failed to answer;
- and which kinds of error the test did not look for.
Without that information, the number may not be wrong. It is simply too indeterminate to support a serious decision.
Sample size imposes another limit. In our simplified example, 99 successes in 100 independent, randomly selected cases would produce an approximate 95% Wilson confidence interval from 94.55% to 99.82%. That interval describes statistical uncertainty if several demanding assumptions hold. It does not tell us whether the questions were legally representative, whether the single failure was catastrophic, or whether the system will behave in the same way after its next update. A mathematical interval cannot repair a badly chosen test set.
A benchmark can answer only the question it was built to ask
Legal work is not one task. Finding a statutory provision, classifying a contract clause, calculating a deadline, determining which version of a rule applied, summarising a judgment and weighing competing legal arguments call for different capabilities.
LegalBench, published at NeurIPS in 2023, sought to reflect that diversity through 162 tasks drawn from 36 data sources. It grouped them into six partly overlapping types of legal reasoning. That is far richer than a single classroom quiz, but it is still not a map of legal work in its entirety.
The suite is predominantly in English, is strongly oriented towards United States law and gives contracts more weight than many other legal fields. It concentrates on tasks for which an expected answer can be specified. It does not encompass whole matters with long files, changing facts, procedural choices and several reasonably arguable legal outcomes. An excellent LegalBench result can be both genuine and useful. It does not, without further evidence, demonstrate an ability to apply Slovenian procedural law in a live case.
LegalBench also revealed something less obvious. Performance belonged not only to the model, but to the instructions and examples supplied to it before the task. Across eight binary tasks, results varied significantly with the choice of few-shot examples; for one task and model, the difference under the reported metric exceeded 20 percentage points. Under the study’s statistical procedure, those differences could not reasonably be dismissed as chance alone.
This does not mean that every evaluation is arbitrary. It means that the name of a model is insufficient. A result belongs to the complete tested configuration: the version, instructions, examples, tools, sources and settings. Change one of those elements and, for evaluation purposes, we may already be dealing with a different system.
Who decides what the correct legal answer is?
Some legal tasks are straightforward to verify. Does a judgment exist? When was a regulation published? What does the third paragraph say? How many days remain if the start date, counting rule and public holidays are known?
Other questions require professional judgment. Does the cited authority actually support the proposition? Which temporal version of the legislation applies? Is a judgment binding, merely persuasive or irrelevant to the issue? Here a reference answer — the predetermined basis against which an output is assessed — cannot be prepared fairly without a lawyer, supporting authorities and written criteria.
A third category is harder still. In open interpretation, the balancing of principles, outcome prediction or strategic choice, well-informed lawyers may reasonably disagree. These tasks are not beyond evaluation. But a single “right” or “wrong” label is often inadequate.
| Type of task | Example | Suitable form of assessment |
|---|---|---|
| Objectively determinable outcome | date, quotation, amount, existence of a judgment | verification against an official source or an unambiguous rule |
| Professionally determinable outcome | temporal version, support for a proposition, legal weight of a source | written criteria, at least two reviewers for a material sample, and a process for resolving disagreement |
| Reasonably contestable judgment | open interpretation, balancing competing positions, strategy | several separate criteria: quality of authorities, treatment of counterarguments, conditions attached to the conclusion, and fair expression of uncertainty |
A reference answer is therefore not always the timeless “ground truth” of law. It is a documented professional decision made for the purpose of a particular evaluation. Its authors, authorities, legal cut-off date, disagreements and correction procedure should be visible.
This is not an academic nicety. If one person writes the questions, sets the reference answers, configures the metric and then declares the winner, the outcome also reflects that person’s choices. If reviewers do not know which system produced an answer, the risk of bias is reduced. If they can discuss disputed cases and record their reasons, the evaluation becomes stronger. If everything remains hidden, a large number demands a large amount of trust precisely where a benchmark was meant to replace trust with evidence.
The law does not treat every error alike
In an ordinary quiz, every wrong answer usually costs one point. Applied to legal work, that arithmetic can have a dangerous side effect: an awkwardly written but legally correct answer and a miscalculated appeal deadline become two identical red marks.
They are not identical.
A legal evaluation must distinguish at least two questions: what failed, and what consequence that failure could have in the intended use.
| Failure mode | Example of a possible consequence |
|---|---|
| incorrect or invented authority | the answer has no valid legal foundation |
| wrong jurisdiction, hierarchy or temporal version | the system applies law that is not determinative or was not applicable to the case |
| missed exception or contrary authority | an apparently clear conclusion may be reversed |
| failure in capture, segmentation or retrieval | decisive text never reaches the part of the system that composes the answer |
| incorrect interpretation or application to the facts | a genuine authority leads to the wrong legal conclusion |
| unwarranted certainty | the user is not warned that a decisive fact is missing or that the authorities conflict |
| unwarranted abstention | the system fails to perform a task despite having an adequate basis |
Severity is not an inherent property of a sentence either. A wrong number may be inconsequential in an internal summary and critical in a deadline calculation. Consequence depends on the intended purpose, the user, the opportunity for human correction and the action taken on the answer.
Error classes and thresholds must therefore be set before the results are seen. The evaluation plan should state which errors are critical, which are material and which are merely editorial. For some uses, the permitted number of critical errors must be zero. For others, the system may be acceptable only as an initial research aid subject to mandatory review. The decision must follow the risk of the use, not a desire to make the eventual result look better.
A weighted overall score does not necessarily solve this problem. If a critical failure is converted into penalty points and then averaged with many easy tasks, it can disappear again. An aggregate result may help with orientation, but the number and description of critical failures must remain visible beside it.
Return to the person in the hundredth case. What matters to that person is not whether the system lost one point or ten. What matters is whether it missed the rule whose omission could cost them a right.
A good retriever is not yet a good legal answer
Modern legal tools often do not answer solely from information stored in a language model. They first retrieve documents or passages. The model then uses that material to compose an answer. In the technical literature, this is called retrieval-augmented generation, or RAG. In plain language: search first, write second.
For most readers, the path hidden behind the acronym matters more than the acronym itself:
legal source → text capture → version identification → retrieval → legal assessment → cited answer → human review
An error can enter at every stage. A page image may be converted into text incorrectly. Segmentation may separate a provision from its exception. Retrieval may return a similar but legally weaker source. The correct provision may be found in the wrong temporal version. The model may misread an accurate source. A citation may lead to a genuine judgment that does not support the proposition made. A reviewer may approve the answer too quickly.
LegalBench-RAG illustrates the distinction precisely because its scope is narrow. It assesses the retrieval component: whether a system finds the exact passage in a document that has been labelled as the basis for an answer. The reported experiments used the smaller version of the dataset, with 776 queries across 72 documents, drawn mainly from contracts and privacy policies.
That result can tell us something important about retrieval. It cannot tell us whether the complete system would identify the correct jurisdiction, use the right temporal version, recognise the legal weight of the passage, find contrary authority and produce a useful final answer. The work was published as a preprint, not as a final peer-reviewed journal article.
A fair evaluation therefore needs both views. Components should be tested separately so that we can identify what failed. The complete path should also be tested because every component can pass its own check while an error arises at the hand-off between them. A correct final sentence may even be a lucky guess based on the wrong authority. If we inspect only that sentence, we cannot distinguish luck from a sound method.
Even silence needs a denominator
A responsible system should not answer at any cost. If a date, document, jurisdiction or sufficiently reliable authority is missing, asking for clarification or withholding an answer may be the best outcome.
But this is another place where a number can mislead. Consider three systems:
| System | Cases answered | Correct among answers given | Correct across all 100 cases |
|---|---|---|---|
| A | 100 | 90% | 90 |
| B | 50 | 98% | 49 |
| C | 10 | 100% | 10 |
System C can advertise perfect accuracy among the answers it gave, while leaving nine tenths of the work undone. System A resolves the most cases but answers ten of them incorrectly. System B is more cautious, yet returns half of the user’s questions without a substantive answer.
The table does not yield a universal winner. In a process where one false statement could cause irreversible harm and a qualified person is ready to take over, the more cautious system may be preferable. In a low-risk initial review of a large body of material, broader coverage may matter more. The intended purpose determines the acceptable balance.
At least four figures should therefore be reported together: the answer rate; accuracy among answered cases, sometimes called selective accuracy; the share of all cases resolved correctly; and abstention quality. The last measure asks whether the system stayed silent when the evidence was genuinely insufficient and whether it needlessly refused tasks it could have completed.
A real study shows why that distinction matters. In Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, published in the Journal of Empirical Legal Studies in 2025, Magesh and co-authors manually assessed answers to 202 challenging questions of United States law. They tested the tools between March and May 2024, so the results are not a current provider ranking.
Under their rules, a response was “accurate and grounded” only if it was substantively correct and appropriate sources supported its material legal propositions. The principal presentation of the results reports 65% for Lexis+ AI, 41% for Westlaw AI-Assisted Research and 19% for Ask Practical Law AI. The article’s introduction gives 42% for Westlaw. That internal discrepancy is better disclosed than concealed behind false precision. Rates of incorrect or misleadingly supported answers ranged from 17% to 33%. Ask Practical Law AI left 62% of cases without a complete, source-supported answer.
The study’s value is not a permanent ordering of brands. The tested versions may have changed substantially since 2024, and the deliberately difficult dataset was not representative of every query real users submit. The authors expressly caution against reading the percentages as a general error rate for legal queries.
The more durable lesson lies in the reporting. Readers could see correct, incorrect and incomplete responses at the same time. Caution was not automatically rewarded as correctness, while confident guessing did not vanish into a single score.
The evaluator must also be evaluated
Answers can be checked by a program, a person, another language model or a combination of all three. Each option has strengths and blind spots.
A program is excellent at checking a date, identifier, calculation or precisely labelled passage. It may reject a substantively correct paraphrase, however, or accept the right keyword embedded in the wrong legal conclusion.
A language model can compare many open-ended answers quickly. In Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, a powerful evaluation model agreed with human preferences in more than 80% of general cases. The same research found that answer order, verbosity and the origin of the model could influence its assessments, while difficult reasoning remained challenging. This was not a study of legal correctness. A model can therefore assist with preliminary triage, but it should not become the sole and final judge of a critical legal error.
Nor is a lawyer an infallible measuring instrument. In Hallucination-Free?, a fourth reviewer independently re-evaluated a stratified random sample of 48 responses. Agreement with the original assessment was 85.4%. Cohen’s kappa, which discounts some agreement expected by chance, was 0.77. That is reasonably strong agreement, but it is not unanimity.
The answer is not to search for a flawless evaluator. It is to use written criteria, relevant legal expertise, blinded evaluation of system identity, double review of a material sample or every critical outcome, a record of disagreement and a procedure for resolution. The version of any evaluation model, its instructions and the order in which it saw the answers should be preserved as carefully as the version of the system under evaluation.
Public benchmarks come at a price
A published test set allows researchers, users and providers to inspect the questions and challenge defective reference answers. That is a major advantage. At the same time, developers can repeatedly tune a system on those very tasks, and some questions or answers may already be present in a model’s training data.
A high score may then partly reflect familiarity with the test rather than competence on a new case. Without access to all training data, that possibility can rarely be excluded completely.
Oren and co-authors developed a statistical method for detecting certain direct overlaps, including in models whose internal workings are not public. Their approach depends on stated assumptions and does not detect every indirect or partial form of contamination. Failure to detect an overlap is not proof that no overlap exists.
A sensible response combines transparency with control: a public portion for professional scrutiny, a sealed portion for final evaluation, fresh tasks, a clear separation between development and testing, an access log, and periodic rotation of examples. A separate regression set of already known failures should check that the next version does not repeat a defect that has already been found.
Even the best-protected dataset ages. Legislation, case law, source collections, the underlying model, instructions and interfaces between components all change. A result without a date and version is a photograph with no record of when or where it was taken.
What the AI Act requires
The European Union’s Artificial Intelligence Act does not prescribe a single universal score for legal reliability. Its logic is different: it connects requirements to a system’s intended purpose, risks, documented testing and the information users need.
In the consolidated text of Regulation (EU) 2024/1689, Article 9 requires continuous risk management for high-risk AI systems and testing against previously defined metrics and probabilistic thresholds appropriate to the intended purpose. Where data sets are relevant, Article 10 requires testing data to be relevant and sufficiently representative, including in relation to the specific setting in which the system is intended to be used. Article 13 requires instructions for use to include the level of accuracy, the relevant accuracy metrics and known or foreseeable circumstances that may affect performance. Article 15 requires appropriate levels of accuracy, robustness and cybersecurity throughout the lifecycle, with the levels and relevant metrics declared in the instructions. Article 17 and Annex IV support documented testing, validation, records and technical documentation.
Put simply, where these provisions apply, the Regulation does not ask for a magic number. It asks for evidence that evaluation fits the purpose and risks of the system. That is a strong basis for purpose-specific, transparent measurement. It is not a statutory formula for adding up legal errors and producing one European percentage.
Scope matters too. Annex III includes certain systems intended to assist judicial authorities in researching and interpreting facts and the law and in applying the law to concrete facts. It does not follow that every tool used by a lawyer is high-risk. Article 6 contains conditions and exceptions, so classification requires an assessment of intended purpose, deployer, impact and concrete function.
Timing also matters. Following the amendments made by Regulation (EU) 2026/1744, the relevant requirements for systems falling under Article 6(2) and Annex III apply from 2 December 2027. For systems under Article 6(1) and Annex I, the relevant date is 2 August 2028. As at the cut-off date of this article, 5 September 2026, those future dates must not be presented as if the requirements already constituted a general obligation for all legal systems.
Nor does the Act always require an external assessment by a notified body for the systems listed in points 2 to 8 of Annex III. Article 43(2) provides for a conformity assessment procedure based on internal control without the involvement of a notified body. Independent expert evaluation may be excellent practice. It would nevertheless be inaccurate to claim that the law already requires it indiscriminately for every legal tool.
Who evaluates the evaluator?
A benchmark is not merely a technical file. It is also an allocation of authority: who chooses which tasks matter, who may see confidential data, who sets the reference answers, who pays for the evaluation and who decides which results are published.
In the peer-reviewed article There Is No Free Benchmark: An Institutional View of Legal AI Benchmarking, published in PNAS in 2026, Guha and co-authors divide benchmarking into six institutional decisions: task specification, dataset composition, execution protocol, metrics, system selection, and transparency and publication.
A provider has the best access to its system, versions and logs, but also has an interest in presenting the result favourably. A buyer understands its own workflow, but its dataset may not transfer elsewhere. A university or non-profit organisation has greater distance, but often less funding, less access to authentic user data and less access to closed products. A public authority has a mandate, but may lack specialist capacity or speed.
No evaluator is automatically ideal. Credibility comes from visible roles and safeguards: funding disclosure, separation of development from final approval, participation by lawyers qualified in the relevant jurisdiction, criteria fixed in advance, a route for challenge and publication of unfavourable results as well as favourable ones. An internal test can be extremely useful if it is described as internal. The word “independent” should describe genuine governance, not merely a different person in the same team.
State the claim before publishing the number
Criticism of weak evaluation should not lead to less measurement. It should lead to measurement precise enough that it cannot be turned, without warning, into an overbroad promise.
The first step is surprisingly simple: before testing, write one sentence stating what the result is intended to establish. For example:
On date Z, version X achieved the reported result under the published rules when retrieving pre-labelled decisive passages from Slovenian administrative decisions in environment Y.
That sentence is less dazzling than “our legal AI is 98% reliable”. It is much stronger as evidence. It immediately exposes the task, setting, time and limit of the claim.
Five connected steps should follow.
1. Freeze the evaluation package
Before a run, record the purpose, user, jurisdiction, language, legal cut-off date, sources and method for selecting cases. Fix the versions of the system, underlying model, instructions, tools and settings. Give every reference answer its authorities, authors and disagreement procedure. Define the metrics, error taxonomy, severity levels, abstention rules, acceptance thresholds and permitted exclusions in advance.
“Frozen” does not mean that a benchmark can never be improved. It means that the rules do not change silently during one evaluation. If a defective reference answer or a new failure mode is discovered, record the change, issue a new version and, where necessary, repeat the comparison.
2. Measure components and the complete path
For a composite legal system, separately test document capture, optical character recognition, source identity and temporal versioning, segmentation, retrieval, legal processing, drafting, citations, abstention and human review. This makes the cause of a failure visible.
Then run a realistic task from the actual input to the final outcome. This reveals failures at the interfaces: a lost temporal label, a broken link between a passage and its source, or a warning that never reaches the user.
3. Report results in disaggregated form
An aggregate percentage may remain, but it needs company. Report absolute counts; results by task type, jurisdiction, period and difficulty; errors by type and severity; answered and unanswered cases; unjustified abstentions; every critical failure; repeated runs for variable outputs; and comparison with the previous version.
Time and cost matter, but they should remain distinct from legal quality. A fast wrong answer is not more correct because it is fast. A slow correct answer may still be unusable in a process that requires an immediate result. Both qualities should be visible instead of being blended, without explanation, into one score.
4. Create an evaluation evidence record
The result should travel with a record that supports professional scrutiny. At a minimum, it should contain the date, version, sufficiently detailed environment, data used, metrics, denominators, outputs, evaluators, disagreements, critical errors, known limitations, funding and the responsible person or institution.
Not every data item can always be published. Legal documents may be confidential, systems may be proprietary, and test cases may be sealed precisely to preserve their value. Privacy does not, however, require secrecy about the method. The public — or an authorised reviewer — can be shown the provenance and selection method of the data, file hashes for the exact materials, rules, aggregate results, limitations and a list of what legitimately remains closed.
5. Make a bounded decision and set retesting triggers
The conclusion need not be simply “pass” or “fail”. A system may be accepted for a narrow task, accepted only with mandatory human takeover, returned to development, or rejected because the evaluation itself was invalid. The decision should state the purpose for which it applies and which change will trigger a new evaluation.
A new model version, legal source collection, instruction, provider or repaired interface between components can change the result. An acceptance label without an expiry or retesting triggers quickly becomes a historical ornament.
Five levels of evidence instead of one grand score
These principles can be organised into a practical ladder. It is not a legal standard or a universal formula. It is a proposed way to build evidential strength one level at a time.
L0 — claim boundary. Before a public number appears, write the proposition the result may support and the propositions it may not support. If that boundary cannot be written, the result is not ready to become a public promise.
L1 — component acceptance. Test a bounded function separately: capture, optical character recognition, retrieval, citation checking or another tool. The result supports only that function under the stated conditions.
L2 — hand-offs between components. Check whether source identity, page, time, warnings and reasons for abstention pass correctly from one stage to the next. Many dangerous failures do not originate inside one component, but at the interface between two of them.
L3 — end-to-end evaluation for a defined purpose. A realistic case travels from documents and facts to the answer, citations, limitations and any handover to a person. Each jurisdiction and type of work needs its own evaluation; success does not transfer automatically.
L4 — the human workflow and post-deployment monitoring. Measure review time, errors detected and missed, corrections, unwarranted reliance, handovers, incidents and the effect of updates. Only this level shows how a technical result behaves inside a real institution.
The ladder is not five points to be added into a new single score. Each level answers a different question. A higher level does not erase a lower one, and the result of one component does not migrate to the complete system.
OpenLegalCore: evidence, boundary and proposal
OpenLegalCore can incorporate this measurement discipline into its development and publication practice. To do so credibly, it must keep three things rigorously separate: what has been demonstrated publicly, what is currently being developed, and what this article merely proposes.
A publicly inspectable example is the acceptance record for Legal OCR Pipeline v0.1.2. OCR stands for optical character recognition: converting images of pages into machine-readable text.
The record covers one run on one private legal filing of 174 pages. All 174 pages were processed in 45.23 seconds, with no retries and no failed pages. The mandatory 40-page sample received 39 PASS labels and one MINOR label on manual review: a legally immaterial variation in a person’s name. The aggregate label was SAMPLE_PASS, meaning that the sample passed the stated acceptance review.
Those figures are useful precisely because the record does not allow them to mean more than they prove. They are not an average across multiple runs, a service-level promise, a review of every page, or a character- or word-error rate measured across several kinds of document. Because the source filing and output remain private, a third party cannot reproduce the complete test independently. Above all, success in optical character recognition does not demonstrate the quality of retrieval, legal assessment or the final answer.
On the public OLC Engine page, version 0.0.7 is labelled PRIVATE · ACTIVE DEVELOPMENT. It is described as the current core for bounded legal-source planning, retrieval from maintained Slovenian legal sources, temporal treatment of legislation, conservative source selection, cited or structured output and controlled delivery. The broader legal-work architecture is separately presented on that page as a future target. This is a statement by the project, not an independent evaluation. It cannot support a whole-system score or a comparison with competing products.
The L0–L4 ladder is therefore a proposal for OpenLegalCore’s future measurement discipline, not a description of completed functionality. Its value will depend on implementation: public rules, disclosed versions, appropriately qualified legal reviewers, a protected sealed test set, publication of critical failures and an avenue for external challenge.
Open code helps because it permits inspection of at least part of the method, and independent replication where the data and environment are also available. It does not prove correctness. Public code can use the wrong authority; a private system can conduct a rigorous internal evaluation. The decisive question is whether each claim is honestly tied to evidence that actually supports it.
OpenLegalCore therefore does not need a large number proclaiming that the system “understands law”. It needs a chain of bounded evidence: what it captured correctly, which source version it preserved, what it retrieved, how it applied a specialised rule, how it supported its conclusion, when it abstained and what a human reviewer confirmed.
This extends the principle developed in A Result Is Not a Method: a result becomes verifiable only when it preserves the path by which it was produced. Evaluation adds one further requirement. That path must also show what the test omitted and what must not be inferred from its result.
Ten questions to ask before trusting a high score
Before a legal AI result becomes a reason to buy, deploy or trust a system, ten questions need answers:
- What did the system actually have to do? Retrieve passages, verify citations, calculate deadlines, interpret law or produce a complete answer?
- For which law was the evaluation designed? Which jurisdiction, language, document types and period?
- How were the cases selected? Do they represent ordinary work, rare critical cases or a deliberately constructed stress test?
- Which version was tested? Which system, underlying model, legal sources, instructions, tools and settings?
- Who set the reference answers? On the basis of which legal authorities, as at what date, and with what process for resolving disagreements?
- Does the result concern one component or the complete path? What remained outside the measurement?
- Which errors does the average conceal? How many were critical, material or minor, and what failed in each case?
- How often did the system abstain? How often was that justified, and how often could it have completed the task?
- Who performed and funded the evaluation? Was it internal, independent or collaborative, and who could prevent an unfavourable result from being published?
- What does the result expressly prove, and when does it expire? Which change triggers a new evaluation?
Without answers to those questions, a percentage is not yet a strong enough basis for a serious legal decision.
Back to the hundredth person
The person in our fictional hundredth case does not need a perfect system. No such system exists. They need a system and a workflow that look for dangerous errors where they can arise, refuse to hide them in an average and alert a person in time when the evidential basis is too weak.
One critical failure does not automatically make a system useless. It does mean that we need to know what failed, why it failed, whether safeguards should have stopped it and for which purposes the system remains acceptable. That is far more useful than arguing about whether 99% is “good” or “bad”.
Benchmarks are essential. They allow progress, comparison and regression testing against failures that have already been corrected. Their strength does not lie in the size of the number. It lies in the precision of the question, the integrity of the procedure and the visibility of the boundary.
Law does not need the score that sounds largest. It needs an evaluation that, even when the score is high, does not forget the person in the case where the system failed.
Selected sources
- Regulation (EU) 2024/1689 — consolidated text as at 27 July 2026
- Regulation (EU) 2026/1744 — amendments to the AI Act timetable and other provisions
- Guha et al., There Is No Free Benchmark, PNAS 2026
- Guha et al., LegalBench, NeurIPS 2023
- Pipitone and Alami, LegalBench-RAG, arXiv v1
- Magesh et al., Hallucination-Free?, Journal of Empirical Legal Studies 2025
- Zheng et al., Judging LLM-as-a-Judge, NeurIPS 2023
- Oren et al., Proving Test Set Contamination in Black-Box Language Models
- NIST, Wilson confidence interval for a proportion
- OpenLegalCore, Legal OCR Pipeline v0.1.2
- OpenLegalCore, canonical OCR acceptance record v0.1.2
- OpenLegalCore, OLC Engine
Legal, research and project sources last verified: 5 September 2026.