How We Measure Contract Risk Detection
Any tool can claim it catches the risks in a contract. The claim worth reading is the one that comes with a method attached: a fixed test, a repeatable count, and a number that would fall if the underlying read got worse. Here is exactly how we measure ours.
A detection rate is a claim about a process, not a single event. It is easy to run a tool once, count what it found, and publish the count as though it were a fact about the tool. It is also, on its own, close to meaningless. Before any number leaves this company, it has to survive being run again. And again. What follows is the discipline behind the numbers we do publish, and the reasoning for the ones we deliberately do not.
Why a single run proves nothing
We learned this the expensive way, on our own internal measurements of a drafting pipeline. The identical code, run five times against the identical set of test agreements, with nothing changed between attempts, returned four passing outputs out of ten on one run, eight out of ten on the next, and seven out of ten on the run after that. Nothing in the system had moved. Only the sample had.
That is the entire lesson, and it is worth stating plainly because it cuts against how most accuracy claims get made. A single run is a draw from a distribution, not a measurement of one. Quote the best run out of five and you have described your best day, not your typical one. Quote the worst and you have described the opposite mistake. Neither is honest, and both are common, because a single run is cheap and a distribution is not.
Our rule since that discovery has been simple to state and hard to skip: nothing gets quoted as a rate until it has been measured across at least five independent runs, and the number we publish is the aggregate across those runs, not the best one, not the most recent one, and not the one that happened to land first. If we cannot show our work across five runs, we do not publish the number at all.
The same rule shapes how we test changes before they ship, not only how we describe results after the fact. When we isolate a variable, we isolate exactly one at a time. A test that runs the same contracts once through the full document pipeline and once through analysis alone, holding the reading step constant while removing the document-parsing step, tells us whether an error came from how a page was read or from how a clause was judged. Change two things in the same test and a result that looks like an improvement can just as easily be two effects canceling out, which is its own way of proving nothing.
What a locked benchmark is
A five-run average only means something if every run is scored against the same yardstick. Ours is a locked benchmark built from real commercial agreements drawn from public company filings: contracts written by real lawyers for real deals, not invented examples built to make a system look good. Legal experts annotated more than thirteen thousand clauses across forty-one categories in that corpus, marking what each clause is and what risk it carries.
The word that matters most in that description is locked. The benchmark does not change when we change our reading pipeline. It is the same set of contracts, the same expert markings, and the same scoring method every single time we test a change, which is precisely what keeps the result honest. A benchmark that shifts to flatter whatever changed most recently is not a benchmark. It is a mirror.
How the number is computed
Two things are measured against that locked corpus, and both have to hold up together. The first is recall of material risk: of the risks the experts actually marked in a contract, how many did the read find. The second is false-alarm control: of the things the read flagged, how many were real. A system that flags every clause in a contract can post a high recall number while being useless, because the noise buries the signal it was supposed to surface. We track both, and neither is allowed to improve at the other’s expense.
Those two measures combine into a single tracked figure, and that figure has an anchor and a floor. The anchor is the current gate of record, roughly nine of every ten material risks caught. The floor sits a few points beneath it and functions as a regression gate. Every change to how the pipeline reads a contract is re-run against the locked benchmark before it ships, and a change that would push the score below the floor does not ship, regardless of what else it improves. The bar can move up. It is not permitted to move down.
The scoring itself does not stop at a single blended average. Every category of clause gets its own precision and recall, computed with statistical intervals sized for a category that might only appear a handful of times in the benchmark, rather than the wider intervals that would be appropriate for a category that appears hundreds of times. The aggregate score is checked with resampling: the same set of contracts is redrawn thousands of times to see how much the overall number would move if the benchmark had been assembled a little differently. And every run is checked for outliers, individual contracts that scored far worse than the rest, because a single badly handled document buried inside a good average is exactly the kind of failure a headline number is built to hide.
The GovCon measurement
The same discipline applies outside the general benchmark, on the narrower question of federal teaming agreements and subcontracts. There, the test set is twenty genuine teaming and subcontract agreements drawn from public securities filings, scored against a published rubric rather than a general clause taxonomy, because the risks that matter in a teaming agreement, most notably affiliation and the ostensible subcontractor rule, are specific to that document type.
On that cohort, detection of affiliation and ostensible-subcontractor risk did not start high. Across three measured iterations of the underlying playbook it moved from finding none of them on the first iteration to nearly nine in ten, and the final step of that improvement was tested for statistical significance rather than assumed: p equals 0.00076. On a second, independent measure in the same cohort, the read correctly identified which party’s side of the deal it was reading in twenty out of twenty documents, with zero inversions. Getting the party’s perspective backward would make every other finding on a contract worse than useless, so we measure it on its own rather than folding it into a single blended score.
The rubric itself measures more than the two headline figures above. It also scores overall coverage against the full set of provisions a careful reviewer would check in a federal prime-to-subcontractor agreement, a stricter and more granular test than the affiliation question alone, and that broader coverage score sits meaningfully below the affiliation figure. We publish the affiliation and party-perspective results because they are the two measures with the clearest stakes for a small business signing this kind of agreement. We do not extend either of them into a general claim about federal contracting competence, because that is not what a twenty-document cohort on one document type was built to prove.
Why we publish without rounding
We say nine in ten. We do not say ninety percent. That is a deliberate choice, not an informal habit. Ninety percent reads as a precise instrument reading, and precision we have not earned is a kind of dishonesty, even when the underlying number happens to round to it. Nine in ten reads as what it is: a measured tendency across a fixed benchmark, not a promise about the next contract that comes through the door.
The same discipline governs everything else on this page. We never publish a single-run rate, no matter how good that run was. We never quote an internal number that has not been checked against the locked benchmark. And we never let a public claim diverge from the internal gate that governs what we ship. The figure a customer reads and the figure that decides whether a change reaches production are the same figure, measured the same way.
What the number does not mean
None of this is legal advice, and none of it is a guarantee about any single document. A detection rate is a statement about how the read performs across a fixed benchmark of many contracts. It is not a statement about the one contract in front of you, which may be better drafted, worse drafted, or structured in a way the benchmark never encountered. Nine in ten material risks caught, measured across thirteen thousand clauses, still leaves room for the one clause in your document that lands in the other tenth.
It also does not mean every category of clause performs identically. A benchmark built from more than forty categories will always have categories the read handles better than others, which is the entire reason we score precision and recall separately for each one rather than reporting a single number and stopping there. A measured average is honest about the whole benchmark. It is not, on its own, a promise about any one row inside it, and we do not present it as one.
That is exactly why the read comes back to you as a report to be read, not a verdict to be accepted. The point of measuring the process this carefully is to earn the right to say, plainly, what the process is good for and what it is not. A measured number you can check is worth more than a confident one you cannot, and it still leaves the reading of your own contract, in the end, to you.
This article describes how BeforeJD evaluates its own analysis process. It is not legal advice, and it is not a guarantee about the outcome of any individual contract review. Consult a licensed attorney for advice specific to your situation.
The full benchmark, the gate, and the principles behind it are on our methodology page. If your document is a federal teaming agreement or subcontract, the measurement built specifically for that document type is on the GovCon page.