Loading...
Loading...
Every number below links to the record it came from, and carries the date that record was measured.
Two different things get called accuracy, and they are not equally strong here. When this tool quotes your contract, the quote is almost always exactly what your contract says, and every quote is checked against your document before you see it. Whether it finds everything worth finding is a separate question, and there it is a first pass rather than the last word. Both numbers are below, each with the standard it was measured against. Read the second before you decide what this is worth to you.
This is the failure that matters most. A tool that invents a clause, or quietly rewords a real one, is worse than no tool: you would go into a negotiation arguing about a term your contract does not contain, or sign past one it does. So every finding carries the exact words it was drawn from, and those words get searched for in your document before the finding reaches you.
That search is the plainest check there is. It takes the quote, takes your file, and looks for one inside the other, allowing for the ways a PDF or a Word export rewraps text and nothing else. No model is asked whether it did a good job. Each quote in your results is marked when it was found word for word, and marked differently when it was not, so you are never left guessing which kind of statement you are reading.
The figures below are that same check, run over a corpus of real contracts, which is how we can tell you what to expect before you upload anything.
109 of 111 clause quotes were verified verbatim (98.2%): 108 matched the source exactly and 1 matched across a page break, leaving 2 transcription slips and 0 invented clauses
Every quote in that run was checked mechanically against the source document, by string matching rather than by asking a model whether it had done well
The 2 slips in that run were transcription errors: a quote that did not match the source character for character, close enough to look right and wrong enough to break the verbatim promise. Neither invented an obligation. Inventing a clause is the far more dangerous failure, which is why its count sits in the figure above rather than in a footnote.
Not all of it, and this is the number most tools do not publish. Semantic completeness measures how much of what a careful reader would flag the review actually flagged.
Semantic completeness averaged 55.8 across 9 contracts
That number needs the standard it was measured against, or it says something worse than the truth. The benchmark is every point a careful reader with no deadline would raise, marked up clause by clause in advance, and the score is how much of that list the review reached on its own. Against that, it surfaced a bit over half.
So: a strong first pass that catches most of the obvious problems and some of the subtle ones, and not a substitute for a lawyer reading your contract. That is why the disclaimer at the bottom of every page says what it says, and why the section above matters as much as this one. What it does hand you is exact: the clauses it raises are really in your document, quoted word for word, and checked.
We publish this because a number that only flatters us is marketing, and because improving it is the current priority. When it improves, the date above changes and the old figure goes away.
There used to be two, and the older one was retired on purpose. It was measured against a test corpus we replaced hours later the same day: the documents got roughly twice as long and the answer keys were rewritten against the new text. Our own record says in writing that scores from before that change cannot be compared to scores after it, so quoting the older, friendlier figures beside the newer ones would have implied a comparison that does not exist.
The retired figures were also better looking than these. That is precisely why they had to go: a number kept because it flatters is not a measurement.
The run that survives is published in full, defects included, at the eval record for 2026-08-04.
Each claim carries the date its measurement ran. A build check fails if any of them goes more than 180 days without being re-measured, so a stale number cannot quietly survive here. Every measured figure lives in one file and is imported into this page rather than typed into it, which means the page cannot say something the measurement does not.
This matters because the alternative is common. An accuracy figure that never changes is a figure nobody is re-checking.
The reason to publish a measurement is so that you do not have to take it on faith. Ours ran on documents we chose. The run that settles anything for you is the one on the document you actually have to sign.
Every finding comes back with the words it was drawn from, already searched for in your file, and marked with whether it was found. You do not have to take that marking on faith either: the quote is right there, and your document is on your screen. Read the completeness section above first, and treat what comes back the way it asks to be treated: a first pass, not the last word. No account, no email.
If a number here does not match the record it links to, that is a bug and we would like to know about it.