We publish what we measure.
Frontier models are already capable of real legal work. The open question is where — precisely — they meet the standard of practicing lawyers, and where they don't. That line is what we measure.
The thesis
Each model generation absorbs more of what makes legal work hard: reading a contract the way a practitioner does, spotting the interaction between provisions rather than clauses in isolation, and drafting language a counterparty will actually accept.
We pressure-test those capabilities on genuine legal tasks, harden the ones that hold up — with playbooks, structured review, and verification — and ship them as products. What doesn't clear the bar stays in the lab.
Real contracts, real stakes
Evaluation sets are built from genuine agreement types — NDAs, SaaS agreements, redemption and release instruments — with the traps practitioners actually see, not textbook hypotheticals.
Blind attorney panels
Practicing attorneys grade AI output and human output on the same contracts, against the same rubric, without knowing which is which.
A bar, not a benchmark
The question is never “did the model do well?” It's “did it meet the standard of the lawyers who do this work today?” A capability ships only when it clears that bar.
Current evaluations
Issue-detection evaluations against a blind attorney panel are underway across our first three contract families. Results are published here when the panels complete — numbers we'd stake the name on, or nothing.
The first capability that cleared the bar: redline.
AI contract analysis and negotiation drafting, entirely over email.