Neuralaw
Research

We publish what we measure.

Frontier models are already capable of real legal work. The open question is where — precisely — they meet the standard of practicing lawyers, and where they don't. That line is what we measure.

The thesis

Each model generation absorbs more of what makes legal work hard: reading a contract the way a practitioner does, spotting the interaction between provisions rather than clauses in isolation, and drafting language a counterparty will actually accept.

We pressure-test those capabilities on genuine legal tasks, harden the ones that hold up — with playbooks, structured review, and verification — and ship them as products. What doesn't clear the bar stays in the lab.

How we evaluate
01

Real contracts, real stakes

Evaluation sets are built from genuine agreement types — NDAs, SaaS agreements, redemption and release instruments — with the traps practitioners actually see, not textbook hypotheticals.

02

Blind attorney panels

Practicing attorneys grade AI output and human output on the same contracts, against the same rubric, without knowing which is which.

03

A bar, not a benchmark

The question is never “did the model do well?” It's “did it meet the standard of the lawyers who do this work today?” A capability ships only when it clears that bar.

Current evaluations

Issue-detection evaluations against a blind attorney panel are underway across our first three contract families. Results are published here when the panels complete — numbers we'd stake the name on, or nothing.

Issue-detection evals · blind panel
Mutual NDA setin progress
SaaS agreementsin progress
Redemption & releasein progress

The first capability that cleared the bar: redline.

AI contract analysis and negotiation drafting, entirely over email.

Open redline