We ran a buy-side deal this summer that included summarizing about 50 of the target’s customer contracts. Because our own homegrown contracts AI isn’t built for summarization, we used GPT-5.5 Pro to do the first pass. GPT-5.5 was the best available OpenAI model at the time we did the review, and “Pro” was its most “think-ey” setting. We then had a Biglaw-experienced corporate lawyer heavily review the outputs. We were surprised by what we found, and thought others might be interested. If you are, read on. Or just flip through the deck. This piece just gives color to the concrete findings in there.

TL;DR: The summaries came back looking great. Then we checked them against the actual agreements, and results looked a lot more mixed. There were a lot more mistakes than we expected, including clauses it missed entirely and (more troublingly!) answers GPT appeared to have simply made up.

Numbers are in the deck. Here’s some context on how we got them.

GPT-5.5 Pro contract summarization deck, slide 1 of 6

Background on the project

One of our flagship offerings at Zuva is diligence reviews, where we pair experienced lawyers with AI to deliver high-quality work product (sell-side pre-diligence reports, disclosure schedules, buy-side diligence reports). Sell-side and disclosure schedule work tends to be more presence/absence reviews (e.g., which agreements require consent on change of control or have restrictive covenants). Buy-side is primarily summarization, with some presence/absence too. This was a buy-side project, and involved a lot of summarization.

Zuva Analyze (ZA), our proprietary contracts AI, runs on the underlying Kira tech (which we got to keep a copy of when we exited Kira) and improved a fair bit since. We find ZA good quality and reliable, but it isn’t built for summarization. So when a project needs summaries, we run documents through ZA and also through GPT or Claude, and use ZA to help check the generative results.

In this project, that meant GPT-5.5 on the Pro setting, i.e., OpenAI’s most advanced model at the time, on its most think-ey setting. (For 50 contracts, this first pass was $75–150 worth of token spend.) This is a fancier setup than most people doing this kind of work would use, and we used a lot more tokens than most people would.

We were searching for eight items across the 50 agreements: indemnification, limitation of liability, an industry-specific term we can’t name here, term and customer termination for convenience, warranty period, IP ownership, change of control, and restrictive covenants.

The primary checker of the summaries was a Biglaw (and AI) experienced corporate lawyer, supported by another heavily experienced teammate. All in, we spent several day of high-end human labor on the project, mostly on the checking (though some on checking other portions of the review beyond this commercial contract part).

Note that some details in this piece are anonymized, since the underlying project is confidential.

What reviewing results found: some things were right, some weren’t

Some results were fine, but a surprising number were not. Making things harder, mistakes were often difficult to spot, because GPT gave results that looked right. They just weren’t right once we went back to the underlying agreements.

We found some fabrications. These unsettled us. For example, GPT told us one contract was terminable on 15 days’ notice. In fact, the customer could terminate at any time on written notice.

Then there were misses. Change of control stands out, even though the numbers look okay. Two of the 50 contracts had a change of control provision, and GPT found one. Admittedly, this missed one was non-standard. Still, on an M&A deal, change of control is pretty important to get right.

Then there were clauses GPT found but described incorrectly. On limitation of liability, it found the clause in 72% of documents and got the summary seriously wrong in 58% of those. An example was a summary reporting a $3M liability cap that turned out to be a $3M insurance requirement.1 There is a big difference between having a maximum downside of $3M (a benefit) and having to carry $3M of insurance (an obligation).

Restrictive covenants scored 100%, but that result doesn’t say much. It turned out that none of these agreements had a restrictive covenant of the type we were looking for, so there was nothing there to miss.

One number in the deck looks worse than it is. The Term row shows an 84% miss rate, but this wasn’t actually too bad. We asked for “term and termination for convenience” and GPT gave us just termination for convenience. This mistake was easy to spot. We asked GPT to re-run and pull the term as well, and it did that fine.

Many of the misses were harder to find than that. Our reviewers found them by going back to the original agreements by hand. On the positive side, once we told GPT it had made mistakes, it was able to go back through its own work product and turn up more of them. Those results still needed further checking, and the process cost significant additional time and token use, but it did work.

If a junior lawyer had made some of these mistakes, I’d be inclined to fire them. Not for the misses, since everyone misses things, but rather for inventing facts, even in spots where they made no material difference. For diligence purposes, there is zero difference between 15-day termination and immediate termination. But made-up numbers are disconcerting, especially when they look like they could be real, since that means every number has to be manually confirmed.

You’re hired?

Despite that I would likely fire a junior lawyer who made these fabrications, GPT saved us a lot of time on this project, and we will use it again.

Doing this manually, summarization like this likely would have taken 1–3 hours per 20–50 page contract, so 50–150 hours across the pool. We spent a few days of senior time on the whole project including all the checking, and that project covered more than just these customer contracts. The AI saved us a lot of time despite being unreliable, and it was relatively quite cheap given the time savings.

While using GPT saved us time, these results needed much heavier human checking than I suspect many people doing this work would think to do. The summaries looked fine at first glance. If we hadn’t looked closer, we would have delivered pretty poor quality work.

This lines up with what we also experienced over the summer as we looked at Harvey’s synthetic LAB data rooms: it’s easy to get overimpressed by GenAI output, because at first glance it looks like what a good human would produce in a fraction of the time, and it’s only on a closer look that the mistakes turn up.

That matters beyond our project, because lawyers and others are using GPT, Claude, and similar models on a fair amount of contract review work these days, both directly as well as inside other tools. (The AI in many legal AI tools is primarily from APIs from GPT, Claude, and Google, now with some home built models perhaps mixed in.) Realistically, most of that work is often running at something less than the top model at maximum thinking, because that’s expensive. We ran the top model at maximum thinking, and even then the result quality had room for improvement.

Some possible objections to our findings

Here are some responses to issues I thought some might raise. I welcome other objections!

Our prompts were bad. Maybe! Prompts can almost always be better. We know it is possible to get good accuracy out of GenAI on contract review tasks. Nonetheless, this review was led by an ex-Biglaw corporate lawyer who has reviewed a lot of contracts with AI support. Could the lawyer/prompter have done better? Probably. Would many lawyers do better than them? Maaaaaaybe.

This experience was with GPT-5.5, but GPT-5.6 or Astra (or Claude, or Kimi, or whatever) would have done better. Fair enough. There are always newer and other models, and they keep improving. It’s totally possible different models would yield better results. Model performance tests show real differences across models across tasks. Both OpenAI and Anthropic have hired senior people recently to focus on the legal vertical, and part of this may be improving performance on contract review tasks. That said, we have been hearing for a while about how good GPT is at contract review, and our experience here was more mixed. For example, Aaron Levie (Box’s impressive CEO) posted in February 2025 that Box measured a 19-point improvement from GPT-4.5 over 4o on single-shot data extraction, tested on CUAD, a set of 510 commercial legal contracts. He posted a 27-point jump for GPT-4.1 two months later. On GPT-5 he wrote that hallucinations had come down dramatically and called this critical for legal work. And on GPT-5.5 specifically, the model we used here, he reported a 10-point accuracy jump on Box’s complex knowledge work evals. To be fair to Levie, single-shot field extraction on a public dataset is a different task than what we did, and I have no reason to doubt any of those measurements (though I don’t think it’s a great idea to test on CUAD, which could easily be in LLM training sets). But that’s sort of the point. Every release brings real, measured improvement on contract data extraction, and we still got these results on a live deal.

GPT caught some of its own mistakes once we pointed out that it had made mistakes, and an agentic setup could automate that loop. True, and an automated re-check could help. But the loop only started because a human read the underlying agreements and found the first errors. Nobody told the lawyer where to look. In fact, the AI results looked good initially. And the AI didn’t catch all its initial mistakes (though we weren’t testing to figure out what proportion of its own mistakes it caught).

We’re biased. We have been building contracts AI since 2011, so maybe we’re shading numbers to make our tech look better in comparison. We like our own tech, and think it’s great for presence/absence work. But it doesn’t do summarization, which is what we evaluated here. And we were evaluating because we used GPT-5.5 in the course of delivering client work, not because we set out to write this up. If GPT-5.5 made less mistakes it would help us, in that we would have to spend less time checking its work.

Our numbers are off. Possible. These are the errors we caught. We’re pretty careful at this work, and our accuracy has come out well in GC-run tests, but we could have made mistakes. Note, though, that if our final work product missed additional items, that means the real GPT accuracy numbers are worse than what we’ve reported, not better. Due to the confidential nature of the underlying project, we can’t share the underlying results for you to check yourself, which is a real limitation of this write-up.

Final thoughts

We didn’t do this project intending to write up results. But the results surprised us. Ideally you’ve found this interesting.

What makes the results compelling to us is that this is a real “Level 4” test, i.e., testing AI against real work product in a real situation. (“Level 4” comes from a recent post we did on the four levels of testing legal AI.) These summaries came about in the course of reviewing contracts on a live M&A deal, not from a test we put together. It’s only one deal, and only 50-ish contracts. But Level 4 results are rare, mostly because confidentiality usually makes them impossible to talk about at all, and we think that makes even a single deal’s worth of them worth sharing.

———–

Thanks to Adam Roegiest and Susan Fox for comments on this piece.


  1. Note that this $3M number has been slightly changed from the original, in an effort to be extra careful on confidentiality (since this was a client project). ↩︎

Share this article: