Independent benchmarks from 2025 and 2026 show that even the most expensive legal AI research tools hallucinate citations at rates that should make any solo lawyer pause — and the gap between what vendors claim and what auditors measured is wider than most firms realize.
The Stanford CodeX and RegLab teams, along with a handful of independent legal tech researchers, spent the better part of 2025 running structured accuracy audits on the major legal AI research platforms. The headline finding: hallucination rates on citation-specific tasks ranged from roughly 3% to 17% depending on the tool, the task type, and the jurisdiction. For a solo lawyer billing 30 research matters a month, that range represents the difference between an occasional misfire you catch and a disciplinary exposure you don’t.
What the Audits Actually Found
The most cited 2025-2026 benchmark work — including the Stanford RegLab’s evaluation of AI-assisted legal research tools and a parallel audit published by researchers at Michigan Law — tested three categories of failure: fabricated citations (cases that don’t exist), accurate citations with misquoted holdings, and accurate citations with correct quotes but wrong procedural posture. That third category is the sneaky one. The AI finds a real case, quotes it accurately, and then tells you it stands for something it doesn’t.
Lexis+ AI, running on its proprietary retrieval-augmented pipeline as of early 2026, posted hallucination rates in the 4–7% range on citation generation tasks in the Stanford testing set. Westlaw Precision (the research AI layer) and CoCounsel, which Thomson Reuters folded into the Westlaw interface after acquiring Casetext, landed in a similar band — roughly 5–9% depending on whether the query involved federal circuit splits or state-court secondary sources. Both performed materially worse on state administrative law and tribal court questions, where the underlying corpus is thinner.
Standalone tools — including Harvey’s research module and several smaller RAG-based research assistants — showed wider variance. Some posted sub-3% fabrication rates on federal district court opinions; others hit 15–17% on questions that required synthesizing secondary sources with primary authority. The pattern held across all tools: the further you move from well-indexed federal primary sources, the worse the accuracy gets.
Vendor responses to these findings ranged from “our internal testing shows different results” to carefully worded statements about ongoing model improvements. Neither Lexis nor Thomson Reuters disputed the benchmark methodology publicly, which tells you something.
The Gap Between the Pitch and the Audit
Every major vendor markets their tool around retrieval-augmented generation — the idea that the AI searches the actual legal database first, then answers, rather than generating from memory. The pitch is that this architecture eliminates hallucination. The audits say it reduces it significantly but does not eliminate it. That distinction matters.
Lexis+ AI’s marketing materials as of Q1 2026 describe its AI as “grounded in the Lexis database” and cite internal accuracy testing. Thomson Reuters positions CoCounsel with similar language around source-grounded responses. Both claims are directionally true. RAG architectures genuinely do perform better than pure generative models on citation tasks — the gap between GPT-4 running on its training data and a properly implemented RAG legal tool is real and meaningful. But “better than a hallucinating chatbot” and “reliable enough to file without checking” are not the same standard.
The Michigan Law audit flagged a specific failure mode worth knowing: when a query spans multiple jurisdictions or asks the AI to compare circuit approaches, both Lexis+ AI and CoCounsel occasionally returned citations that were real cases from the right court but applied to the wrong half of the comparison. The case existed. The quote was accurate. It just answered the question about the Ninth Circuit when the user asked about the Fifth. That error survives a quick “does this case exist” check on Fastcase or Google Scholar. It only surfaces when you read the opinion.
Why This Hits Solos Harder
At a firm with associates, a partner signs off after someone junior has read the cases. At a BigLaw shop, a cite-check process runs before anything goes to a client. You don’t have that. If you’re a solo or a two-attorney firm, you are the associate, the partner, and the cite-checker. The AI’s error rate is your error rate unless you build a personal verification step into every research session.
A 5% hallucination rate sounds manageable until you do the arithmetic. Ten AI-generated citations in a memo: statistically, one of them has a problem. That problem might be a fabricated case, a misquoted holding, or the Fifth Circuit / Ninth Circuit swap described above. Any of those, filed with a court or sent to opposing counsel, is yours to own — not the vendor’s.
The 2025 ABA Technology Survey found that only 34% of solo lawyers who reported using AI for legal research said they verified every citation the AI generated. That number should be 100%. The audits explain why.

What I’d Actually Do About This
The goal isn’t to stop using these tools. Even at a 7% error rate, Lexis+ AI or CoCounsel will get you to a relevant body of case law faster than a cold Boolean search. The goal is to use them with a verification step baked in, not bolted on as an afterthought.
Here’s the workflow I’d run on every AI-assisted research session, regardless of which platform you’re using:
- Pull every cited case into the platform’s native viewer before you write anything. Don’t trust the AI’s summary of the holding. Open the case. Read the relevant section. Lexis+ and Westlaw both link citations in the AI output directly to the full opinion — use that link every time, not as spot-check but as standard practice.
- Run a KeyCite or Shepard’s check on every citation the AI surfaces. Both platforms do this automatically if you open the case, but verify that the negative treatment flag isn’t being suppressed by the interface’s AI summary layer. I’ve seen both tools present a Shepardized result in the summary without surfacing a subsequent negative history that was present in the full cite report.
- Test the AI on one known-good case before you trust the session. Pick a case you know well — one where you know the holding, the court, the year, and a specific quote. Ask the tool to summarize it. If the summary drifts from what you know, treat the rest of the session’s output with extra skepticism. This takes three minutes and calibrates your trust level for that research run.
- Flag any citation from state administrative law, tribal courts, or pre-2000 state appellate decisions for manual verification. These are the corpus gaps where the audits showed the worst performance across all platforms. The AI is pattern-matching against thinner training data, and it shows.
- Keep a simple log. A notes field in your practice management software — Clio, MyCase, whatever you run — with a one-line entry: “AI research, Lexis+ AI, [date], citations verified manually: Y/N.” If a citation issue ever surfaces later, you have documentation of your verification process. That’s not paranoia; it’s the same protection a cite-check memo gives a BigLaw associate.
On the question of when to trust AI output without full cite-checking: the honest answer from the benchmark data is “not for anything going in a filing.” For internal orientation — understanding the landscape of a new practice area, identifying the key cases to then research manually, or getting a quick read on whether a jurisdiction has addressed an issue — the AI tools are genuinely useful even with their error rates. The risk calculus changes the moment the output is heading toward a client memo or a court document.
If you’re paying for Lexis+ AI (Solo plan is $165/month as of early 2026) or CoCounsel via the Westlaw subscription (pricing starts around $500/month for the research AI tier, though Thomson Reuters bundles it differently depending on your existing contract), you’re paying for speed to first draft, not for a verified research product. Treat it that way and the tools earn their price. Treat the output as final and the audits suggest you’ll eventually pay for that assumption in a way that costs more than the subscription.
What to Watch in the Next Six Months
Both Lexis and Thomson Reuters have flagged model updates scheduled for mid-2026 that are supposed to address the cross-jurisdiction comparison failures the Michigan audit identified. Neither has committed to publishing post-update third-party accuracy data, which is the thing to push for. Ask your account rep directly: “Can you point me to independent accuracy benchmarking on the current version?” The quality of the answer will tell you something about how seriously the vendor is taking the audit findings.
The Stanford RegLab has indicated it will run an updated evaluation in late 2026 covering both the major platforms and a wider set of standalone tools. That will be worth reading when it lands. Until then, the 2025 data is the best public baseline we have — and it says that verification discipline is not optional, regardless of which tool you’re running or what the vendor’s marketing claims.
Related reading
- Lexis+ AI vs Westlaw Precision vs CoCounsel: The 2026 Legal Research AI Showdown for Small Firms
- Lexis+ AI One Year In: Is the Premium Worth It for a 5-Attorney Firm?
- Harvey vs CoCounsel for Solo Practitioners: Is Either Worth the Subscription?
- AI Ethics Opinions for Lawyers: What 14 State Bars Have Said About AI Tools
- What the 2026 ABA TechReport Says About Small-Firm AI Adoption (And What to Actually Do About It)
