U

Unnamed Am Law 100 Firm

Am Law 100 firm cuts document review time by 67% with GenAI on 126,000-document government investigation

Curated & reviewed by Peter Korpak, Founder & Chief Analyst, 100SignalsHow we verify
50–67% reduction (one-quarter of the personnel)Document Review Time Reduction
126,000 documents in under 24 hoursDocuments Coded
90%+ accuracy, at or above first-level attorney reviewer benchmarksAI Accuracy Rate

Vendor-reported figures — source: www.everlaw.com

The Challenge

In fall 2024, a three-attorney team at a leading Am Law 100 firm faced a defining e-discovery challenge: reviewing 126,000 documents for production in a large-scale government investigation tied to potential civil litigation, under a compressed timeline and constrained budget. The document set spanned emails, email attachments, and Microsoft Teams messages requiring coding for responsiveness and tagging across nearly two dozen distinct issue codes. At the time, document review represented the dominant cost driver in litigation — accounting for nearly 80% of total litigation spend. The conventional approach would have required approximately 20 contract attorneys over four weeks, a staffing and cost profile the matter could not support.

The Solution

The firm deployed EverlawAI Assistant Coding Suggestions, an LLM-powered e-discovery tool built on large language model technology, in partnership with litigation managed services provider Right Discovery. The implementation followed a structured three-stage workflow: first, developing initial code criteria with full case context at the code, category, and case level; second, iterating prompts against three sample document subsets to validate accuracy before scale deployment; third, running Coding Suggestions across all 126,000 documents once validation thresholds were met. The tool analyzed each document against natural-language instructions and returned a four-tier classification — Yes, Soft Yes, Soft No, No — along with a written rationale for each coding decision, enabling attorneys to concentrate human review only on the uncertain middle tiers rather than the full corpus.

Results

The team coded 126,000 documents in approximately one day with five team members — compared to the 20 contract attorneys over four weeks that a traditional managed review would have required. Key outcomes:

  • 50–67% reduction in document review time, achieved with one-quarter of the personnel
  • 90%+ accuracy, meeting or exceeding first-level attorney reviewer benchmarks on precision and recall
  • Precision score of 0.63, recall of 0.85, F1 of 0.72 on the final validation set — comparable to established TAR benchmarks
  • Greater consistency in coding decisions than human first-level reviewers
  • Prompts and code criteria are reusable for future productions in the ongoing investigation

Key Takeaways

  • Prompt quality is the primary determinant of AI accuracy in LLM-based document review — the 15-hour prompt development and iteration investment was what made processing 126,000 documents in under 24 hours defensible, not just fast.
  • A four-tier classification (Yes/Soft Yes/Soft No/No) is operationally superior to binary coding: it lets teams concentrate human review on uncertain documents while treating high-confidence results as settled.
  • Validating AI performance using the same precision, recall, and F1 metrics applied to TAR provides an auditable, court-defensible record of review quality before scaling to the full document set.
  • Coding criteria and prompts built for one production become durable assets — directly reusable when a dormant investigation resurfaces with new document requests.

Share:

Details

Company Size
Enterprise
Company
Unnamed Am Law 100 Firm
Quality
Curated
Last verified
Jul 28, 2026

Have a similar implementation?

Share your customer's AI results and link it to your vendor profile.

Submit a case study →