U

Unnamed Entity

Immigration legal AI platform cuts lawyer review time 85% and ships agents 3x faster with Judgment Labs evaluators

Curated & reviewed by Peter Korpak, Founder & Chief Analyst, 100SignalsHow we verify
85%+Lawyer Review Time Reduction
100+ hoursMonthly Hours Saved
3x fasterAgent Deployment Speed

Vendor-reported figures — source: www.judgmentlabs.ai

Anonymous Immigration Legal AI Platform
Metric Before After Impact
Document quality identification accuracy 52% (baseline) 97% 87% improvement, matching lawyer-level accuracy
Lawyer review time per document Baseline (100%) 15% of original 85%+ reduction
Monthly hours saved 0 100+ hours/month 100+ hours saved per month
Agent deployment cadence 1x 3x faster 3x improvement, 2 releases shipped 3 months ahead of schedule

The Challenge

Immigration law is documentation-intensive by nature: O-1A, H-1B, and green card petitions require precise, citation-accurate drafts where a single factual error can delay or sink a case. To handle growing caseload volume, this startup built agentic workflows that generate rough document drafts for lawyers to review and finalize. The model worked — until iteration velocity became the enemy. Every update to prompts, underlying models, or retrieval tools required lawyers and engineers to manually compare old and new outputs side by side. The process was inconsistent, dependent on who happened to be reviewing, and expensive. Post-deployment, regressions still slipped through, surfacing as quality failures that consumed even more lawyer time downstream.

The Solution

Judgment Labs addressed the regression testing gap by building custom post-trained LLM evaluators trained directly on the platform's internal data — specifically, anonymized pairs of AI-generated rough drafts and lawyer-finalized versions. The final drafts encoded the systematic mistakes the writing agents made at scale; Judgment used LLMs to bootstrap reasoning traces explaining each lawyer revision, ran ablations to filter for high-quality traces, then distilled that reasoning into a judge model via post-training. To further improve reliability, they deployed an LLM-as-jury system: N evaluators with different base models and post-training ablations vote on document quality, with weighted majority vote as the final decision. The judge was integrated into an automated regression testing pipeline via Judgment Labs' judgeval package, and later repurposed as an Agent Behavior Monitoring (ABM) system to triage low-quality documents before they reach lawyers.

Results

Judgment's model system correctly identified the higher-quality document 97% of the time on hold-out sets, compared to a 52% baseline — matching or exceeding lawyer-level accuracy. That evaluator performance translated directly into operational gains:

  • 85%+ reduction in lawyer review time per document
  • 100+ hours saved per month across the full caseload
  • 2 new agent releases shipped 3 months ahead of schedule
  • 3x faster agent deployment cadence post-integration
  • 20% more caseload handled with the same team size

The ABM monitoring layer, live for only three weeks at publication, had already flagged over 40 cases containing factual contradictions, misquoted citations, and misconstrued evidence — issues that previously reached lawyers undetected.

Key Takeaways

  • Training an LLM judge on real expert edit pairs (rough draft → finalized draft) captures domain-specific quality signals that generic out-of-the-box evaluators consistently miss.
  • An LLM-as-jury ensemble with weighted majority voting meaningfully reduces evaluator variance and improves alignment with human expert judgment at scale.
  • The same judge model built for regression testing can be repurposed as a production monitoring system — eliminating the need to build separate quality infrastructure.
  • Automated evaluation pipelines remove lawyers from the critical path of engineering iteration, compounding throughput gains over time without adding headcount.

Share:

Details

Company Size
Startup
Company
Unnamed Entity
Quality
Curated
Last verified
Jul 28, 2026

Have a similar implementation?

Share your customer's AI results and link it to your vendor profile.

Submit a case study →