Pakistan's judiciary nets $38.50 per dollar invested by training judges on an AI legal assistant
Pakistan's Trial Courts (District Judiciary) deployed Large Language Models & Generative AI for Legal Research & Case Law in Litigation & Disputes. As reported by the-decoder.com: $38.50 saved per $1 invested return on investment.
Source-reported figures — cited source: the-decoder.com
What Pakistan's Trial Courts (District Judiciary) was trying to fix
Pakistan has fewer than two judges per 100,000 residents (versus 22 in the EU and 30 in England and Wales), and by the end of 2024 had a very large backlog of pending cases, with the vast majority in trial courts. Judges work with minimal technology and no support staff, and before the study only about 25% had ever used a large language model like ChatGPT.
What Pakistan's Trial Courts (District Judiciary) deployed
Researchers from ETH Zurich, Imperial College London, and the New Economic School ran a randomized field experiment across 1,559 judges in 118 courts (roughly half of Pakistan's trial court judges), testing JudgeGPT, a GPT-4-based assistant using retrieval augmented generation over 129,235 documents (128,292 rulings and 943 laws) to generate cited answers. One group received the tool plus six 90-minute training sessions taught by ETH Professor Elliott Ash; a second group got the same tool access but only a general seminar; a control group attended the seminar with no AI access.
Results
Trained judges used JudgeGPT about four times as much as the untrained-but-equipped group (nearly 60 logins and 200+ prompts after 40 weeks versus about 20 logins and fewer than 50 prompts). Districts with more trained judges resolved roughly 1,848 more cases per year (a 6.3% increase) at moderate exposure, with even bottom-quartile districts clearing about 616 more cases. Researchers estimate a return of about $38.50 saved per dollar invested (at least $10 under conservative assumptions), while ruling quality, appeal rates, and work hours held steady or improved, with no detected increase in gender or religious bias.
Key Takeaways
- Tool access alone barely moved usage or outcomes — targeted training on which tasks suit the AI (and which don't) was the real multiplier, driving substantially higher usage and measurable case-resolution gains.
- Training shifted judges toward lower-risk tasks like editing and summarizing and away from open-ended legal questions where hallucination risk is higher, while keeping final decisions in human hands.
- The study ran on GPT-4, a pre-reasoning model, suggesting today's more capable models could yield even larger productivity gains.
Evidence for Pakistan's Trial Courts (District Judiciary)'s Legal Research & Case Law deployment
- Reported outcome metrics
- 3 cited below
- Cited source
- the-decoder.com
- Last updated
- Source published
- Source link checked
Explore Related
Details
- Industry
- Litigation & Disputes
- Use Case
- Legal Research & Case Law
- AI Technology
- Large Language Models & Generative AI
- Company Size
- Enterprise
Have a similar implementation?
Share your customer's AI results and link it to your vendor profile.
Submit a case study →