Vendor-reported figures — source: the-decoder.com
Pakistan has fewer than two judges per 100,000 residents (versus 22 in the EU and 30 in England and Wales), and by the end of 2024 had a very large backlog of pending cases, with the vast majority in trial courts. Judges work with minimal technology and no support staff, and before the study only about 25% had ever used a large language model like ChatGPT.
Researchers from ETH Zurich, Imperial College London, and the New Economic School ran a randomized field experiment across 1,559 judges in 118 courts (roughly half of Pakistan's trial court judges), testing JudgeGPT, a GPT-4-based assistant using retrieval augmented generation over 129,235 documents (128,292 rulings and 943 laws) to generate cited answers. One group received the tool plus six 90-minute training sessions taught by ETH Professor Elliott Ash; a second group got the same tool access but only a general seminar; a control group attended the seminar with no AI access.
Trained judges used JudgeGPT about four times as much as the untrained-but-equipped group (nearly 60 logins and 200+ prompts after 40 weeks versus about 20 logins and fewer than 50 prompts). Districts with more trained judges resolved roughly 1,848 more cases per year (a 6.3% increase) at moderate exposure, with even bottom-quartile districts clearing about 616 more cases. Researchers estimate a return of about $38.50 saved per dollar invested (at least $10 under conservative assumptions), while ruling quality, appeal rates, and work hours held steady or improved, with no detected increase in gender or religious bias.
Have a similar implementation?
Share your customer's AI results and link it to your vendor profile.
Submit a case study →