Pakistani Judges Give Their Verdict on JudgeGPT
OpenAI
A large-scale trial in Pakistan found that a custom AI tool for judges, combined with training, boosted case resolutions by 6.3% without reducing decision quality. The tool, built on OpenAI's GPT-4 and Pakistani legal data, was used by over 1,500 judges.
In Pakistan, a trial involving 1,559 trial judges tested a custom AI tool called JudgeGPT, which combines OpenAI's GPT-4 with a database of 128,292 judicial opinions and 943 statutes. The tool helped judges with legal research and drafting, leading to a 6.3% increase in cases resolved per district, with no corresponding drop in decision quality. The study, led by economist Sultan Mehmood, also provided training to some judges, which proved crucial: judges who received JudgeGPT-specific training logged in 56 times and sent 212 prompts on average, compared to 10 logins and 25 prompts for those with generic training, and those with no training often abandoned the tool. The researchers used retrieval-augmented generation to reduce hallucinations, and they assessed judgment quality by having OpenAI's GPT-5-mini compare pre- and post-training judgments, with experienced lawyers agreeing with the model 70.6% of the time. While the tool saved time for judges, some concerns remain about the quality of justice and the risk of over-reliance on AI.
- Abbreviations
- LLM = Large Language Model — большая языковая модель
- RAG = Retrieval-Augmented Generation — генерация с дополнением поиском
Source: IEEE Spectrum AI —
original
