Researchers who tested on a large scale an artificial intelligence (AI) tool for judges in Pakistan found that the use of the system, combined with specific training, increased by 6.3% the number of cases resolved in first-instance courts, with no apparent drop in the quality of decisions.

According to the study coordinated by Sultan Mehmood, an economist at the New Economic School in Moscow, and collaborators, this is the first large-scale independent evaluation of the continuous use of AI by judges. The experiment, launched in 2024, offered the tool to 1,559 magistrates, about half of the country's first-instance judges.

Pakistan faces a backlog of 2.26 million cases and has fewer than two judges per 100,000 inhabitants, compared with 22 in the European Union and 8 in Brazil. To reduce the overload, the team developed JudgeGPT, based on OpenAI's GPT-4 model, with a knowledge base of 128,292 judicial decisions and 943 Pakistani laws. The system's responses include notes and links to the cases and laws consulted.

Of the total, 1,197 judges underwent six 90-minute online training sessions on how large language models work, their limitations, risks of bias and hallucinations, and the importance of verifying responses. Another 180 received generic technology training, and one group received no training.

Results

After nine months, the median district recorded a 6.3% increase in resolved cases, and the effect was greater where there were more trained judges. Appeal rates fell slightly, suggesting that faster resolution did not result in worse decisions. Trained judges accessed JudgeGPT an average of 56 times and sent 212 requests, compared with 10 accesses and 25 requests among those who had generic training. Without training, usage fell after about a month.

"Just giving people the technology doesn't make them use it persistently," Mehmood said. The researchers calculated that a trained judge began resolving 38.5 more cases per month, representing savings of US$ 38.50 in judicial costs for every dollar spent on operating the tool.

To assess the quality of rulings, the team asked the GPT-5-mini model to compare decisions by the same judge before and after training. The later ones were chosen 59% of the time. Two experienced Pakistani lawyers evaluated 90 pairs and agreed with the model in 70.6% of cases, close to their agreement rate of 73%.

The study found that about one-fifth of requests to JudgeGPT involved "substantial delegation," such as asking the tool to decide or draft the reasoning with little participation from the judge. Training reduced the proportion of this type of use.

Limits and risks

The study's co-author, Elliott Ash, a professor of law, economics, and data science at ETH Zurich, said the solution to hallucinations lies not only in smarter models, but in connecting them to tools that search for and verify sources. "There are risks in using these AIs, even with safeguards, but at some point you have to put judges in the strongest possible position," he said.

David Autor, a professor of economics at the Massachusetts Institute of Technology (MIT), said the productivity increase is "credible" and should improve with broader use of the tool. John Zeleznikow, a professor of law and technology at La Trobe University in Australia, however, cautioned that efficiency is not the only criterion for justice and that it is unclear whether the "quality of justice" improved.

More from Radar