TechNewsReel
Live

LLMs excel at factual extraction but struggle with expert economic judgment

A study of IMF reports reveals that while AI can automate routine document review, it remains over-optimistic compared to human economists.

TechNewsReel Newsroom · August 16, 2026

Frontier large language models (LLMs) can reliably extract facts from complex financial documents but lack the nuanced judgment required for professional economic evaluation. This gap suggests that while AI can accelerate policy surveillance, it cannot yet replace the critical eye of a human expert.

Researchers Ganum and Atashbar tested the alignment between LLMs and human economists by analyzing 543 IMF Article IV staff reports spanning advanced, emerging, and low-income economies. The study compared various GPT-family models against human benchmarks, revealing a stark divide between factual accuracy and qualitative assessment. In 2024, LLMs achieved exact match rates of 76% to 81% on binary yes/no factual questions. However, the predicted probability of a model matching a human dropped to approximately 7% for complex rating questions.

The Evolution of Model Accuracy

Performance improved significantly across model generations. According to the researchers, GPT-5.5 reached 77% rating accuracy in 2024, a notable increase from the 59% accuracy recorded for GPT-4o. Stronger models generally reached between 71% and 75% accuracy on qualitative ratings, with 91% to 98% of those ratings falling within one point of the human benchmarks.

Despite these gains, the study identified a systemic bias in AI evaluations. LLMs exhibit a tendency toward "over-optimism," frequently assigning higher and less dispersed ratings than their human counterparts. This suggests that AI tends to view the quality of macrofinancial coverage more favorably than professional economists do.

Implications for Policy Institutions

This research is part of a larger effort to determine if generative AI can move beyond simple text summarization to support expert judgment in technical, multi-step professional reviews. The findings point toward a "human-machine complementarity" model for public institutions. By automating the routine extraction of data and factual verification, LLMs can significantly reduce the time economists spend on initial document reviews.

However, the inability of AI to replicate the depth and integration of human analysis remains a critical hurdle. In high-stakes surveillance, where contextual judgment is paramount, the risk of AI over-optimism could lead to overlooked vulnerabilities if left unchecked.

The Path Forward

As models continue to evolve, the gap in qualitative alignment may narrow, but the fundamental role of the economist remains secure for now. As Ganum and Atashbar noted, LLMs are becoming good enough to support routine extraction, but they are not yet dependable enough to own difficult judgment. Future surveillance workflows will likely integrate AI as a first-pass filter, leaving the final, critical evaluation to human experts.

Sources

Get a notification when a big story breaks. A few a day at most — no spam.