Can Frontier AI Build a Statistical Register? A Benchmark and Research Programme on Vietnam's Thermal Power Fleet
Minh Ha-Duong ()
Additional contact information
Minh Ha-Duong: CIRED - Centre International de Recherche sur l'Environnement et le Développement - Cirad - Centre de Coopération Internationale en Recherche Agronomique pour le Développement - EHESS - École des hautes études en sciences sociales - AgroParisTech - Université Paris-Saclay - CNRS - Centre National de la Recherche Scientifique - ENPC - École nationale des ponts et chaussées - IP Paris - Institut Polytechnique de Paris
Working Papers from HAL
Abstract:
Energy policy models need asset-level data that are complete, current, historical and prospective, and traceable to authoritative primary sources. Yet across much of the world these data arrive late or incomplete. We introduce a computable benchmark for this task: reconstruct the register of Vietnam's 177 thermal power plants and score the result against a hand-compiled, per-cell-sourced reference. We find that the task is hard for today's AI systems. Fourteen language models answering from memory recovered much less than the Wikipedia lists contain, even though these lists had presumably been in their training corpora for years. We observed that a reference-free signal, the within-run variability of the reported capacities, correlates with accuracy, letting us screen out weak runs without the reference. We then tested the commercial frontier of mid-2026: four agentic systems with web access and extended reasoning. Only one improved coverage over the memory-only query, and a multi-step harness degraded three of the four. A curated document set improved coverage, most for the weaker agents. Fusing the lists from multiple runs more than doubles recall. We conclude that an evergreen statistical-quality register requires deliberate knowledge engineering. We propose deriving datasets as dated snapshots from a sourced, auditable knowledge base, mechanically updated from a periodically harvested corpus, with humans in the loop to vet sources and resolve the hard tail.
Keywords: AI benchmark; Large language models; Agentic systems; Retrieval-Augmented Generation; Energy statistics; Knowledge engineering; Vietnam; Thermal power (search for similar items in EconPapers)
Date: 2026-06-16
Note: View the original document on HAL open archive server: https://hal.science/hal-05658462v1
References: Add references at CitEc
Citations:
Downloads: (external link)
https://hal.science/hal-05658462v1/document (application/pdf)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:hal:wpaper:hal-05658462
Access Statistics for this paper
More papers in Working Papers from HAL
Bibliographic data for series maintained by CCSD ().