Return
A multi-stage agentic AI system for extracting information from large digital archives: case study on the Czechoslovak year 1968 in CIA's FOIA collection
J
K
L
DOI:10.1108/EL-06-2025-0272.png)
Abstract
En 中文
Purpose - This study aims to design, implement and evaluate a conceptual multi-stage artificial intelligence (AI) system for the systematic analysis of large, unstructured digital archives. Using the 1968 Prague Spring and subsequent Soviet invasion of Czechoslovakia as a case study, the paper demonstrates how such a system can automate the extraction of historical intelligence from declassified documents, creating a time-resolved narrative from non-machine-readable sources. Design/methodology/approach - A multi-stage agentic system comprising eight specialized agents was developed to deconstruct the historical research workflow. The system was applied to the corpus of declassified President's Daily Briefs from 1968 to 1969, sourced from the CIA's FOIA Electronic Reading Room. The methodology integrates optical character recognition (OCR) and expert-guided prompt engineering and introduces a novel evaluation framework to quantitatively and qualitatively compare the performance of four distinct LLMs (GPT-5, Claude Sonnet 4.5, Grok 4 and Magistral Medium) across the 2,122-page corpus. Findings - The system produced three key outputs: a comprehensive monthly summary of intelligence reporting, a structured list of key named entities and a thematic quantification of the content. Critically, the comparative analysis reveals significant trade-offs in performance: GPT-5 achieved the highest output quality (F1 score: 0.731), while Claude Sonnet 4.5 offered superior cost-efficiency and processing speed. Moreover, Claude Sonnet 4.5 and Grok 4 demonstrated flawless operational stability, while Mistral Magistral Medium proved most effective at text reduction. These findings underscore that while AI enhances efficiency, expert human oversight remains essential for ensuring interpretive nuance.Research limitations/implicationsA primary limitation is the data acquisition process, due to the lack of a public API for the canonical data source, which affects long-term reproducibility. Furthermore, the reliance on OCR introduces a layer of potential error into the source text. The study implies that fully automated historical analysis is not yet fully feasible; rather, a human-in-the-loop, collaborative approach is essential for credible results. Practical implications - The proposed framework provides a replicable model for historians, archivists, librarians and intelligence analysts to unlock insights from vast, unstructured document collections. It streamlines labor-intensive tasks (e.g. data discovery, text extraction, summarization), allowing researchers to focus on higher-level analysis and interpretation. Originality/value - This paper's primary novelty lies in presenting one of the first comprehensive frameworks for evaluating and benchmarking competing LLMs on complex historical analysis tasks using declassified intelligence documents. It moves beyond single-LLM case studies by offering a complete, end-to-end workflow that not only processes intelligence documents but also provides a replicable methodology for assessing the practical trade-offs (quality, cost and speed) between different AI models in a digital humanities' context.
Keywords:
Artificial intelligence
History
Cold war
Czechoslovakia
Multi-agent system
Large language models
Central intelligence agency
Freedom of information act
President's daily brief
Soviet invasion
Journal
E
IF:
1.5
Papers:
53
Citations:
1.2K
