Digital history did not fail in isolation—it failed just before the world discovered how much it needed it. As artificial intelligence systems attempt to interpret vast bodies of text, they confront problems of evidence, disagreement, and representation that historians have spent centuries learning to navigate. The future of AI may depend not on replacing the humanities, but on rediscovering what they know.
But here’s the problem. At the very moment when artificial intelligence is reorganizing the economy and the university, its central challenges are, in fact, historical ways of knowing. In computer science, tasks such as retrieval, benchmarking, and the establishment of ground truth all depend on the selection, comparison, and evaluation of documents within large corpora. Determining whether outputs correspond to reliable bodies of evidence is not a purely technical question. It is an interpretive one.
Historians have long developed methods for addressing precisely these challenges; it was these approaches that I review in the opening chapters of The Dangerous Art of Text Mining, showing how information scientists who attempt large-scale analysis without historians’ understandings of the bias of archives and algorithms has led to retracted journal articles and unsupportable findings. I argued that historians’ methods, paired with NLP strategies, could make a “smarter” data science, which matched inferences, algorithms, and sources .
Today, a growing body of work by historians shows that historians are on the forefront of the task of retooling AI to improve its retrieval mechanisms, confronting messy sources on historians’ terms. In “Data retrieval from local heritage books—Is artificial intelligence the solution?”, Robert Stelter and Rafael Biehler test LLM-based extraction against code-based and manual methods and show both the promise and the limits of AI when working with irregular historical books, reminding us that retrieval from the archive is never just a technical matter but a problem of source structure, error, and judgment. In Stewart Spencer Dean and Sanskriti Sinha’s “Retrieving information from unstructured historical sources using large language models”, the authors show that LLMs designed to mirror historians’ workflows can extract structured information from difficult historical documents more flexibly than older pipelines, while also underscoring how fragile reproducibility and transparency become when retrieval depends on opaque model behavior. In “Mapping the Latent Past: Assessing Large Language Models as Digital Tools through Source Criticism”, Daniel Hutchinson demonstrates how to evaluate LLMs as historians would evaluate sources: not only for fluency or accuracy in isolation, but for provenance, distortion, omission, and the ways they reshape historical memory. Taken together, these studies suggest that historians are not merely future users of AI tools; they are helping define what reliable retrieval, evaluation, and interpretation should mean in the first place.
The selection of the most relevant examples from a large corpus is an obvious example of an instance where historians turn to their own methods when selecting cases, assembling archives, and balancing typical and exceptional examples without collapsing their differences.
The ‘alignment problem,’ often framed as ensuring that AI systems produce trustworthy outputs, also presents a problem of managing agreement and disagreement across documents. At stake is not only accuracy but the structure of knowledge itself. As AI systems scale, they tend to normalize—to compress disagreement into dominant narratives, privileging what appears most frequent, coherent, or statistically central. Without explicit intervention, AI risks erasing this plurality, producing a flattened account of human experience that undermines democratic reasoning. Embedding historical methods into AI systems—methods attuned to comparison, contradiction, and the coexistence of multiple valid perspectives—ensures that technological scale does not come at the expense of epistemic diversity. In this sense, history is not simply a content domain for AI, but a necessary foundation for building systems that can represent the past, and reason about it, in ways that remain open, plural, and accountable.
Here, too, historians have solutions. Historians have long developed methods for identifying consensus, tracing dissent, and preserving plurality without reducing it to a single dominant narrative. They have worked on preserving small-scale differences of “little” details even within the large-scale arc of meaning-making typical of so-called “big history.” Historians have shown that what matters most in the record of the past is often not consensus but dissent: competing interpretations, marginalized voices, and unresolved conflicts that shape political and cultural change.
Beyond AI, the challenges of curating data extend across the disciplines that have engaged with data science. As scholars such as Arthur Spirling have noted, fields like political science increasingly rely on large-scale textual analysis but often lack robust frameworks for ensuring that their corpora are complete and representative.
This is not a trivial issue. To study labor history, one cannot rely on a single archive or even a handful of sites. One must assemble sources that reflect different classes, migrant and ethnic experiences, institutional perspectives, and temporal changes shaped by major events. The problem is not simply one of scale, but of completeness: bringing together heterogeneous materials into a coherent evidentiary base.
This challenge extends to law, medicine, sociology, and beyond. The interpretation of large textual corpora depends on assembling archives that are sufficiently comprehensive, balanced, and historically grounded. The expertise required to build and evaluate such archives—understanding what is missing, what is overrepresented, and how sources relate—has traditionally been concentrated in history.
The expertise, in other words, already exists. However, it is dispersed across institutions such as Saskatchewan, SMU, George Mason, Clemson, and Waterloo, rather than integrated into the central infrastructures of major research universities.
Yet the university as a whole increasingly depends on precisely this kind of knowledge. History can contribute to AI by providing methods for retrieval, evaluation, and the preservation of plurality. It can contribute to the world by enabling large-scale understanding of problems such as climate governance, where decades of negotiation must be analyzed across multiple actors and perspectives. It can contribute to other disciplines by offering frameworks for assembling and interpreting complete and representative archives.
The next step is clear. What is needed is not another generation of small projects, but a coordinated effort—a kind of archival and analytical moonshot—that brings together historians, computational experts, and domain specialists to build comprehensive, multi-source, multi-lingual corpora and the methods to interpret them.
Such an effort must be led by historians, because the central problem is not simply one of computation, but of judgment: what counts as evidence, what counts as completeness, and how meaning is constructed across time.
The digital breakthrough will happen when those questions are treated not as peripheral, but as foundational.




This post raises a key question: how do we create a diverse benchmark to test LLMs on really hard historical thinking? AI systems perform well on topics with deep secondary literatures, but at the edge of historical knowledge, where the evidence is fragmentary, contradictory, or absent, they tend to confabulate confident answers rather than acknowledge uncertainty. How do you benchmark nuance and the ability to say ‘we don’t know’? And on a more practical side, how do we populate a benchmark with new unpublished historical research to test the models when we also want to publish our historical research (decaying the benchmark)?