Discussion about this post

User's avatar
Jim Clifford's avatar

This post raises a key question: how do we create a diverse benchmark to test LLMs on really hard historical thinking? AI systems perform well on topics with deep secondary literatures, but at the edge of historical knowledge, where the evidence is fragmentary, contradictory, or absent, they tend to confabulate confident answers rather than acknowledge uncertainty. How do you benchmark nuance and the ability to say ‘we don’t know’? And on a more practical side, how do we populate a benchmark with new unpublished historical research to test the models when we also want to publish our historical research (decaying the benchmark)?

1 more comment...

No posts

Ready for more?