TEI 2026
Creating Connections, Unsettling Practices
August 10-14, 2026
University of British Columbia, Vancouver, BC, Canada
Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
|
Daily Overview |
| Session | |
|
Poster Session Lightning Talks Location: Koerner 5th floor common area Session Chair: Syd Bauman, Northeastern University | |
| Presentation 19 | |
ID: 155
/ Poster Session Lightning Talks: 19
Late-breaking Poster Keywords: LLM assisted encoding, TEI P5, Workflows Unsettling Manual Encoding and Creating Connections in Historical Natural History Texts: LLM-Assisted Entity Annotation in a TEI Edition of Eighteenth-Century Texts Academy of Sciences and Humanities in Lower Saxony, Germany This poster presents a series of experiments in LLM-assisted TEI P5 annotation carried out on Johann Friedrich Blumenbach’s “Handbuch der Naturgeschichte” within the broader context of Blumenbach-Online, a long-running digital edition and cultural heritage project at the Academy of Sciences and Humanities in Göttingen. The project has demonstrated both the scholarly value and the practical limits of deeply manual TEI workflows: while expert encoding ensures high semantic precision, it is costly, slow, and difficult to scale across large corpora of editions, translations, and related materials. Therefore we argue for hybrid workflows combining expert review with computational assistance. Our poster concentrates on experiments in which a selection of large language models was used to support Level 5 semantic markup of an eighteenth-century German text from the Natural History domain. The task included identifying and encoding persons, places, and natural-historical objects in TEI P5 XML; assigning authority links for persons (GND) and places (Getty TGN); generating stable IDs for recurring natural-historical entities and validating output against a project-specific Relax NG Compact schema. The quality of the outputs was assessed using a multi-dimensional evaluation framework against the human-guided annotations. Experiments showed that the LLMs were useful in their ability to follow custom TEI schema rules, given the proper few-shot examples. Zero-shot prompting however, led to non-compliant XML outputs. At the same time, the workflow also exposed typical weaknesses of generative systems: limited context windows, overfitting and occasional hallucinatory references. The results suggest that LLMs are most productive not as autonomous encoders, but as interactive editorial assistants embedded in a tightly constrained TEI workflow. In this model, LLMs improve recall and accelerate exploratory markup, while scholarly supervision remains essential for validation, disambiguation, and epistemically responsible encoding. We argue that such hybrid workflows can create help new connections and encoding in efficient and productive ways. | |
