TEI 2026
Creating Connections, Unsettling Practices
August 10-14, 2026
University of British Columbia, Vancouver, BC, Canada
Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
|
Daily Overview |
| Session | |
|
A1 Long Papers: TEI Futures Location: BUCH B208 Session Chair: Raff Viglianti, University of Maryland | |
| Presentation 1 | |
ID: 119
/ A1 Long Papers: 1
Long Paper Keywords: Text Encoding Initiative, large language models, low-resource languages, editorial bias, open data Commentary as Infrastructure: The Algorithmic Afterlives of TEI Annotation Aarhus University, Denmark Digital scholarly editions are becoming AI training data. Grundtvig's Works, a comprehensive digital scholarly edition of the published writings of N.F.S. Grundtvig (1783–1872), exemplifies this shift. Close to 700 texts are currently published, and all materials are released under a CC0-license. In 2025, the project uploaded 632 of these TEI-encoded texts to Hugging Face, positioning the edition not just as a reader-facing publication but as research infrastructure for historical Danish language processing (Danish Foundation Models 2025; Enevoldsen et al. 2025). The project is now preparing to release its commentary layer: 160,000 explanatory notes across more than 21,000 pages that gloss archaic terms and contextualize nineteenth-century usage. Drawing on empirical analysis of the commentary layer’s structure, density, and editorial logic (Vad, Rasmussen & Baunvig 2024; Vad, Rasmussen & Baunvig forthcoming), this paper asks what happens when scholarly annotations become algorithmic infrastructure and proposes an infrastructure-aware editing approach. For a low-to-medium resource language like Danish, where historical language data is scarce, these annotations could carry outsized importance. While one edition is modest, multiple annotated editions could significantly expand training data for historical Danish. Yet annotations are not neutral metadata: they encode editorial judgment about what requires explanation and how to frame context. When a note explains “højskole” – a term Grundtvig transformed from “university” to folk high school – it teaches language models interpretive frameworks shaped by scholarly tradition and institutional priorities. This creates both opportunity and risk. Commentaries could help AI systems understand historical language and avoid hallucination. Yet they transmit editorial biases and where alternatives are scarce, such bias is amplified. The structure of digital cultural heritage reflects this asymmetry: ‘cultural saints’ like Grundtvig receive comprehensive editions with rich annotation, while ‘textual masses’ – newspapers, marginalized voices – remain poorly digitized (Baunvig 2023). This creates what Lassen et al. (2025) term a “vicious cycle of silencing”, where biases in data collection compound through the data science pipeline, systematically excluding certain groups as legitimate knowledge sources (p. 11). Current discussions of AI in digital textual scholarship focus predominantly on what language models can do for editions: accelerating transcription, suggesting collations, identifying patterns (Pollin 2025). This paper reverses the lens, examining what happens when editions become training data for AI, with particular attention to how explanatory notes function as a mediating layer between historical text and algorithmic systems (Vad et al., forthcoming). It pursues two interconnected questions:
This positions TEI as a site for ethical AI development in multilingual contexts. The paper proposes infrastructure-aware editing: a reflexive open-science model that releases annotations while documenting the editorial frameworks that shaped them, making curatorial decisions legible to AI researchers rather than invisible in training pipelines. | |
