TEI 2026
Creating Connections, Unsettling Practices
August 10-14, 2026
University of British Columbia, Vancouver, BC, Canada
Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
|
Daily Overview |
| Session | |
|
A2 Long Papers: TEI Infrastructures Location: BUCH D316 Session Chair: Elli Bleeker, Huygens Institute | |
| Presentation 2 | |
9:00am - 9:30am
ID: 126 / A2 Long Papers: 2 Long Paper Keywords: digital scholarly editing; distributed modeling; workflow design; interoperability From Distributed Tagging to Distributed Modeling: Creating Connections Across Environments in the Research Portal BACH Saechsische Akademie der Wissenschaften zu Leipzig, Germany The long-term project "Forschungsportal BACH" (since 2023), a collaboration between the Saxon Academy of Sciences and Humanities in Leipzig and the Bach Archive Leipzig, aims to document, digitally process, and publish all surviving archival documents relating to the musically active members of Johann Sebastian Bach's family from the late 16th to the early 19th century. The corpus is structurally and semantically heterogeneous and includes private and official correspondence, legal and administrative documents, teaching materials, account books, petitions, and composite manuscript units with transcripts and appendices. Letters represent only a small portion of this material. Unlike edition projects that focus on a single genre, the diversity and scope of this corpus require an encoding strategy that is both flexible and sustainable. Rather than conceptualizing TEI encoding as a single-stage editorial process, our project has developed a distributed modeling approach in which encoding decisions are temporally and technically divided across multiple environments. Our workflow includes digitizing documents in archives, storing both data and images in our database, recognizing handwritten text and structural annotations in Transkribus, exporting to PAGE XML, automatically converting to TEI P5 via XSLT, and then adding semantic enrichment in a mostly customized TEI Publisher annotation environment. Structural and layout-related features are primarily captured in the HTR environment, where document segmentation and formal features are identified early in the process. These decisions are then mapped to TEI during automated conversion. In a later phase, editors perform more detailed semantic annotation which is not restricted to conventional named-entity tagging but aims to maximize the informational yield of each document through differentiated semantic markup. The information in the respective documents is automatically supplemented by metadata from the associated database. Furthermore, for correspondence, additional CMIF data are generated to ensure interoperability with correspSearch. This paper discusses that such a workflow shifts TEI encoding from a singular editorial moment to a multi-stage, tool-dependent process. Automated routines help manage task transitions across stages, ensuring a smooth progression from one environment to the next. Three methodological implications of this redistribution are examined: (1) the impact of early structural annotation in the HTR environment on later processing flexibility; (2) the role of automated transformation scripts in mediating between layout-oriented and semantics-oriented representations; (3) the balance between efficiency and scholarly control in large-scale heterogeneous corpora. Drawing on concrete encoding cases, we evaluate how decisions made at different stages influence validation, consistency, and interoperability. Consistency is ensured through controlled transformation routines and structured annotation guidelines across environments. By viewing encoding as a distributed and iterative process, the paper contributes to current discussions on TEI workflows in complex archival contexts. It suggests that large-scale heterogeneous edition projects require a reconceptualization of encoding as a multilayered process shaped as much by technical infrastructures and tool logics as by editorial interpretation. | |
