TEI 2026
Creating Connections, Unsettling Practices
August 10-14, 2026
University of British Columbia, Vancouver, BC, Canada
Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
|
Daily Overview |
| Session | ||
A2 Long Papers: TEI Infrastructures
| ||
| Presentations | ||
8:30am - 9:00am
ID: 112 / A2 Long Papers: 1 Long Paper Keywords: TEI and Beyond, Archives, Graphs The Michael Sperberg-McQueen Digital Nachlass as a basis for the research and study of text modelling 1: University of Cologne, Department for Digital Humanities; 2: University of Cologne, Data Center for the Humanities; 3: University of Cologne, Center for Data and Simulation Sciences The Michael Sperberg-McQueen Digital Nachlass ensures the preservation of Sperberg-McQueen’s digital legacy by creating an accessible archive, built to modern standards of digital preservation. All materials will be archived according to international best practices, providing multiple access layers, and made public for research projects as well as for didactic activities to the extent we legally can. The CMSMcQ archive was created in 2025 and is currently operating based on a limited amount of institutional support from the University of Cologne and private donations. We are already working on a funding bid to further develop the project and continue expanding—and as far as possible fulfilling—the potential of the digital archive. The initial phase of our project focuses on the publication of a bibliography of Sperberg-McQueen’s work, with links to online versions of his published work whenever possible: research articles, webpages, blogs, standards, source code, and running systems. Taking the cue from the modelling discussions found in the CMSMcQ archive, as part of the long tradition of text modelling, we ask what the best formalism for representing complex texts is. TEI has always used graph formalisms to represent text. Both SGML and TEI represent hierarchical, rooted, ordered, simple, undirected, connected, acyclic graphs, that is, a type of trees. Already in the 1990s, other graphs formalisms were suggested, but had a limited application due to a lack of formal validation systems and also a lack of useful tools. While some sort of cyclic structure can be expressed in XML using XML ID/IDREF links, and this is highly useful for TEI, this is rather an addition to the XML structure than an extension of it, especially when it comes to validation. Using a less strict graph formalism for text encoding, for instance by including cycles, one can still use significant parts of TEI. We will show examples of encoded documents where the element names and their meaning (as expressed in the text of the TEI Guidelines) are mostly kept. To what extent the meaning changes is one of our research questions we look forward to discussing at the conference. We will also show how the development in user friendly graph encoding and visualisation systems has made the use of less restricted graph formalisms more available, also outside experimental setting. Further to that, we will suggest some consequences for the integration with external ontologies and other knowledge representation systems. How to develop meaningful translation paths from graph based encoding to XML based TEI? We will show some examples and discuss to what extent the semantics of the encoding can be kept, and to what extent the result will be human readable. We will also discuss how the very conception of a less restricted graph is a scholarly endeavour that uncovers new dimensions of the materials. The flexibility granted by the chosen framework, paired with the development of a rule system for our graphs which enable syntactic graph validation, allow for a higher degree of intellectual creativity and bottom-up modelling compared to other methodologies. 9:00am - 9:30am
ID: 126 / A2 Long Papers: 2 Long Paper Keywords: digital scholarly editing; distributed modeling; workflow design; interoperability From Distributed Tagging to Distributed Modeling: Creating Connections Across Environments in the Research Portal BACH Saechsische Akademie der Wissenschaften zu Leipzig, Germany The long-term project "Forschungsportal BACH" (since 2023), a collaboration between the Saxon Academy of Sciences and Humanities in Leipzig and the Bach Archive Leipzig, aims to document, digitally process, and publish all surviving archival documents relating to the musically active members of Johann Sebastian Bach's family from the late 16th to the early 19th century. The corpus is structurally and semantically heterogeneous and includes private and official correspondence, legal and administrative documents, teaching materials, account books, petitions, and composite manuscript units with transcripts and appendices. Letters represent only a small portion of this material. Unlike edition projects that focus on a single genre, the diversity and scope of this corpus require an encoding strategy that is both flexible and sustainable. Rather than conceptualizing TEI encoding as a single-stage editorial process, our project has developed a distributed modeling approach in which encoding decisions are temporally and technically divided across multiple environments. Our workflow includes digitizing documents in archives, storing both data and images in our database, recognizing handwritten text and structural annotations in Transkribus, exporting to PAGE XML, automatically converting to TEI P5 via XSLT, and then adding semantic enrichment in a mostly customized TEI Publisher annotation environment. Structural and layout-related features are primarily captured in the HTR environment, where document segmentation and formal features are identified early in the process. These decisions are then mapped to TEI during automated conversion. In a later phase, editors perform more detailed semantic annotation which is not restricted to conventional named-entity tagging but aims to maximize the informational yield of each document through differentiated semantic markup. The information in the respective documents is automatically supplemented by metadata from the associated database. Furthermore, for correspondence, additional CMIF data are generated to ensure interoperability with correspSearch. This paper discusses that such a workflow shifts TEI encoding from a singular editorial moment to a multi-stage, tool-dependent process. Automated routines help manage task transitions across stages, ensuring a smooth progression from one environment to the next. Three methodological implications of this redistribution are examined: (1) the impact of early structural annotation in the HTR environment on later processing flexibility; (2) the role of automated transformation scripts in mediating between layout-oriented and semantics-oriented representations; (3) the balance between efficiency and scholarly control in large-scale heterogeneous corpora. Drawing on concrete encoding cases, we evaluate how decisions made at different stages influence validation, consistency, and interoperability. Consistency is ensured through controlled transformation routines and structured annotation guidelines across environments. By viewing encoding as a distributed and iterative process, the paper contributes to current discussions on TEI workflows in complex archival contexts. It suggests that large-scale heterogeneous edition projects require a reconceptualization of encoding as a multilayered process shaped as much by technical infrastructures and tool logics as by editorial interpretation. 9:30am - 10:00am
ID: 140 / A2 Long Papers: 3 Long Paper Keywords: events, history, semantic analysis, ontologies, London But What Is an Event? Modeling Historical Data with TEI and CIDOC-CRM Bucknell University, United States of America The adoption of the eventName element by the TEI in Fall 2023, promised a new way for researchers encoding historical documents to annotate occurrences in ways similar to those we use to recognize people, places, and organizations. Understandably, there was some question at the time about what constitutes an event, or what rises to the level of an event. Ultimately, among the examples chosen when the element was published included what might be described as significant events (those worthy of 'proper noun' status): World War II, the 1618 Defenestration of Prague, and the TEI 2019 Conference (!) I would suggest that these are events by consensus – occurrences that are notable because many people agreed to confer named fame to them. And yet, encoding archival materials that are contemporaneous with such famous phenomena (in my case performances connected to political occasions in early modern London) there is a lack of explicit references to 'proper noun' events like the Coronation of Queen Elizabeth I, the English victory over the Spanish Armada, or the Marriage of Princess Elizabeth Stuart and Frederick V. Rather, in correspondence, personal diaries, municipal and financial records studied as part of the REED London project, I have realized that people are more likely to focus on what are really nested or micro-events undertaken in relation to those phenomena, but which are contextually important because of the micro-event itself: production of a masque, procession through the streets of London, a mock sea-battle on the River Thames. All of these performance events were planned, funded, performed by, and witnessed by unnamed Londoners as well as notable people. The records and eyewitness accounts that describe the performances sometimes refer obliquely to a super-event (coronation, victory, wedding festivities) but not always, and the focus of the people who wrote the letters, organized the dances, and tracked the expenses clearly was on the business of the production. But are these events that rise to the level that the TEI Council envisioned? At what point do the details about planning, production, and observation of a performance collapse under the weight of a tagged document? In the past year I have collaborated with a group of researchers extending the Linked Art ontology to include to better represent performing arts data (which is itself an extension of the CIDOC Conceptual Reference Model). Using the Masque of the Inner Temple and Gray's Inn as the first test of a very complex example of this type of nested event, I have been trying to create lemma-level connections between the richly-encoded TEI files that reveal components of the preparation and production of the Masque evidenced in the REED London records and the structured LAPA data model. In this paper I will present the progress I am making in making meaningful connections between TEI event structures and the LAPA modeling techniques. In doing so I will consider how far we can push the eventName element to sustain such connections, and when those connections necessarily break, or must be adjusted. 10:00am - 10:30am
ID: 131 / A2 Long Papers: 4 Long Paper Keywords: critical semantic modeling, linked data, knowledge representation, feminist data epistemology, cultural heritage metadata Modeling memory as a constellation: traveling with ontologies, capturing archives, representing tacit knowledge University of Ottawa, Canada This paper focuses on the effect of categorization on historical tacit knowledge, in line with the TEI 2026 themes of “Data Interoperability" (or “Cross-Domain Application”). Within the constellation of my research, I propose to shift the center of gravity away from “Let’s categorize!” to “What happens when we shape memory this way?”. This study examines how applying core ontologies to the contributions of Canadian women of impact to history and archaeology, in an equity-aware approach, reshapes how cultural knowledge is modeled and shared across disciplinary boundaries. While linked data standards are increasingly used in digital heritage infrastructures, the humanities data remain underrepresented and difficult to model in interoperable ways[1]. Despite large-scale initiatives and DH infrastructures, there is limited evidence on how existing models accommodate humanities-specific data types, especially those carrying tacit knowledge with them, such as personal archives and local place-based knowledge, and how these are encoded with minimal structural biases[2] [3]. This gap is addressed by using data from a set of archaeological archive collections held in Canadian institutions, as a testbed for assessing the possibilities and limits of modeling interoperable cultural data on material documenting heritage management activities. These documents include notebooks, correspondence, diaries, and extensive photographic materials capturing landscapes, artefacts, and field teams. Together, they record sites and objects but also research practices, collaborations, and relationships with the land and the communities over several decades. How can models represent people, places, events, and objects in these documents in ways that respect their historical and cultural specificity while enabling cross-structures interoperability? How do we investigate the limits encountered when encoding and assigning persistent identifiers to entities drawn from multilingual, historically layered archival descriptions; which CRM extensions can accommodate concepts specific to specific times, languages, and cultures within a broadly shared ontology? This project combines methods from qualitative data modeling with practical TEI encoding and RDF conversion workflows. It unfolds within a theoretical space structured by (1) critical semantic modeling, turning archival description into structured markup (what counts as “entity”, how to register uncertainty, contestation?), and (2) interpretive data modeling (how entities or their relationships will be modeled, understood, read). The analysis focuses on two intertwined dimensions: (1) semantic adequacy, where CIDOC CRM for example, provides a strong fit for events or artefact production and cases where it struggles (roles exceeding standard actors categories or nuanced distinctions); and (2) interpretive/ethical implications, as we adopt an intersectional perspective to examine how modeling choices inherited from archival descriptions, and authority structures shape who and what becomes visible [4] [5]. Preliminary findings suggest that ontologies (CIDOC CRM) significantly improve interoperability across documents, but the modeling process reveals persistent gaps: historically significant distinctions important to archaeologists and communities are difficult to represent with TEI, express in CIDOC CRM and map in RDF. Additionally, authority-based identifiers often reproduce existing asymmetries in whose names and places are standardized. These observations resonate with concerns in FAIR humanities data about the risk of losing thick description and about the dependency of researchers on external infrastructures and policies. | ||
