TEI 2026
Creating Connections, Unsettling Practices
August 10-14, 2026
University of British Columbia, Vancouver, BC, Canada
Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
|
Daily Overview |
| Session | ||
B2 Long Papers: Writing Systems and the TEI
| ||
| Presentations | ||
10:45am - 11:15am
ID: 106 / B2 Long Papers: 1 Long Paper Keywords: Chinese Paleography, Text-encoding, Unicode Standard, <g> element, automized text collation The Implementation of the TEI-Guidelines on Early Chinese Excavated Administrative Documents: With a Focus on Paleographic Issues Meiji University, Japan This paper aims to implement the TEI-Guidelines on early Chinese excavated administrative documents. As the handling of these documents stretches over the academic fields of both Chinese paleography and paleo-diplomatics, the respective traditional descriptive tools of these two fields need to be reflected. Being still at a preliminary stage, this paper focuses on the paleographic features only, which early Chinese excavated administrative documents share with excavated literature. In praxis, Chinese paleography is the art of transcribing and collating early Chinese manuscripts. In order to store information about the transcribers’ understanding of the original text within the modern transcriptions, Chinese paleography has developed a peculiar markup language. Interestingly, this traditional markup language can be easily translated into existing TEI elements and attributes. In this regard, this paper proves the high sophistication of the TEI-Guidelines rather than adding something new to them. By contrast, a more theoretical branch of Chinese paleography, which is occupied with linguistic issues like the changes in character usage over time and in space, poses more complex semantic challenges. Basically, the paleographic understanding of Chinese characters differs hugely from that of the Unicode standard, which ultimately represents a compromise to meet multifold diverging practical demands. This gap leads to the necessity of redefining characters even in cases where Unicode seemingly provides clearly distinguishable “ideographs”. For this redefinition, this paper attempts to change the semantic meaning of <g>(gaiji) elements, linking them to paleographic character definitions within the <charDecl> element in the TEI header or in a stand-off data format. Classically, a Chinese character is understood as the combination of form, reading, and meaning. Studies indicate stable links between reading and meaning, whereas form frequently undergoes quicker, seemingly arbitrary alterations. Certain Chinese linguists suggest concentrating on reading and meaning units as a more suitable tool for linguistic analysis of changes in character usage. Shào Yǒnghǎi邵永海 calls these units “character positions (zì wèi字位)”, the concept of “position” referring to a relatively fixed position in the semantic space in contrast to the easily changing forms of characters. This paper supports this more historical approach to Chinese characters and attempts to realize it by separating the description of “character positions” from “character forms” in the <char> and <glyph> elements within the <charDecl> element respectively. In practical application, this redefinition of the <g> element results in all characters in the main body and the attached apparatuses being turned into <g> elements, using @ref and @ana attributes to associate them with the corresponding form and character position definitions in the <glyph> and <char> elements. This has the beneficial side effect of eliminating arbitrary text and element node mixing, converting text into homogeneous sequences of <g> elements. Combined with stand-off solutions for annotations and strict stratification of different layers of annotation, this furthermore facilitates automated collation of encoded text and complex arrangements of overlapping annotations as will be shown by means of a simple example from excavated literature. 11:15am - 11:45am
ID: 111 / B2 Long Papers: 2 Long Paper Keywords: digital editions, historical texts, methodology, digital library, infrastructure Manuscriptorium Full-Text Module: Building a Digital-Edition Infrastructure National Library of the Czech Republic, Czech Republic (Czechia) After several years of development, the transition to the new version of the Manuscriptorium digital library was completed last autumn. The development, however, is still ongoing—the most significant and recent part of which is the TEI-compatible full-text module. As a result, Manuscriptorium offers its users an alternative core that shifts the central perspective from visual representation of digitized documents to textual content, with images serving as a complement to the texts they contain. The overall objective is to build a solid and stable infrastructure, enabling the community of scholars as well as the general public to read, study, and research the historical texts within a wider Central European digital context. At present, the main focus lies on converting older editions of historical texts of Bohemian origin. The editions are then correlated not only with the documents forming their textual basis but also with other known witnesses, thereby reviving such editions in a broader and more dynamic context. In the extended process of establishing the exact method of data preparation, their subsequent processing, and visualization workflows, we have come across several methodological ‘nooks and crannies’ for which the current TEI Guidelines offer no proper solution. Thus, a rather creative approach was at times necessary, and even a few changes and extensions of the TEI standard in use proved unavoidable, primarily in relation to the problem of directly connecting digitized fragmentary witnesses of varying extent and content with the full texts of the editions. From the end-user perspective, the edition module consists of two components: an application for advanced viewing of the texts integrated with the digitized images of original documents, and a catalog enabling high-granularity searching across the edition database. The essential data persistence has already been secured by the overarching Manuscriptorium resolver infrasctructure. By means of full IIIF compatibility, the texts are correlated with the digitized documents from Manuscriptorium's content core as well as from external image repositories. Parallel yet differentiated browsing of texts and images, along with the additional visualization of modern translations or transcriptional variants, has also been successfully implemented. The next developmental phase will be primarily dedicated to processing and visualizing complex critical apparatuses. On the whole, compared with other historical-text databases, our principal purpose is to provide users with highly curated—and thus reliable—textual content and to make it as accessible as possible, not least through an engaging and easy-to-navigate user interface. In the paper, all the above-mentioned aspects of the development and the current state of the module will be discussed, with the presentation of methodological problems relating to the markup, its specifics and broader applicability, among the key objectives. Quite naturally, adhering as closely as possible to the TEI Guidelines, we aim to invite critical feedback to our work from the community best suited for it. 11:45am - 12:15pm
ID: 123 / B2 Long Papers: 3 Long Paper Keywords: EpiDoc, Rust, Digital Epigraphy, Middle Persian, Interoperability BEDA: A High-Performance Rust Ecosystem for Binary-Extended EpiDoc and Complex Script Digitization Sapienza University of Rome, Italy Traditional digital epigraphy has long relied on TEI-EpiDoc XML as the primary standard for data modeling and interchange. In the current technological landscape, digital epigraphy often navigates the apparent oxymoron between the descriptive complexity required for archiving and the high-performance efficiency demanded by responsive web interfaces. This paper introduces BEDA (Binary Epigraphic Data Archive), a Rust ecosystem that addresses this tension through a two-layer architecture: a binary archive format for metadata and digital surrogate management (including zoomable facsimile and RTI), and a dedicated epigraphic DSL for authoring critical editions. 12:15pm - 12:45pm
ID: 128 / B2 Long Papers: 4 Long Paper Keywords: Multilingual text encoding, tonal languages Encoding Tone in TEI: Expanding Standards for Cross-Linguistic Applicability University of Graz, Austria The TEI has served as the foundation for digital scholarly editing across diverse cultural and linguistic contexts for over three decades. Despite this broad adoption, the Guidelines contain a significant gap: there is no standardized mechanism for encoding tone, a phonological feature present in approximately 70% of the world’s languages (Yip 2002; Maddieson 2013). This absence reflects structural biases within digital humanities infrastructure, where existing standards do not fully accommodate the application to non-European language families (Bowers 2020; Lukaniec & Holmes 2023; Ngué Um 2017; Thieberger & Tuohy 2017). Providing a dedicated mechanism for tone would further TEI’s commitment to internationalization and multilingualism, strengthening its legitimacy as a global standard. Tone carries lexical and grammatical meaning across a vast range of languages, distinguishing words that would otherwise be homophonous and adding grammatical functions. Failing to represent tone is not a matter of omitting a diacritic—it is omitting meaning itself, leaving the representation linguistically incomplete. While existing TEI mechanisms, e.g., feature structures, could serve this purpose, no general recommendation exists. This results in project-specific workarounds that undermine interoperability, complicate computational analysis, and limit scholarly reuse. Drawing on our experience creating a digital scholarly edition of a 17th-century Chinese-Spanish dictionary from colonial Manila (Döhla et al. 2022; 2024), we present both the practical consequences of this encoding gap and a concrete proposal for addressing it. The “Bocabulario de lengua sangleya por las letraz de el A.B.C.” (British Library Add ms. 25.317) documents Hokkien as spoken by Chinese immigrants in early Spanish Philippines (Döhla 2025). Because tone distinguishes lexical meaning, accurate linguistic interpretation depends on tone annotation. The current implementation employs @ana attributes linked to a SKOS-based controlled vocabulary, following the reconstructed seven tone categories for Early Manila Hokkien (Klöter 2003). This solution is functional, but tone—as a prominent feature in the world’s languages— warrants a dedicated recommendation within the Guidelines, ensuring visibility. We therefore propose a <toneDecl> element within <encodingDesc> for declaring tone inventories, following established TEI patterns for analytical frameworks. Individual <toneDef> entries would specify @xml:id for unique identification and @type distinguishing level from contour tones based on pitch patterns. Each definition includes human-readable <label> and <desc> elements, with optional <pitch> specifications using established notational systems such as Chao tone letters. A corresponding @toneRef attribute links tone-bearing segments to their tonal definitions. The minimal attribute should accommodate typological diversity without imposing analytical assumptions, while the declaration-reference pattern maintains TEI’s separation of metadata and content. Community deliberation is required to assess whether the proposal is truly applicable cross-linguistically and suitable for different levels of linguistic description, since tone by definition applies to different aspects like phonology or morphology. This work directly addresses TEI 2026’s call for “Creating Connections, Unsettling Practices.” Challenging Eurocentric biases in infrastructure through concrete technical proposals contributes to decolonizing digital scholarship. Enabling the encoding of fundamental linguistic features allows previously underrepresented textual traditions to receive appropriate digital representation and opens new possibilities for scholarly connection. | ||
