TEI 2026
Creating Connections, Unsettling Practices
August 10-14, 2026
University of British Columbia, Vancouver, BC, Canada
Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
|
Daily Overview |
| Session | |
|
B2 Long Papers: Writing Systems and the TEI Location: BUCH D316 Session Chair: Dimitra Grigoriou, Austrian Academy of Sciences | |
| Presentation 4 | |
12:15pm - 12:45pm
ID: 128 / B2 Long Papers: 4 Long Paper Keywords: Multilingual text encoding, tonal languages Encoding Tone in TEI: Expanding Standards for Cross-Linguistic Applicability University of Graz, Austria The TEI has served as the foundation for digital scholarly editing across diverse cultural and linguistic contexts for over three decades. Despite this broad adoption, the Guidelines contain a significant gap: there is no standardized mechanism for encoding tone, a phonological feature present in approximately 70% of the world’s languages (Yip 2002; Maddieson 2013). This absence reflects structural biases within digital humanities infrastructure, where existing standards do not fully accommodate the application to non-European language families (Bowers 2020; Lukaniec & Holmes 2023; Ngué Um 2017; Thieberger & Tuohy 2017). Providing a dedicated mechanism for tone would further TEI’s commitment to internationalization and multilingualism, strengthening its legitimacy as a global standard. Tone carries lexical and grammatical meaning across a vast range of languages, distinguishing words that would otherwise be homophonous and adding grammatical functions. Failing to represent tone is not a matter of omitting a diacritic—it is omitting meaning itself, leaving the representation linguistically incomplete. While existing TEI mechanisms, e.g., feature structures, could serve this purpose, no general recommendation exists. This results in project-specific workarounds that undermine interoperability, complicate computational analysis, and limit scholarly reuse. Drawing on our experience creating a digital scholarly edition of a 17th-century Chinese-Spanish dictionary from colonial Manila (Döhla et al. 2022; 2024), we present both the practical consequences of this encoding gap and a concrete proposal for addressing it. The “Bocabulario de lengua sangleya por las letraz de el A.B.C.” (British Library Add ms. 25.317) documents Hokkien as spoken by Chinese immigrants in early Spanish Philippines (Döhla 2025). Because tone distinguishes lexical meaning, accurate linguistic interpretation depends on tone annotation. The current implementation employs @ana attributes linked to a SKOS-based controlled vocabulary, following the reconstructed seven tone categories for Early Manila Hokkien (Klöter 2003). This solution is functional, but tone—as a prominent feature in the world’s languages— warrants a dedicated recommendation within the Guidelines, ensuring visibility. We therefore propose a <toneDecl> element within <encodingDesc> for declaring tone inventories, following established TEI patterns for analytical frameworks. Individual <toneDef> entries would specify @xml:id for unique identification and @type distinguishing level from contour tones based on pitch patterns. Each definition includes human-readable <label> and <desc> elements, with optional <pitch> specifications using established notational systems such as Chao tone letters. A corresponding @toneRef attribute links tone-bearing segments to their tonal definitions. The minimal attribute should accommodate typological diversity without imposing analytical assumptions, while the declaration-reference pattern maintains TEI’s separation of metadata and content. Community deliberation is required to assess whether the proposal is truly applicable cross-linguistically and suitable for different levels of linguistic description, since tone by definition applies to different aspects like phonology or morphology. This work directly addresses TEI 2026’s call for “Creating Connections, Unsettling Practices.” Challenging Eurocentric biases in infrastructure through concrete technical proposals contributes to decolonizing digital scholarship. Enabling the encoding of fundamental linguistic features allows previously underrepresented textual traditions to receive appropriate digital representation and opens new possibilities for scholarly connection. | |
