TEI 2026
Creating Connections, Unsettling Practices
August 10-14, 2026
University of British Columbia, Vancouver, BC, Canada
Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
|
Daily Overview |
| Session | |
|
B1 Long Papers: Reflections on TEI Standards and Models Location: BUCH B210 Session Chair: Janelle Jenstad | |
| Presentation 3 | |
ID: 114
/ B1 Long Papers: 3
Long Paper Keywords: TEI Alternatives, Binary Encoding, Unicode, Minority Writing Systems TEI Beyond TEI Unaffiliated, United States of America We conventionally think of computer text and document formats in terms of hierarchically layered functionality: Unicode encodes individual glyphs, HTML specifies the abstract structure of underlying plain text, and CSS concretely styles an HTML document. Unfortunately, these tidy abstractions collapse upon contact with the actual technologies. In reality, Unicode, HTML, and CSS can each (partially) encode, structure, and style text. What we require is a proper accounting of the remit of a digital text or document format. The problem lies in there being no obvious boundary between contentful text in itself and formatting or structure imposed upon it. Bolding may be a mere visual flourish in a novel, but essential to a mathematical treatise. So, should bolding be handled at the markup or character set level? TEI P5 offers some guidance here, but both positions come down to taste. This stratification of text into character sets and markup is, I argue, unnatural and historically contingent; no single, universally applicable boundary between text and formatting or character and markup can exist. The salience of concrete, visual characteristics or structure to the meaning of text is always indexed to a particular text. Instead, we might imagine the space of document formats as a continuum between the fully concrete, such as a typeset PDF, and fully abstract, such as an application-specific XML schema. TEI is unique in its ability to run this gamut to nearly both extremities within a single mechanism, and in this respect sits at an undertheorized position in this space. Indeed, TEI even includes a facility to define new characters altogether and thus is, in a remote sense, a character set encoding unto itself! While TEI serves as an embryonic model of what a truly universal text and document format could be, it is unsuited to this role in its present form due to, among other things, the rather extreme syntactic overhead of XML itself. In this talk, I introduce Telecode: a new, single standard for encoding, structuring, and styling text, harmonizing the capabilities of Unicode and TEI. Telecode includes both a new character set and an XML-like data model which subsumes XML-based document formats such as TEI—all TEI documents can be expressed as Telecode text—but expands its scope to cover encoding all text at all levels of formatting. I further describe a novel binary encoding for Telecode text, which is competitive with UTF-8 Unicode in compactness. I finally give an overview of Telecode tooling, including an open source editor and software library that I have authored. Telecode's design is particularly attuned to the needs of minority and minoritized writing system communities. Text as a category is not and cannot be a static thing. Text encodings must meet the needs of an ever increasing number of computerized scripts, but the expansion of many such technologies in functionality is metered by standardization bodies. Telecode unifies character encoding itself with a more advanced form of TEI's character declaration features and is thus readily adaptable to new scripts in a decentralized manner. | |
