Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
|
Daily Overview |
| Session | ||
Poster Session Lightning Talks
| ||
| Presentations | ||
ID: 133
/ Poster Session Lightning Talks: 1
Poster Keywords: P6 “Ceci n'est pas une pipe” or, “William” is not a person or, the three TEIs Technical University Darmstadt, Germany About his famous painting, “La trahison des images“, René Magritte said: ‘The famous pipe. How people reproached me for it! And yet, could you stuff my pipe? No, it's just a representation, is it not? So if I had written on my picture "This is a pipe", I'd have been lying!’ Other interpretations notwithstanding, both painting and quote show the importance to make a clear distinction between a representation and an actual thing; or a “signifié” and “signifiant“, to use Saussure’s terms. Regarding the TEI and a possible P6, this to my mind means that we need a clearer distinction between “the three TEIs”. We have a set of elements and attributes for schema description/definition, for text description, and for meta data (in the broadest sense). The problem is that these three sets that have very distinct uses share a lot of elements and attributes. We have cases that state what an element or attribute is about when annotating a text, just to add that it has a completely different meaning when used in a schema defintion. This makes reusing the data much more complicated as one cannot be certain that, e.g., a word contained in a tei:p is actually part of a text. My poster will present this, admittedly radical, approach how to fully separate the three areas, thus reducing ambiguity, improving interoperability, and making information retrieval easier. It is, however, aimed at retaining as much compatibility with P5 as possible by trying to reuse existing definitions wherever that is possible. I hope to spark a discussion about whether this is a good approach and how a new P6 would be drawn up. ID: 124
/ Poster Session Lightning Talks: 2
Poster Keywords: textual variance, shared vocabulary, knowledge organisation, variants, digital editing A Taxonomy for Textual Variation: Updates from the VIDIT Working Group 1: Huygens Institute; 2: University Of Nebraska–Lincoln; 3: Università di Parma; 4: Ca' Foscari University of Venice; 5: Universität Wien; 6: Freie Universität Berlin; 7: Università di Torino When working with variant texts, scholars often struggle to express textual complexities in a formal, machine-processable manner: it is the familiar balance between standardization and flexibility (Eide 2015). In textual scholarship, this is further complicated by linguistic, cultural, and disciplinary differences (Spadini et al. 2016; Bleeker et al. 2025a). As of yet, there is no shared vocabulary to describe textual variation, tools and datasets lack interoperability (Martignano 2023a), and visualizations are often project-specific and not reusable (Franzini et al. 2019). The international working group “VIDIT” (Visualising and Investigating Differences in Texts) aims to address these challenges by (1) developing a taxonomy of core concepts related to textual variation, and (2) establishing best practices for collation and visualization tools (Bleeker and Nava, 2025b). This poster provides an update of VIDIT’s activities, focusing on the methodological approach and the first taxonomy results. The taxonomy development has three stages. First, core concepts of text variation are collaboratively identified. For each concept, examples are gathered from texts from different periods and genres, along with definitions from scholarly publications, tool documentation, and encoding guidelines. Rather than imposing a single definition, we identify the range of meanings and contexts of use. Second, these concepts are structured following knowledge organization principles: identifying hierarchies, clarifying overlaps and distinctions, and documenting conditions under which different definitions apply. Third, the taxonomy is aligned with existing ontologies in the field (Giovannetti 2019, Martignano 2023b, Spadini 2019[2023]). Our methodology has raised several questions, among others relating to the balance between comprehensiveness and usability. When and how to prioritise consensus over diversity? How to ensure the taxonomy’s relevance as digital methods evolve, and how to couple the taxonomy with the TEI Guidelines? We welcome feedback on the above questions as well as on our methodological approach and the first taxonomy results. ID: 113
/ Poster Session Lightning Talks: 3
Poster Keywords: graffiti, image, Mexico, gender-based violence, feminist From Image to Text: Encoding the (Dis)appearance of Missing and Murdered Women in the Streets of Mexico University of British Columbia, Canada This poster reports the preliminary findings from my field work to Mexico City, one of the leading places for the creation of artivism (art + activism) that seeks to address social inequalities and oppression. During the trip, I photographed street art (graffiti, banners, mural paintings) produced in the context of International Women’s Day on March 8th, 2026. The present research is part of my doctoral dissertation that tracks how contemporary literary works by Mexican women are increasingly engaging with horror and the gothic mode to grapple with gender-based violence. To complement my analysis of literary representations, I propose that we turn to other art forms and spaces through which women denounce heteropatriarchal, racial, and colonial violence. Given graffiti’s discursive power but also its ephemeral nature, I approach its textual study as a practice of recovery that seeks to amplify marginalized voices and make visible women’s concerns. To standardize my research, I visited three specific locations before and after the social mobilizations of March 8. Tracing the same locations at different times provides insights into censorship and the erasure of women’s demands for justice. After the trip, I transcribed, encoded, and annotated a sample of visual materials following TEI guidelines. Some preliminary parameters I have identified include location; temporality; the surface (for example, fence, cement sidewalk, monument); and themes (femicide, reproductive rights, domestic abuse, among others). The purpose of cataloguing these images as text is twofold: to explore how text encoding enriches our understanding of the often-underrepresented street narrative and how these textual examples expand our use of TEI through a feminist lens. Ultimately, my research aims to contribute to current activism and debates that are in the process of creating a language that allows us to recognize, name, and address gender violence and its aftermath. ID: 149
/ Poster Session Lightning Talks: 4
Poster Keywords: medical corpus analysis, CMC corpora, data modeling, AI agents, chat logs Longevity Chats in XML-TEI for Medical Corpus Analysis University of Rostock, Germany An interdisciplinary project on the topic of ethics and AI is currently being prepared at the University of Rostock. Researchers from medicine, computer science, ethics, interdisciplinary research, and DH aim to jointly explore the role of AI agents in settings of co-medical reasoning (Salloch and Eriksen, 2024, Porsdam Mann et al., 2024) and situatedness (Troqe, Lakemond and Holmberg, 2024). To this end, AI agents are developed and assessed in the area of longevity and dementia studies, focusing on ethical implications. The DH subproject aims to model and evaluate chat logs for medical consultations as text corpora. Of particular interest is linking the texts with situation-related parameters (e.g., when and where someone used the chat), information about the patients (gender, age, health status), and other protocols of human conversation (between doctor and patient). The aim is to map the collaborative process of consultation, anamnesis, and diagnosis (co-medical reasoning) and the co-situating of AI agents with human participants. We are currently evaluating whether and to what extent modeling the data in XML-TEI would be useful for the project (using elements for encoding CMC corpora, transcriptions of speech, and the profile description in the TEI header). This is a novel approach in the field of medical and AI research on this topic, but it is promising in view of the desired modeling of co-situatedness. First, fictitious chat logs are created with Chat-GPT on the topic of longevity and counseling, which are used for modeling purposes, as the protection of personal data must be taken into account when using real chats (see examples at https://github.com/hennyu/longevity-chats). The poster presents the current state of preparatory work for the project and invites discussion with the TEI community. ID: 137
/ Poster Session Lightning Talks: 5
Poster Keywords: methods, evaluation, ontologies, correspondence editions Mapping correspondence editions: a preliminary look at the themes, methods, and work practices of a global DH community Rutgers University, United States of America Digital editions of correspondence are vital projects for the analysis of historical and cultural figures and events in addition to being popular training grounds for students. While the TEI Correspondence SIG has codified editorial standards and the web service correspSearch and the review journal RIDE have improved discoverability and methodological transparency, knowledge of the range of possibilities with source encoding and presentation remains difficult to acquire. What is the current state of the digital scholarly edition of correspondence in the global DH community? There are significant differences in text encoding projects emerging from various geopolitical contexts. Greedy of time, funding, expertise, and labor, such projects have historically benefited from large teams, particularly in the historic "center" of digital humanities activity in North America and Europe. Given increased resource constraints and a diminished funding outlook, scholars in what we call the "South" but also increasingly in Northern contexts, have led a turn towards minimal computing frameworks. Are there lessons to be learned from examining projects emerging from the centers and margins of DH knowledge production that will improve our collective understanding of the planning and execution of these projects? In this poster presentation, I report preliminary findings from a thematic and methodological analysis of a diverse corpus of correspondence editions. I will track variables including: the historical period and language(s) of the sources, disciplinary and theoretical frameworks, the DH methods used, the encoded entities, technical infrastructure (static or dynamic site software), as well as project longevity. By aggregating and analyzing these features, this study will establish a framework for identifying current methodological patterns and build awareness for the diversity of approaches possible within this community of practice. ID: 110
/ Poster Session Lightning Talks: 6
Poster Keywords: TEI CMC, social media, Japanese language corpora, microblogging, messaging Modeling Japanese Social Media with TEI CMC: Integrating Microblogging and Messaging in BCCWJ2 National Institute for Japanese Language and Linguistics, Japan This poster presents the integration of social media data into the Balanced Corpus of Contemporary Written Japanese 2 (BCCWJ2) and examines how heterogeneous computer-mediated communication (CMC) environments can be modeled within the TEI CMC framework. The original Balanced Corpus of Contemporary Written Japanese (BCCWJ), released in 2011 by the National Institute for Japanese Language and Linguistics, is a publicly available, genre-balanced corpus of modern written Japanese, including books, magazines, white papers, and web texts. It has served as a foundational resource for Japanese linguistics. However, most of its materials date from 2005 or earlier. To better reflect contemporary language use, BCCWJ2 (2024–2028) aims to expand the corpus to approximately 200 million words by incorporating texts published between 2006 and 2025, including digital communication that now plays a central role in everyday written practices. A key challenge is the inclusion of social media data. Our project integrates both microblogging platforms (Bluesky and Misskey) and a messaging service (LINE). While the TEI CMC Guidelines provide models for Twitter-like services and for messenger-style interaction, actual platforms present service-specific complexities that do not always align with text-centered encoding assumptions. Particular issues arise in encoding platform-specific reactions and visual elements. Misskey supports custom, non-Unicode emoji used as reactions, and LINE employs stickers and rich reaction systems that cannot be reduced to plain text. These semiotic elements function as discourse acts yet resist straightforward textual representation. By comparing microblogging and messaging ecologies within a single corpus design, this project evaluates how far TEI CMC can accommodate diverse social media practices and where customization may be required. We argue that constructing a balanced corpus including such data contributes to documenting contemporary Japanese linguistic life and promotes sustainable, interoperable, and globally accessible Japanese language resources. ID: 120
/ Poster Session Lightning Talks: 7
Poster Keywords: TEI Processing Model, ODD, web components, machine-assisted annotation, publication Modular, reusable, sustainable and community-oriented edition workflows with TEI Publisher 1: e-editiones; 2: Jinntec TEI Publisher[1] provides a flexible and sustainable toolbox which enables scholars to publish their material without forcing a one-size-fits-all framework. It lays out a smooth and fast entry path for editors with little technical experience while remaining endlessly flexible and extensible for advanced users. With the newest version, based on Jinks application manager,[2] the key TEI Publisher idea of assembling an edition from modular "lego" blocks is taken to the next level. The core of TEI Publisher itself is now decomposed into a set of small, modular profiles, which can be combined and configured to assemble concrete applications. Each profile provides end-to-end implementation for a particular aspect of a digital edition, e.g. support for a certain input format, integration of facsimile images, display of timelines or maps. TEI Publisher now comes with a central application manager, Jinks, which manages all profiles and all the applications generated from it. Thus, the creation of an application is no longer a one-time generative step, but it can be approached iteratively: new features may be added or removed, and the configuration for the existing ones can be modified. The same user interface which first allowed us to select, pre-configure and generate our application, can be used to reconfigure and regenerate it at any later point. As new profiles are added to the public library—or new versions are released—this application manager is able to update profiles and custom applications under its control. Updates can largely be carried out automatically with the number of necessary manual interventions reduced to a minimum and most updates can be applied with a single click. Long-term maintenance of our editions is thus guaranteed and easily achieved. [1] https://teipublisher.com ID: 109
/ Poster Session Lightning Talks: 8
Poster Keywords: Buddhism, Chinese Buddhism, Indic script, Unicode, transcription Siddhaṃ into TEI: Recording Indic and East Asian Complexities 1: International Institute for Digital Humanities; 2: Keio University; 3: Musashino University As part of a project to convert the Taishō Tripiṭaka—the scholarly standard edition of the Chinese Buddhist canon—into TEI-compliant XML, we are migrating Siddhaṃ script data from an internal transcription rule to Unicode to facilitate better data exchange. Siddhaṃ is a medieval Brahmic script spread to East Asia approximately between the 6th and the 8th centuries as a medium for Buddhist texts; it has been preserved there as a liturgical alphabet for transcribing Indic syllables. Beyond its complex ligature system inherent to Indic abugidas, Siddhaṃ has acquired ideographic or iconic qualities in East Asia. Due to its religious role and the influence of the logographic Chinese writing system, specific variants have become associated with fixed meanings. In the Japanese Esoteric Buddhist (Mikkyō) tradition, this differentiation of variants is particularly systematic in the use of shuji (種字; “seed syllables”) to represent specific deities. During the migration process, we compared our original notations in the Taishō Tripiṭaka against facsimile images and other Siddhaṃ resources. This examination revealed various non-standard glyphs, including irregular akṣaras (syllables), idiosyncratic variants, and possibly unencoded symbols and conjuncts. This presentation provides a preliminary overview of Siddhaṃ usage in the Taishō Tripiṭaka, evaluates current Unicode and font support, and discusses optimal encoding strategies for academic Siddhaṃ texts. ID: 160
/ Poster Session Lightning Talks: 9
Late-breaking Poster Keywords: Computable text, Verse text, Data modeling <l>10 PRINT"Why Code Is Nothing Less Than Sweetest Poetry"</l>: Computable Texts, Verse Structures, and the Limits of TEI Semantics University of Wuerzburg, Germany Born-digital heritage initiatives increasingly preserve software as cultural artefacts, yet the textual status of source code itself remains comparatively underexplored. This poster proposes a deliberately provocative, TEI-centered experiment: what happens if computer code is encoded not merely as functional notation, but as a literary and line-based textual form comparable to verse? The approach focuses on early home-computer listings, especially from languages such as BASIC, whose syntax was explicitly designed to resemble natural language. Especially in these sources, but also in other programming languages, code is not only executable instruction, but also a material writing practice: visually arranged (through indentation and lineation), rhythmically structured (through loops and iterations), stylistically distinctive (particularly when deviating from normative coding conventions), and often strongly associated with individual authorship. Like poetry, code is usually processed line by line; lineation structures interpretation, pacing, hierarchy, and semantic grouping. Against this background, the poster explores several experimental directions within the TEI. The central question is whether source code, when understood as authored and lineated text, can be represented through structures traditionally associated with verse. Possible approaches include the repurposing, conceptual broadening, or even renaming of elements such as At the same time, the poster explicitly addresses objections to this approach: the danger of collapsing distinctions between functional and aesthetic texts, violating established TEI semantics, or obscuring computational structures already represented through parsers and formal grammars. Rather than advocating an immediate standardization effort, the poster presents these encodings as experimental interventions intended to test the conceptual boundaries of TEI itself. Ultimately, the poster asks a broader methodological question relevant to textual scholarship and digital philology alike: is verse defined by aesthetics, by material layout, or by structural segmentation? And what happens to TEI semantics when executable code is treated as authored, lineated text? #sky { color: blue; border: none; } ID: 159
/ Poster Session Lightning Talks: 10
Late-breaking Poster Keywords: language code, language tag, language variation, standardization, metadata A Massive Language Code Expansion Is Coming! 1: International Institute for Digital Humanities; 2: National Institute for Japanese Language and Linguistics, Japan This presentation reports on a very recent development, following the May 2026 meeting, regarding ISO's initiative for the coding of language varieties. The outcome will directly affect the IETF language tag format (BCP 47), which is employed as the value of ID: 156
/ Poster Session Lightning Talks: 11
Late-breaking Poster Keywords: Arabic, Japanese, Text orientation, Yoan Udagawa, 18th century At the junction of the two orientations: TEI for the Arabic inspired by Yoan Udagawa Okayama University, Japan This study focuses on the limitations of right-to-left (RTL) editing in the editing software oXygen and proposes practical methods for improving annotation input of Arabic text in TEI. Although the software provides an RTL display mode, the tagging workflow remains optimized for left-to-right languages, making rapid markup difficult and hindering researchers from creating TEI/XML of Arabic materials. To address this usability gap, this project explores layout-based strategies such as inserting line breaks at the word level, performing bottom-to-top range selections, and dividing text into smaller, directionally stable units to improve the efficiency of RTL annotation. From the 18th century onward, Japan began to adopt practical knowledge, such as medicine, from the West, particularly from countries like the Netherlands. During this process, dictionaries of Japanese corresponding to Western languages were written, but there was a problem: Japanese is written vertically and read from right to left, while Western languages are written horizontally and read from left to right. Japanese scholars such as Yoan Udagawa experimented with various layouts to accommodate languages with different writing orientations on the same page. This progressed from the first stage, where Western languages were written horizontally and Japanese vertically on vertically ruled notebooks, to the fourth stage, still in use today, where Western languages were written horizontally and Japanese horizontally from left to right on horizontally ruled notebooks. While it can be argued that the traditional Japanese writing style has been lost, it offers great convenience, as both languages can be read in the same orientation. Meanwhile, Yoan Udagawa conducted experiments on various layouts for Arabic text. This study, drawing inspiration from his ideas, employs methods such as arranging Arabic texts vertically with line breaks between each word and evaluates their usability to explore better methods for inputting mixed Eastern and Western language texts with TEI. ID: 158
/ Poster Session Lightning Talks: 12
Late-breaking Poster Keywords: TEI, Large Language Models, Human-in-the-Loop, Date Normalisation, Shōsōin Document Date Normalisation at Semantic Document Boundaries: LLM Capabilities and Limits in Historical TEI Encoding 1: National Museum of Japanese History, Japan; 2: Historiographical Institute, The University of Tokyo, Japan Accurate date normalisation—mapping partial or abbreviated date expressions to canonical era–year–month–day forms—is a prerequisite for reliable `<date>` TEI encoding, and a task where LLMs offer evident potential in sparse-data historical settings. Yet systematic evidence of where automated inference succeeds and where it requires human intervention remains absent, hampering principled deployment in scholarly editing. This challenge corresponds to D4—semantic recognition—in the TEI encoding quality framework of Strutz (2026), the dimension least amenable to automation. ID: 154
/ Poster Session Lightning Talks: 13
Late-breaking Poster Keywords: cultural heritage materials, graduate education, DH ecosystem From Rare Materials Digitization to TEI Practice: Building a Text Encoding Learning and Practice Environment at Keio University Keio University, Japan Keio University has a long history of Digital Humanities activities rooted in the digitization and scholarly use of cultural heritage materials, beginning with the launch of the HUMI Project in 1996. Since 2001, DH-related teaching at Keio has expanded from undergraduate courses and postgraduate XML edition projects to graduate-level education. This institutional history should be understood against the broader background of Japanese DH, which, unlike Western contexts, where TEI has been a foundational technology of the DH ecosystem, has been characterized by a long "image-centered" period, due to the significant technical challenges of creating text data of East Asian scripts. Even before the recent advances in character recognition technologies, TEI had begun to gain wider visibility and use in Japan. More recently, those technological advances have made large-scale East Asian text data increasingly accessible. Reflecting this shift, in these two years, Keio’s DH education has begun to shift from digitization-centered activities toward a more explicit engagement with text encoding. TEI-oriented instruction now helps students understand humanities materials as structured, shareable, and reusable research data, rather than only as images or archival objects. This poster reports on the emerging transition from rare materials digitization to TEI practice. It focuses on the development of a learning and practice environment in which text encoding is connected to coursework, workshops, international exchange, and early-career researcher training. It also introduces TEIKeM, a newly established initiative at Keio Museum Commons, which aims to provide an institutional setting for sustaining TEI-related education and practice beyond individual courses. Rather than presenting a specific TEI data model or textbook, this poster examines how TEI practice can be embedded within a broader Digital Humanities ecosystem. ID: 162
/ Poster Session Lightning Talks: 14
Late-breaking Poster Keywords: Coptic, manuscript, TEI, digital scholarly edition, apocrypha Making of TEI Digital Scholarly Editions of Coptic Apocrypha: Case Studies on the Apocalypse of Elijah and the Gospel of Judas 1: University of Tsukuba, Japan; 2: Tsukuba Institute of Advanced Research, Japan The Apocalypse of Elijah and the Gospel of Judas are among the most consequential non-canonical writings preserved in Coptic, the latest stage of the Egyptian language, written in a Greek-derived alphabet and attested in several dialects (notably Sahidic, Bohairic, Akhmimic, Lycopolitan, Fayyumic). The Apocalypse of Elijah, an Egyptian Christian apocalypse possibly drawing on earlier Jewish material, survives in Akhmimic (P. Heid. Kopt. 600), Sahidic (notably the Chester Beatty codex), and Greek fragments, mostly produced in the fourth to fifth centuries; it offers a rare window onto Egyptian eschatology, the Antichrist figure (the "Son of Lawlessness"), and martyrological discourse under Roman persecution. The Gospel of Judas, recovered in the Sahidic Codex Tchacos and published in 2006, has reshaped scholarship on early "Gnostic" Christianity. Both survive only in codicologically fragile witnesses—ideal yet demanding candidates for TEI encoding. This poster presents two in-progress TEI P5 editions. Editorial interventions, such as lacunae, restorations, uncertain letters, nomina sacra, scribal corrections, are explicitly marked with <gap>, <supplied>, <unclear>, <damage>, <add>, and <subst>, each with @reason and @cert. Multiple witnesses are collated via <app>/<rdg> against a <listWit> apparatus. The header documents a custom POS taxonomy in <classDecl>, named entities in <listPerson>/<listPlace> linked to Pleiades and Trismegistos, and an explicit Coptic Scriptorium provenance for the linguistic layer. Every token is wrapped in <w> with part-of-speech (@type: ACAUS, ACONJ, ADV, N, NPROP, V…), Coptic @lemma, and @xml:lang="Greek" on loanwords, generated by the Coptic Scriptorium NLP pipeline and harmonized with our parallel editions of the Gospel of Judas and Pistis Sophia. The Apocalypse of Elijah base text required ~2,000 systematic OCR corrections (ⲛ̄/ⲙ̄ normalization, spurious supralinear strokes stripped, ⲍ↔ϩ and ⲭ↔ϫ confusions resolved against Wintermute's English translation), all transparently recorded in <editorialDecl>. The poster invites discussion on TEI customizations for under-resourced ancient languages and the philological auditability of OCR-derived editions. ID: 157
/ Poster Session Lightning Talks: 15
Late-breaking Poster Keywords: Digital humanities, Encoding inscriptions, TEI, EpiDoc, Early Irish Og(h)am: Harnessing digital technologies to transform understanding of ogham writing, from the 4th century to the 21st. Maynooth University, Ireland The OG(H)AM project involved harnessing digital tools from different fields to transform scholarly and popular understanding of ogham—an ancient script unique to Ireland and Britain that consists of strokes and notches. There are over 400 examples of large stones with ogham inscriptions from Ireland and Britain. Additionally, there are about two dozen small, portable ogham-bearing objects still extant. Although the original ogham script was probably devised for use on wood, it was later applied to stone. Later still, the writing system was adapted to suit the format of the manuscript page and survives in manuscript sources dating from as early as the ninth century and continuing up to the modern period. While the majority of evidence for ogham is found in Ireland, there are important clusters of ogham inscriptions found in Britain, predominantly in Wales, Devon and Cornwall, Scotland, and the Isle of Man. Previously, scholarship has studied the ogham evidence from Ireland and Britain independently of each other. The OG(H)AM project was a collaboration between Glasgow and Maynooth universities to produce the first digital corpus of all ogham inscriptions known to date from both Ireland and Britain. As a three-dimensional writing system, the ogham script poses several interesting challenges for text encoding. Regarding the orientation of the inscriptions: usually ogham is engraved vertically up the edges and horizontally across the top of stone pillars. However, the reading direction may vary for each inscription. Moreover, there are intriguing examples of spiral and circular inscriptions which further complicate interpretation. The ogham script serves as an invaluable case study for testing the capabilities and limits of text encoding to date. The comparison of ogham evidence from both Ireland and Britain creates connections which significantly enlighten our understanding of this unique writing system and highlights the complex linguistic heritage and cultural exchanges of these neighbouring islands. ID: 153
/ Poster Session Lightning Talks: 16
Late-breaking Poster Keywords: oral history interview, ethnographic coding, standOff, interview transcript StandOff and Ethnographically Coded Oral History Interviews Burnaby Village Museum, Canada At Burnaby Village Museum, we have a collection of ethnographically coded, oral history interview transcripts. We are applying TEI encoding to illustrate and emphasize this thematic analysis of each transcript, rather than descriptive, in-text encoding. Here are some features of TEI we are making use of:
We are using <standOff/> to encode our thematic analyses within each interview. We are using this because:
Using the codes, we intend for these TEI encoded files to become:
Things still being considered:
ID: 170
/ Poster Session Lightning Talks: 17
Late-breaking Poster Keywords: accessibility TEI Semantics for Web Inclusion University of Maryland, United States of America Within the digital humanities and cultural heritage sectors, the Text Encoding Initiative (TEI) has expanded public access to primary sources, breaking down the financial and physical barriers of traditional scholarly editions. However, within both this democratization effort and the popular FAIR (Findable, Accessible, Interoperable, and Reusable) data paradigm, the term "accessibility" typically refers to either broad public dissemination or to machine-retrievable data. Accessibility for people with disabilities remains somewhat overlooked in DH research (Pirrone et al, 2023) and the complex interactive user interfaces of TEI-powered digital editions are often a hindrance to web accessibility. Nonetheless, TEI is by design well positioned to facilitate web accessibility through its semantic markup. Granular encoding provides a somewhat untapped opportunity to make digital scholarly editions more accessible thanks to rich encoding rather than despite their complexity. This poster will showcase preliminary experimentation in this direction at the Scholarly Editing open access journal. The journal publishes small scale TEI-powered digital scholarly editions focused on recovering and expanding access to suppressed, appropriated, or understudied material text. Starting with Vol 43, published in 2026, Scholarly Editing micro-editions are compliant with WCAG 2.1 Level AA and with the Americans with Disabilities Act Title II Regulations. They also experiment with surfacing editorial (such as <supplied> , <corr>) and semantic (such as <persName>) markup to screen readers and keyboard navigation relying directly on TEI data in HTML5 custom elements via CETEIcean (Cayless and Viglianti 2018). References:
ID: 161
/ Poster Session Lightning Talks: 18
Late-breaking Poster Keywords: Waka literature, Waka poetic vocabulary, Utakotoba, Shōji Godo Hyakushu, TEI Guidelines TEI-compliant Markup Method for Waka Poetic Vocabulary: A Case Study of Shōji Godo Hyakushu Keio University, Japan Waka literature, one of the major genres of classical Japanese literature, is a verse form consisting of five phrases in a 5-7-5-7-7 moraic pattern. Waka have often been handed down as anthologies, following the model of Kokin Wakashū established in the early 10th century. Waka anthologies normally include contextual information before and after each poem, for example, Kotobagaki (headnotes, explaining the context), Kadai (poetry themes), and authorship, according to standardized formats. In recent years, foundational research into structuring TEI-compliant text data for Waka literature based on these traditional formats has progressed, and several texts have been made publicly available. However, there has been insufficient investigation into the markup of the vocabulary that forms the core of Waka expression. This specialized vocabulary, known as Utakotoba (Waka poetic vocabulary), is a crucial element that carries specific imagery and rhetorical functions within the linguistic space of Waka. In the commercial database currently used by many researchers, it is difficult to investigate the vocabulary based on the meaning, because search function of the database typically rely on simple keyword matching. Furthermore, dictionaries and indexes for Utakotoba are available only in print, with no free, open-access databases. As a case study of how to solve these problems, this presentation proposes a TEI-compliant markup method for Utakotoba, using the Shōji Godo Hyakushu (The second of two anthologies of one hundred Waka poems compiled during the Shōji era in the early 13th century). Furthermore, with the aim of compiling a dictionary of Utakotoba linked directly to the Waka text, this proposes a method for structuring vocabulary entries and their contents as data in accordance with the TEI Guidelines by classifying the vocabularies based on meaning, assigning unique XML IDs. ID: 155
/ Poster Session Lightning Talks: 19
Late-breaking Poster Keywords: LLM assisted encoding, TEI P5, Workflows Unsettling Manual Encoding and Creating Connections in Historical Natural History Texts: LLM-Assisted Entity Annotation in a TEI Edition of Eighteenth-Century Texts Academy of Sciences and Humanities in Lower Saxony, Germany This poster presents a series of experiments in LLM-assisted TEI P5 annotation carried out on Johann Friedrich Blumenbach’s “Handbuch der Naturgeschichte” within the broader context of Blumenbach-Online, a long-running digital edition and cultural heritage project at the Academy of Sciences and Humanities in Göttingen. The project has demonstrated both the scholarly value and the practical limits of deeply manual TEI workflows: while expert encoding ensures high semantic precision, it is costly, slow, and difficult to scale across large corpora of editions, translations, and related materials. Therefore we argue for hybrid workflows combining expert review with computational assistance. Our poster concentrates on experiments in which a selection of large language models was used to support Level 5 semantic markup of an eighteenth-century German text from the Natural History domain. The task included identifying and encoding persons, places, and natural-historical objects in TEI P5 XML; assigning authority links for persons (GND) and places (Getty TGN); generating stable IDs for recurring natural-historical entities and validating output against a project-specific Relax NG Compact schema. The quality of the outputs was assessed using a multi-dimensional evaluation framework against the human-guided annotations. Experiments showed that the LLMs were useful in their ability to follow custom TEI schema rules, given the proper few-shot examples. Zero-shot prompting however, led to non-compliant XML outputs. At the same time, the workflow also exposed typical weaknesses of generative systems: limited context windows, overfitting and occasional hallucinatory references. The results suggest that LLMs are most productive not as autonomous encoders, but as interactive editorial assistants embedded in a tightly constrained TEI workflow. In this model, LLMs improve recall and accelerate exploratory markup, while scholarly supervision remains essential for validation, disambiguation, and epistemically responsible encoding. We argue that such hybrid workflows can create help new connections and encoding in efficient and productive ways. | ||