TEI 2026
Creating Connections, Unsettling Practices
August 10-14, 2026
University of British Columbia, Vancouver, BC, Canada
Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
|
Daily Overview |
| Session | |
|
Poster Session Lightning Talks Location: Koerner 5th floor common area Session Chair: Syd Bauman, Northeastern University | |
| Presentation 6 | |
ID: 110
/ Poster Session Lightning Talks: 6
Poster Keywords: TEI CMC, social media, Japanese language corpora, microblogging, messaging Modeling Japanese Social Media with TEI CMC: Integrating Microblogging and Messaging in BCCWJ2 National Institute for Japanese Language and Linguistics, Japan This poster presents the integration of social media data into the Balanced Corpus of Contemporary Written Japanese 2 (BCCWJ2) and examines how heterogeneous computer-mediated communication (CMC) environments can be modeled within the TEI CMC framework. The original Balanced Corpus of Contemporary Written Japanese (BCCWJ), released in 2011 by the National Institute for Japanese Language and Linguistics, is a publicly available, genre-balanced corpus of modern written Japanese, including books, magazines, white papers, and web texts. It has served as a foundational resource for Japanese linguistics. However, most of its materials date from 2005 or earlier. To better reflect contemporary language use, BCCWJ2 (2024–2028) aims to expand the corpus to approximately 200 million words by incorporating texts published between 2006 and 2025, including digital communication that now plays a central role in everyday written practices. A key challenge is the inclusion of social media data. Our project integrates both microblogging platforms (Bluesky and Misskey) and a messaging service (LINE). While the TEI CMC Guidelines provide models for Twitter-like services and for messenger-style interaction, actual platforms present service-specific complexities that do not always align with text-centered encoding assumptions. Particular issues arise in encoding platform-specific reactions and visual elements. Misskey supports custom, non-Unicode emoji used as reactions, and LINE employs stickers and rich reaction systems that cannot be reduced to plain text. These semiotic elements function as discourse acts yet resist straightforward textual representation. By comparing microblogging and messaging ecologies within a single corpus design, this project evaluates how far TEI CMC can accommodate diverse social media practices and where customization may be required. We argue that constructing a balanced corpus including such data contributes to documenting contemporary Japanese linguistic life and promotes sustainable, interoperable, and globally accessible Japanese language resources. | |
