TEI 2026
Creating Connections, Unsettling Practices
August 10-14, 2026
University of British Columbia, Vancouver, BC, Canada
Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
|
Daily Overview |
| Session | |
|
D2 Short Papers: Macro, Standards, Big Projects Location: BUCH B210 Session Chair: Elisa Beshero-Bondar, Penn State Erie | |
| Presentation 3 | |
ID: 125
/ D2 Short Papers: 3
Short Paper Keywords: text encoding, text analysis, women and gender studies Transforming TEI Documents for Text and Network Analysis Northeastern University, United States of America This paper will demonstrate some ways that the rich information modeled in TEI documents can be used for text and network analysis. The Women Writers Project (WWP), a long-term research project focused on text encoding and early women’s writing, has recently begun developing methods for using encoded texts in computational analyses. As part of the WWP’s work on making word embedding models more accessible for humanists, we published a set of routines that use XSLT and XQuery to prepare TEI files for computational text analysis. These routines provide a set of fine-tuning options that are targeted to the needs of natural language processing applications. For example, the contents of notes can be placed either at their points of anchor or in a separate section in the plain-text output files; the transformation routines also enable researchers to select and omit any elements whose contents are likely to distort results in trained models. In a current project studying the impacts of women on the development of the sciences, the WWP team has built two new transformation routines, both in XQuery. One expands on the plain-text generation process described above to allow selection of researcher-defined “documents” for topic modeling, based on how texts’ structures are modeled in the markup. The other uses bibliographic markup to extract information on co-citation within texts and produce edge lists for network analysis. This paper will share some insights that the WWP has developed through our efforts to realize the potential of TEI documents for computational analysis. It will also offer some strategies for projects interested in taking up such work with their own data, with approaches for adapting existing resources, making the most of provisional solutions, and iterating between data retrieval and research design. | |
