Corpus Linguistics with Semantically Tagged Data: A Case Study of the Human Remains Digital Library

Gribomont, Isabelle
University of Liverpool, United Kingdom
isabelle.gribomont@liverpool.ac.uk

Understanding the history of exhumation and reburial is key to shaping sound ethical practices and policies today. Nevertheless, the textual material narrating the history of Britain's dead has never been studied systematically. The Human Remains Digital Library project aims to gather and analyse the documents related to exhumation in Britain from the 7th to the 19th century. Ultimately, the project will inform policymaking regarding the handling of human remains in Archaeological and Heritage Management settings. We will be collaborating with various stakeholders such as cathedral foundations which are interested in acquiring new knowledge about their burial grounds and establishing research-informed practices when it comes to handling human remains. Our Digital Library will include thousands of files ranging from newspaper articles and scientific papers to pamphlets, ecclesiastical documents and historical accounts. The database will be made available for scholars and practitioners alike at the end of the project.

The project, currently in the corpus building and pilot analysis stage, is using digital methods from the fields of Corpus Linguistics and Natural Language Processing to analyse the data. Such methods allow us to uncover significant patterns of meaning throughout the entire corpus and within subcorpora of particular interest. We identify trends in the language used to describe, justify and argue about the exhumation of human remains. More specifically, we aim to mine the opinions and sentiments triggered by such issues and map their evolution over time. Because debates and conversations about exhumation cut across spiritual, ethical and scientific concerns, our corpus affords a unique and privileged window into the tensions between religious beliefs on the one hand and scientific inquiries on the other. Identifying the textual strategies used to encode such attitudes and understandings, as well as mapping the ways in which these tensions were negotiated linguistically, provide unique insights in historical sensitivities.

We are particularly interested in the possibility of using corpus linguistics techniques with semantically tagged historical data, whether via pre-existing corpus linguistics software or Python code. As argued by Alexander et al. (2015), “[t]ruly effective searching of text is currently hindered by a need to search using word forms, while in actual use almost all searches are aimed at the ‘meaning’ behind that word form”. Tagging a corpus with semantic categories allows for big picture patterns to emerge, without being diluted by synonymity or lexical complexity. Although significant advances have been made in recent years, including the UCREL Semantic Analysis System (USAS) developed by Rayson et al. (2004) and the Historical Thesaurus Semantic Tagger developed by Piao et al. (2017), it is not the norm for corpus linguistics studies to make use of deep semantic tagging of each word or lexical unit.

This poster will focus on a pilot study of a sample corpus comprised of modern English translations of medieval texts related to exhumation. The sample corpus has been tagged with the USAS tagger. The outcome of the preliminary corpus analysis will be presented, with an emphasis on the insights brought by the semantic tags. In addition to questions related to ethics, morality and science, the analysis zooms in the language used to make sense of the exhumation and re-burial processes. For instance, we are investigating the use of anatomical language, expressions of surprise or expectation regarding conservation and decay, the ritualisation of the process of exhumation, the presence and role of 'supernatural' elements, the description of sensory experiences of the body, and the emotional dimension of the accounts.

Moreover, the poster will outline the performance assessment of the tagger in the context of our data. This assessment is based on the manual evaluation of the tags for the most common word forms and a representative sample of all semantic categories in our corpus by three subject-specialists.

With this poster, we hope to spark a practical conversation between scholars working with computational text analysis methods, with a particular interest on the present and future of automated semantic tagging in Digital Humanities.

Appendix A

Bibliography
  1. Alexander, Marc / Dallachy, Fraser / Piao, Scott / Baron, Alistair / Rayson, Paul (2015): "Metaphor, Popular Science, and Semantic Tagging: Distant Reading with the Historical Thesaurus of English", in: Digital Scholarship in the Humanities 30: 16-27.
  2. Piao, Scott / Dallachy, Fraser / Baron, Alistair / Demmen, Jane / Wattam, Steve / Durkin, Philip / McCracken, James / Rayson, Paul / Alexander, Marc (2017): "A Time-Sensitive Historical Thesaurus-based Semantic Tagger for Deep Semantic Annotation", in: Computer Speech & Language 46: 113-135.
  3. Rayson, Paul / Archer, Dawn / Piao, Scott / McEnery, Tony (2004): "The UCREL semantic analysis system", in: LREC (ed.): Proceedings of the workshop on Beyond Named Entity Recognition Semantic labelling for NLP tasks in association with 4th International Conference on Language Resources and Evaluation (LREC 2004), 25th May 2004, Lisbon, Portugal 7-12 <https://www.lancaster.ac.uk/staff/rayson/publications/usas_lrec04ws.pdf> [21.08.2021].