Multi-format publishing and repurposing of historical linguistics data

Bermúdez-Sabel, Helena
University of Neuchâtel, Switzerland
helena.bermudez@unine.ch

Table of contents

1. Introduction

This paper stems from a project that studies modality in Latin from a diachronic perspective (WoPoss 2019).1 We present the outputs produced by the project and how they are tailored to different purposes. Our strategy can be easily implemented by other projects dealing with similar issues: heterogeneous sources and multiple layers of annotation that require different workflows (e.g., the combination of automatic and manual annotation).

In recent years, much focus was put into the definition of data management practices in the scientific context. A significant milestone in the creation of a common framework for the construction and administration of scientific data was reached with the publication of the FAIR principles (Wilkinson et al. 2016). In addition to these cross-disciplinary guidelines, specific recommendations that cover the particularities of Humanities research have been presented (Harrower et al. 2020). Thus, complying with the principles of open science, our project has been implementing data sharing policies from its early stages (Dell’Oro et al. 2020).

We will first present the different phases in which our data processing is articulated. We will then focus on how the publication of our materials in multiple formats increases its reusability and interoperability.

2. An open science workflow

Our data collection begins by gathering an initial dataset from different open, online resources. Besides taking into consideration the diachronic perspective, we are interested in semantic change that may have emerged in any of the Latin varieties. This means that our corpus needs to be representative in terms of chronology, provenance, source typology and textual genre, among other sociolinguistic factors. This requirement compels us to collect texts from heterogeneous resources that provide their data in different formats. The collected texts are converted into plain text, but pseudo-markup is added to preserve relevant information that was previously described in the markup.

These plain text files are then automatically annotated using Stanza (Qi et al. 2020). The results are output in CONLL-U format (Universal Dependencies contributors 2020). These files are then uploaded to the annotation platform INCEpTION (Klie et al. 2018). Using this platform, we annotate each modal passage according to our Guidelines (Dell’Oro 2019). A modal passage is identified on the basis of the presence of a predefined list of modal markers (eg. the verb possum or the gerundive suffix -ndum) (Fig. 1).

After the annotation is curated, the results are exported in the UIMA CAS XMI format; the advantages of this format to formalize linguistic annotations are discussed in Eckart de Castillo et al. (2017). UIMA provides a standard for annotating unstructured data that enables the interaction between multiple layers of annotation, and CAS is the data structure through which all UIMA components communicate.

Example of modal passages annotated in INCEpTION

Although UIMA CAS XMI is a format used in the NLP community, its use is not as widespread among Humanities researchers as other formats. Thus, to increase the reusability of our dataset, we perform a post-processing that transforms this format into TEI. Our initial pseudo-markup also gets formalized as TEI elements at this stage. Metadata about the work (including information about the author) is also included, together with philological information provided by our annotators during the annotation process (in particular, in the case the used text deviates from the edition of reference we have chosen).

Together with the annotation, we also make a detailed study of the modal markers focusing on their semantic change. The purpose of this study is twofold: First, enriching the annotation by adding the most ancient modal meaning of each marker. Second, creating diachronic interactive visualizations of the semantics of each marker (Bermúdez-Sabel et al. 2020). These maps will be enhanced with the results of our empirical corpus-based analysis, once the annotation will be completed.

In addition, we will create an ontology, compliant with the Linking Latin (LiLa) model (Passarotti et al. 2020). Our goal is to convert our annotations to an RDF serialization that can be integrated in the LiLa Knowledge Base.

The complete annotated dataset will be accessible through a user-friendly interface (GUI) with faceted search, which will exhibit and exploit the fine-grained annotation. All the code and the resources that are used during development are open-source (WoPoss 2020). Our own programs are made openly available, and all external software the project relies on is also open source.

3. Multi-format publishing as a reusability strategy

The use of open standards is key for data sharing (Harrower et al. 2020: 19–20). We publish our data in different formats so that they can be repurposed according to the needs of different communities:

4. Conclusion

In this paper we discussed the advantages of publishing in different formats, at every step of development, the data constructed in the context of a historical linguistics project. We argue that this strategy increases the possibility of repurposing data, thus becoming a more useful asset to the scientific community.

Appendix A

Bibliography
  1. Bermúdez-Sabel, Helena / Dell’Oro, Francesca / Marongiu, Paola (2020): "Visualisation of Semantic Shifts: The Case of Modal Markers", in: Estill, Laura / Guiliano, Jeniffer (eds.): 15th Annual International Conference of the Alliance of Digital Humanities Organizations, DH 2020. Conference Abstracts, Ottawa, Canada July 20-25, 2020 DOI: 10.17613/SCY4-BR70 <https://dh2020.adho.org/wp-content/uploads/2020/07/165_Visualisationofsemanticshiftsthecaseofmodalmarkers.html> [25.05.2021].
  2. Dell’Oro, Francesca (2019): WoPoss Guidelines for Annotation. DOI: 10.5281/zenodo.3560950.
  3. Dell’Oro, Francesca / Bermúdez Sabel, Helena / Marongiu, Paola (2020): "Implemented to Be Shared: The WoPoss Annotation of Semantic Modality in a Latin Diachronic Corpus", in: Gabay, Simon / Herrmann, Berenike / Hodel, Tobias / Schulthess, Sara / Spadini, Elena / Tasovac, Toma (eds.): Sharing the Experience: Workflows for the Digital Humanities. Proceedings of the DARIAH-CH Workshop 2019 (Neuchâtel), Neuchâtel, Switzerland, December 2019 DOI: 10.5281/zenodo.3739440.
  4. Eckart de Castilho, Richard / Ide, Nancy / Lapponi, Emanuele / Oepen, Stephan / Suderman, Keith / Velldal, Erik / Verhagen, Marc (2017): "Representation and Interchange of Linguistic Annotation. An In-Depth, Side-by-Side Comparison of Three Designs", in: Association for Computational Linguistics (ed.): Proceedings of the 11th Linguistic Annotation Workshop. Valencia, Spain, April 2017 67–75 DOI: 10.18653/v1/W17-0808.
  5. Harrower, Natalie / Maryl, Maciej / Biro, Timea / Immenhauser, Beat (2020): Sustainable and FAIR Data Sharing in the Humanities: Recommendations of the ALLEA Working Group E-Humanities. Berlin: ALLEA Working Group E-Humanities DOI: 10.7486/DRI.TQ582C863.
  6. Klie, Jan-Christoph / Bugert, Michael / Boullosa, Beto / Eckart de Castilho, Richard / Gurevych, Iryna (2018): "The INCEpTION Platform: Machine-Assisted and Knowledge-Oriented Interactive Annotation", in: Association for Computational Linguistics (ed.): Proceedings of the 27th International Conference on Computational Linguistics: System Demonstrations. August 2018 5–9 <http://tubiblio.ulb.tu-darmstadt.de/106270/> [25.05.2021].
  7. Passarotti, Marco / Mambrini, Francesco / Franzini, Greta / Cecchini, Flavio M. / Litta, Eleonora / Moretti, Giovanni / Ruffolo, Paolo / Sprugnoli, Rachele (2020): "Interlinking through Lemmas. The Lexical Collection of the LiLa Knowledge Base of Linguistic Resources for Latin", in: Studi e Saggi Linguistici 58, 1: 177–212 DOI: 10.4454/ssl.v58i1.277.
  8. Qi, Peng / Zhang, Yuhao / Zhang,Yuhui / Bolton, Jason / Manning, Christopher D. (2020): "Stanza: A Python Natural Language Processing Toolkit for Many Human Languages", in: Association for Computational Linguistics (ed.): Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, July 2020 101–108 <https://nlp.stanford.edu/pubs/qi2020stanza.pdf> [25.05.2021].
  9. Siegel, Erik / Retter, Adam (2014): EXist: A NoSQL Document Database and Application Platform. Sebastopol, California: O’Reilly Media.
  10. Universal Dependencies contributors (2020): CoNLL-U Format <https://universaldependencies.org/format.html> [25.05.2021].
  11. Wilkinson, Mark D. et al. (2016): "The FAIR Guiding Principles for Scientific Data Management and Stewardship", in: Scientific Data 3: 160018 DOI: 10.1038/sdata.2016.18.
  12. WoPoss (2019): A World of Possibilities. Modal Pathways over an Extra-Long Period of Time. The Diachrony of Modality in the Latin Language <https://woposs.unine.ch> [25.05.2021].
  13. WoPoss (2021): The Interactive Visualisation of Semantic Modal Shifts <https://woposs.unine.ch/semantic-modal-maps.php> [25.05.2021].
  14. WoPoss (2020): WoPoss – GitHub <https://github.com/WoPoss-project> [25.05.2021].
Notes
1.
This work is supported by the Swiss National Science Foundation [176778] and led by Francesca Dell’Oro at the University of Neuchâtel.