Interlinking through Lemmas. Linking Linguistic Resources to a lexically-based LLOD Knowledge Base

Passarotti, Marco Carlo
Università Cattolica del Sacro Cuore, Italy
marco.passarotti@unicatt.it

Franzini, Greta
Eurac Research, Italy
greta.franzini@eurac.edu

Litta, Eleonora
Università Cattolica del Sacro Cuore, Italy
eleonoramaria.litta@unicatt.it

Mambrini, Francesco
Università Cattolica del Sacro Cuore, Italy
francesco.mambrini@unicatt.it

Sprugnoli, Rachele
Università Cattolica del Sacro Cuore, Italy
rachele.sprugnoli@unicatt.it

Table of contents

1. Objectives

The tutorial aims to introduce the architecture, use and enhancement of the LiLa Knowledge Base of interlinked linguistic resources for Latin, developed in the context of the LiLa: Linking Latin project. In particular, the tutorial will present how the Linked-Data model adopted by LiLa is used to connect distributed lexical and textual resources, to ensure their interoperability. We show how, via lemmatisation, texts become part of this network of resources. We provide participants with a theoretical introduction to the architecture of LiLa, as well as hands-on support in their interaction with the LiLa Knowledge Base.

The proposed tutorial falls within the field of Linguistic Linked Open Data (LLOD: Cimiano et al. 2020). While we focus on Latin, the methods discussed are language independent and thus have a much wider application, proving useful for similar initiatives on other languages.

Starting from the experience of LiLa, we introduce the audience to some of the most relevant topics in current digital textual studies and linguistic resources, including:

2. Detailed description

We propose a 1-day tutorial, divided in two sections.

2.1. Morning (conceptual foundations)

In this session the LiLa project will serve as a use-case to illustrate a typical LLOD workflow, that is:

  1. inguistic annotation/processing (Gleim et al. 2019; Sprugnoli et al. 2020)
  2. data modelling (McCrae et al. 2017)
  3. linking (Declerck et al. 2012).

2.2. Afternoon (practical activity)

Participants will be divided into groups, supervised by one or more presenters/assistants. We provide them with a selection of Latin raw texts from various types and genres, and guide them through the stages leading up to the connection of texts to the LiLa Knowledge Base. Specifically, we focus on how to perform automatic tokenisation, part-of-speech tagging and lemmatisation, run a custom tool to automatically transform lemmatised output into the RDF, link it to the LiLa Knowledge Base and, finally, query the Knowledge Base with SPARQL (DuCharme 2013).

In closing, all groups will come together to present their results and draw conclusions.

Data and tools necessary to participate in the tutorial will be provided ahead of the event.

2.3. Tentative schedule

09:00-09:45 – General introduction (Passarotti)

09:45-10:30 – Data model (Mambrini)

10:30-10:45 – Coffee Break

10:45-11:30 – Processing (Cecchini) and currently linked resources (Litta, Sprugnoli)

11:30-12:00 – Questions and group formation

12:00-13:30 – Lunch

13:30-16:00 – Hands-on work

16:00-17:00 – Group presentations and conclusions

3. Instructors

Presenters and assistants (all affiliated with Università Cattolica del Sacro Cuore):

Assistants:

4. Target audience and expected attendance

This tutorial is intended for those who wish to explore solutions to publish texts using LOD. Although we focus on our experience with Latin, we welcome any participant interested in the theme of textual resources and LOD; prior knowledge of Latin and/or LOD is not required.

We expect around 20-30 participants from different backgrounds: computational linguists, theoretical linguists, classicists, philologists.

5. Proposed budget

As the tutorial will be run by the members of the LiLa team, no costs to pay for the instructors is foreseen.

Neither publication of proceedings nor invitation of keynote speakers is planned. The website of the tutorial (if required to be external to that of EADH 2021) will be hosted on that of LiLa, free of charge.

Should the tutorial take place in loco, the costs for the coffee breaks will be covered by the tutorial fees.

6. Requirements and format

Should the tutorial take place in loco, we would require a projector and a whiteboard with markers.

Should it take place online, we would adopt Zoom as our preferred platform for its “Breakout Rooms” functionality.

Appendix A

Bibliography
  1. Cimiano, Philipp / Chiarcos, Christian / McCrae, John P. / Gracia, Jorge (2020): Linguistic Linked Data: Representation, Generation and Applications. Cham: Springer.
  2. Declerck, Tierry / Lendvai, Piroska / Mörth, Karlheinz / Budin, Gerhard / Váradi, Tamás (2012): "Towards linked language data for digital humanities", in: Chiarcos, Christian / Nordhoff, Sebastian / Hellmann, Sebastian (eds.): Linked Data in Linguistics. Representing and Connecting Language Data and Language Metadata. Berlin / Heidelberg: Springer 109–116 DOI: <https://doi.org/10.1007/978-3-642-28249-2_11>.
  3. DuCharme, Bob (2013): Learning Sparql. Querying and Updating with Sparql 1.1. Sebastopol, CA: O’Reilly.
  4. Gleim, Rüdiger / Eger, Steffen / Mehler, Alexander / Uslu, Tolga / Hemati, Wahed / Lücking, Andy / Henlein, Alexander / Kahlsdorf, Sven / Hoenen, Armin (2019): "A practitioner’s view: a survey and comparison of lemmatization and morphological tagging in German and Latin", in: Journal of Language Modelling 7, 1: 152 DOI: <https://doi.org/10.15398/jlm.v7i1.205>.
  5. McCrae, John P. / Bosque-Gil, Julia /Gracia, Jorge / Buitelaar, Paul / Cimiano, Philipp (2017): "The Ontolex-Lemon model: development and applications", in: Kosem, Iztok / Tiberius, Carole / Jakubíček, Miloš / Kallas, Jelena / Krek, Simon / Baisa, Vít (eds.): Electronic lexicography in the 21st century. Proceedings of eLex 2017 conference. Brno: Lexical Computing 19–21.
  6. Sprugnoli, Rachele / Passarotti, Marco / Cecchini, Flavio M. / Pellegrini, Matteo (2020): "Overview of the EvaLatin 2020 Evaluation Campaign", in: Sprugnoli, Rachele / Passarotti, Marco (eds.): Proceedings of LT4HALA 2020. 1st Workshop on Language Technologies for Historical and Ancient Languages. Marseille: ELRA 105–110.