Costa, Rute
NOVA CLUNL, Centro de Linguística da Universidade NOVA de Lisboa
rute.costa@fcsh.unl.pt
Salgado, Ana
NOVA CLUNL, Centro de Linguística da Universidade NOVA de Lisboa; NOVA CLUNL, Centro de Linguística da Universidade NOVA de Lisboa
anacastrosalgado@gmail.com
Bruno Almeida, Bruno
NOVA CLUNL, Centro de Linguística da Universidade NOVA de Lisboa; ROSSIO Infrastructure
brunoalmeida@fcsh.unl.pt
Lexicography has radically changed over the last two decades, especially with the ongoing transition to digital resources. This change has led to a paradigm shift with the advancement of digital humanities, the growing importance of standards, and the need to interconnect available lexicographical data. The application of computing and modelling has become inevitable in lexicography.
Lexicography has traditionally been understood as the art and craft of compiling general language dictionaries (Landau 2001), which is often seen as a branch of applied linguistics. There is, however, a more holistic approach, which embraces lexicography’s relationships with lexicology, terminology, encyclopaedias and information science. According to this broader view, metalexicography “should be regarded as part of information science” (Wiegand 2013: 14). More than describing the lexicon of languages, the purpose of lexicography is to “resolve specific types of information needs detected in society” (Trap-Jensen 2018: 22). Indeed, it can be argued that lexicography aims “in a more general way at the production of information tools” (Bergenholtz / Gouws 2012: 40), i.e. reference works currently focused on “enhanced information retrieval” (ibid.). Information science is seen as an interdisciplinary field concerned with “the origination, collection, organization, storage, retrieval, interpretation, transmission, transformation and use of information” (Borko 1968: 3). Information includes all encoded representations (in natural language or other modalities) which can be transmitted, stored and organised for subsequent retrieval. While the ties between terminology science and information science were noted from the start of terminology as a contemporary subject of inquiry, the ties between information science and lexicography may be less obvious.
Lexicographical reference works are legitimately part of digital humanities. Historical lexicographical resources are valuable resources for historical and diachronic linguistics, whose heritage must be preserved and made available online in open access to the community of researchers in the humanities. One example of such a resource is the Vocabulário Ortográfico da Língua Portuguesa [Orthographic Vocabulary of the Portuguese Language] (VOLP-1940), the first orthographic vocabulary published by the Academia das Ciências de Lisboa (ACL) in 1940, which was the basis for the common Portuguese orthographic vocabulary adopted in Portuguese speaking countries. This research also aims at filling a gap in Portuguese lexicography, given that lexicographical resources are still scarce, especially legacy dictionaries made available on the Web as open access resources (Williams 2019: 83).
This paper falls within the domain of application of digital lexicography in the context of a scholarly editing project and is based on a set of methodological and theoretical assumptions about which we will make some considerations. We will focus on the organisation of linguistic information of a lexicographical nature within the field of digital humanities, emphasising the development of a cross-disciplinary methodology that combines lexicography, information science and computational methods. We propose to present a mixed methodology, the objective of which is to associate the organisation of knowledge with the processing of lexicographical data, to allow the sharing of data with the goal of exporting them in a digital format.
The work described in this paper is part of a larger project that involves the digitisation of all vocabularies of the ACL and their computational analysis. This will allow to better assess the importance of the ACL vocabularies for the evolution of the Portuguese language and will contribute to the contemporary movement of creating innovative and data-driven computational methods for text digitisation, encoding, analysis and organisation of lexicographical data. The digitisation of VOLP-1940 allows for the creation of a lexicographical resource encoded in TEI Lex-0 (Tasovac et al. 2018), a subset of the TEI Guidelines for encoding dictionaries, with structured information in SKOS. Adherence to the FAIR – Findable, Accessible, Interoperable, Reusable – principles (Wilkinson et al. 2016) will ensure the future connection of VOLP-1940 to other existing systems and resources, in particular from the Portuguese-speaking world.
The front matter of VOLP-1940 includes a list of abbreviations and conventional signs (VOLP-1940: LXXXIX–XCII). We have focused on organising this list for computational processing using SKOS and TEI Lex-0 to ensure future interoperability. In the paper version, this list is sorted alphabetically and is divided in two parts: (i) list of abbreviations and (ii) list of conventional signs. The list shows the abbreviation or conventional sign followed by its full form. Our analysis focused on the first part, from which we drew up a classification consisting of the 220 abbreviations that make up the list. Although well organised in two columns, this list is static and has some limitations inherent to the paper format. A simple alphabetical list, whose original form is kept on the website of this project, was organised and modelled for the digital environment and enriched with linguistic annotations.
After a thorough analysis of the abbreviations, the following types have been identified: part of speech, onomastic classification, grammatical gender, grammatical number, language, register, tense, etymology, word formation, and others. These categories thus constitute what we call the typological organisation of the list of abbreviations. We have further built a SKOS model of the above-mentioned categories, as will be shown in this paper. In the transition to digital, we reorganised the original content of the list of abbreviations for processing reasons and to ensure its future interoperability. From the complete list, we have isolated those related to word classes. Based on this list, and for interoperability with other lexicographical resources, we have made the correspondence between word classes and the Universal Dependencies Part of Speech tags, which will be exemplified in this paper. The SKOS model will supply URIs for each part of speech, and remaining categories, which will be used in the TEI Lex-0 encoding of the lexicographical entries.
We consider that the work we have done so far highlights the need to change traditional lexicographical practices. Many of the principles now defined will be used for the annotation of the remaining entries and will be applied in the retrodigitisation of subsequent works, as they share several of the typographic conventions now identified. With the methodology described in this paper, we intend to represent the ever-increasing synergy between lexicographers, terminologists, computational linguists, information experts and digital humanists that we so keenly advocate.
The work described in this paper has the following objectives: (i) create a new online lexicographical resource, accessible to the entire scientific community and the general public; (ii) improve the consistency of the original metadata, following an exact linguistic annotation strategy in line with TEI recommendations, while ensuring data are accessible and reusable; (iii) organise metadata information according to SKOS; (iv) describe the linguistic annotation for further semantic enrichment of the database; (v) add new metadata, namely domain names, information that will be recovered from other lexicographical works that contain this annotation, and make the connection between several synonymous units that are included in the work’s word list.
This paper will cover 4 main sections:
1. The first part will address the paper’s theoretical framework where we will analyse the relationship between digital humanities, information science and lexicography, exploring new possibilities and some challenges. The notions of interoperability and standards will be discussed, and we will argue for their importance in lexicography.
2. In the second part, we will report the development of the methodology we used to analyse the list of abbreviations.
3. In the third part, we address the modelling in SKOS and TEI Lex-0 encoding.
4. In the last part, we will discuss the methodology used and the results obtained.
Keywords: digital humanities, information science, lexicography, Portuguese language