Ferger, Anne
Universität Paderborn / Musikwissenschaftliches Seminar Detmold/Paderborn
anne.ferger@uni-paderborn.de
Jettka, Daniel
Universität Paderborn / Musikwissenschaftliches Seminar Detmold/Paderborn
daniel.jettka@uni-paderborn.de
The collaborative creation of high quality linguistic corpora is accompanied by multiple practical and methodological challenges, some examples are the absolutely basic prevention of data loss, control of single file changes and global corpus versions. For corpora, which contain manually compiled transcription and translation layers, annotations, and complex metadata, considerations regarding the quality and consistency of the inter-dependent information play a crucial role in creating scientifically relevant resources.
This contribution1 deals with the implementation and optimization of workflows for the creation of complex corpora for the documentation of indigenous languages, and particularly discusses benefits of automating quality assurance mechanisms. The approach is extended by a quantification method which can be used to measure its impact on the overall quality of the corpora.
Linguistic corpora are usually created collaboratively. Consistent version control mechanisms facilitate the collaboration in corpus creation processes and are a basic necessity considering the amount of time and work spent on building scientific resources. Minimizing the risk of losing results and maximizing re-usability and reproducibility have to be priorities.
In the project presented here numerous linguistic corpora are created for documenting indigenous languages spoken on the territory of the Russian Federation. The corpora contain transcriptions of spoken and written language, translations into Russian, English, and German, morphologic, syntactic and semantic information, as well as complex speaker and corpus metadata, all of which are mainly derived in a costly manual process.2
The allocation and completion of particular tasks in the corpus creation process is organized with an established project management tool connected to the applied version control software. The consistent and reproducible data creation is supported by measures for optimizing existing workflows and enhancing the corpus quality (cf. Hedeland / Ferger 2020).
The notion of quality for linguistic corpora used in this context includes general FAIR metrics, as well as consistency and structure in metadata, transcription conventions and annotation schemes, allowing for reliable automatic processing (cf. Hedeland 2020).
Assuring the quality of produced corpus data as well as optimizing the underlying workflows lead to faster production and higher quality of resources. Work on quality assurance of research data is for instance also conducted in the projects Conquaire (Hermann et al. 2021) and QUEST (Arkhangelskiy et al. 2020), the latter building upon the framework discussed here. The framework for quality assurance was developed using an open-source, extensible software system3 (Ferger et al. 2020) and integrating it into existing versioning workflows. The functionality of the software consists of methods assisting with manual tasks as well as applying automatic changes and corrections to the data, for enhancing the quality of the created resources.
The necessary manual improvement of the data quality (which is unavoidable in manually created data) is supported by the automatic identification of problems and the generation of reports and correction lists, which for instance can be used for navigating to the problematic phenomena directly in the transcription software. The automatic check mechanisms are applied periodically and information about absolute numbers of found errors is documented in overviews reflecting the current status of the corpora (see Figure 1). At the moment, rule-based heuristics are applied (e.g. checking against controlled metadata vocabularies, transcription conventions, annotation schemes, or custom patterns)4.

Corpus curation status showing error numbers in different creation phases
In addition to reporting, certain recurring systematic errors are identified and fixed automatically. This for instance includes the adherence to basic transcription conventions (e.g. use of whitespaces and punctuation) and file naming conventions.
The consistent version control, including information about automatically executed corrections, enables the tracking of error and correction rates for different versions of the corpora over time. The possibility to re-instantiate all editing states of the corpora also allows for the extraction of this information for corpus versions for which the number of applied fixes was not tracked yet.5 While it is possible to re-generate the needed numbers retroactively, it is obviously easier to collect them directly during the fixing operations.
Thus the quality assurance measures were complemented by saving counts of automatic fixes done to the corpus data in a structured format6 that can directly be used for further analysis. The collected data contains the types of fixes that caused changes to the data, the date of their application and the number of the corrections performed. For the exploration of the quantitative and temporal information about the quality assurance and further analysis an interactive dashboard was implemented that allows for creating different visualizations. Overviews of checks and fixes can be displayed chronologically, or sorted by type. This way it is possible to see if certain types of errors are typical for certain corpus creation phases.

Dashboard for quality assurance evaluation
This contribution shows an approach to evaluate specific measures for quality assurance in the creation of complex linguistic corpora for the documentation of indigenous languages. Besides giving a basic idea about the applied methods, a quantification mechanism is presented that complements the positive intuition of the significance of the efforts taken to optimize the overall quality of the created corpora. By tracking information about error and automatic correction rates and importing it into a search engine, it is possible to get clearer insights into the development of the corpus quality and the actual effect of the quality assurance methods.
After demonstrating the concept of building a quantification mechanism for enhancements in quality control measures, the approach should now be widened to take into account not only directly accessible information from existing error reports, but also previous corpus versions or even other corpora which were not yet covered by systematic reporting.
Further operationalization should take place in future work on this topic. This can be as simple as correlating the size of a corpus to the error and correction rates, or more complex like taking into account temporal, infrastructural, or personal factors for the creation of corpora.