<?xml version="1.0" encoding="UTF-8"?>
<TEI xmlns="http://www.tei-c.org/ns/1.0">
    <teiHeader>
        <fileDesc>
            <titleStmt>
                <title type="full">
                    <title type="main">Evaluating the benefits of a quality assurance
                        framework</title>
                </title>
                <author>
                    <persName>
                        <surname>Ferger</surname>
                        <forename>Anne</forename>
                    </persName>
                    <affiliation>Universität Paderborn / Musikwissenschaftliches Seminar
                        Detmold/Paderborn</affiliation>
                    <email>anne.ferger@uni-paderborn.de</email>
                </author>
                <author>
                    <persName>
                        <surname>Jettka</surname>
                        <forename>Daniel</forename>
                    </persName>
                    <affiliation>Universität Paderborn / Musikwissenschaftliches Seminar
                        Detmold/Paderborn</affiliation>
                    <email>daniel.jettka@uni-paderborn.de</email>
                </author>
            </titleStmt>
            <editionStmt>
                <edition>
                    <date>2021-06-09T18:13:21.877602918</date>
                </edition>
            </editionStmt>
            <publicationStmt>
                <publisher>Elisabeth Burr, University of Leipzig</publisher>
                <address>
                    <addrLine>Beethovenstr. 15</addrLine>
                    <addrLine>04107 Leipzig</addrLine>
                    <addrLine>Germany</addrLine>
                    <addrLine>Elisabeth Burr</addrLine>
                </address>
            </publicationStmt>
            <sourceDesc>
                <p>Converted from an OASIS Open Document</p>
            </sourceDesc>
        </fileDesc>
        <encodingDesc>
            <appInfo>
                <application ident="DHCONVALIDATOR" version="1.22">
                    <label>DHConvalidator</label>
                </application>
            </appInfo>
        </encodingDesc>
        <profileDesc>
            <textClass>
                <keywords scheme="ConfTool" n="category">
                    <term>Paper</term>
                </keywords>
                <keywords scheme="ConfTool" n="subcategory">
                    <term>Short Paper</term>
                </keywords>
                <keywords scheme="ConfTool" n="keywords">
                    <term>Quality Assurance</term>
                    <term>Corpus Creation</term>
                    <term>Project Management</term>
                    <term>Data Analysis</term>
                </keywords>
                <keywords scheme="ConfTool" n="topics">
                    <term>Discovering</term>
                    <term>Gathering</term>
                    <term>Designing</term>
                    <term>Cleanup</term>
                    <term>Structural Analysis</term>
                    <term>Modeling</term>
                    <term>Identifying</term>
                    <term>Collaboration</term>
                    <term>Meta: ProjectManagement</term>
                    <term>Infrastructure</term>
                    <term>Visualisation</term>
                    <term>Methods</term>
                    <term>Metadata</term>
                    <term>ResearchProcess</term>
                    <term>ResearchResults</term>
                    <term>English</term>
                </keywords>
            </textClass>
        </profileDesc>
    </teiHeader>
    <text>
        <body>
            <div type="div1" rend="DH-Heading1">
                <head> Introduction </head>
                <p>The collaborative creation of high quality linguistic corpora is accompanied by
                    multiple practical and methodological challenges, some examples are the
                    absolutely basic prevention of data loss, control of single file changes and
                    global corpus versions. For corpora, which contain manually compiled
                    transcription and translation layers, annotations, and complex metadata,
                    considerations regarding the quality and consistency of the inter-dependent
                    information play a crucial role in creating scientifically relevant resources. </p>
                <p> This contribution<note xml:id="ftn0" place="foot" n="1"> Parts of this
                        contribution have been created in the project INEL
                        (https://inel.corpora.uni-hamburg.de) within the context of the joint
                        research funding of the German Federal Government and Federal States in the
                        Academies’ Programme, with funding from the Federal Ministry of Education
                        and Research and the Free and Hanseatic City of Hamburg. The Academies’
                        Programme is coordinated by the Union of the German Academies. </note> deals
                    with the implementation and optimization of workflows for the creation of
                    complex corpora for the documentation of indigenous languages, and particularly
                    discusses benefits of automating quality assurance mechanisms. The approach is
                    extended by a quantification method which can be used to measure its impact on
                    the overall quality of the corpora. </p>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head> Corpus creation </head>
                <p>Linguistic corpora are usually created collaboratively. Consistent version
                    control mechanisms facilitate the collaboration in corpus creation processes and
                    are a basic necessity considering the amount of time and work spent on building
                    scientific resources. Minimizing the risk of losing results and maximizing
                    re-usability and reproducibility have to be priorities. </p>
                <p>In the project presented here numerous linguistic corpora are created for
                    documenting indigenous languages spoken on the territory of the Russian
                    Federation. The corpora contain transcriptions of spoken and written language,
                    translations into Russian, English, and German, morphologic, syntactic and
                    semantic information, as well as complex speaker and corpus metadata, all of
                    which are mainly derived in a costly manual process.<note xml:id="ftn1"
                        place="foot" n="2"> The corpora were created in the project INEL
                        (https://inel.corpora.uni-hamburg.de). For a sample of the design and
                        content of a particular corpus see Arkhipov et al. (2019) and
                        http://hdl.handle.net/11022/0000-0007-DA6E-9 </note>
                </p>
                <p>The allocation and completion of particular tasks in the corpus creation process
                    is organized with an established project management tool connected to the
                    applied version control software. The consistent and reproducible data creation
                    is supported by measures for optimizing existing workflows and enhancing the
                    corpus quality (cf. Hedeland / Ferger 2020).</p>
                <p>The notion of quality for linguistic corpora used in this context includes general FAIR metrics, as well as consistency and structure in metadata, transcription conventions and annotation schemes, allowing for reliable automatic processing (cf. Hedeland 2020).</p>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head> Quality assurance and workflow optimization </head>
                <p>Assuring the quality of produced corpus data as well as optimizing the underlying
                    workflows lead to faster production and higher quality of resources. Work on
                    quality assurance of research data is for instance also conducted in the
                    projects Conquaire (Hermann et al. 2021) and QUEST (Arkhangelskiy et al. 2020),
                    the latter building upon the framework discussed here. The framework for quality
                    assurance was developed using an open-source, extensible software system<note
                        xml:id="ftn2" place="foot" n="3"> The Java-based software ‘Corpus Services’
                        is available under MIT license at https://doi.org/10.5281/zenodo.4725655 and
                        is open for collaborative development, extensions, and adaptations by
                        further projects. </note> (Ferger et al. 2020) and integrating it into
                    existing versioning workflows. The functionality of the software consists of
                    methods assisting with manual tasks as well as applying automatic changes and
                    corrections to the data, for enhancing the quality of the created resources. </p>
                <p>The necessary manual improvement of the data quality (which is unavoidable in
                    manually created data) is supported by the automatic identification of problems
                    and the generation of reports and correction lists, which for instance can be
                    used for navigating to the problematic phenomena directly in the transcription
                    software. The automatic check mechanisms are applied periodically and
                    information about absolute numbers of found errors is documented in overviews
                    reflecting the current status of the corpora (see Figure 1). At the moment,
                    rule-based heuristics are applied (e.g. checking against controlled metadata
                    vocabularies, transcription conventions, annotation schemes, or custom
                        patterns)<note xml:id="ftn3" place="foot" n="4"> An example for metadata
                        checking is the function ComaKmlForLocations which compares names and
                        coordinates of settlements added to a map with the corresponding instances
                        in the metadata. For a full overview of functions see the "List of corpus
                        functions" in Ferger et al. 2020. </note>. </p>
                
                    <figure>
                        <graphic url="Pictures/9f39caa48a329abac5a45bb7b811af7b.png"/>
                     </figure>
                <p>Corpus curation status showing error numbers in different creation phases</p>
                <p>In addition to reporting, certain recurring systematic errors are identified and
                    fixed automatically. This for instance includes the adherence to basic
                    transcription conventions (e.g. use of whitespaces and punctuation) and file
                    naming conventions. </p>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head> Evaluation of enhancements </head>
                <p>The consistent version control, including information about automatically
                    executed corrections, enables the tracking of error and correction rates for
                    different versions of the corpora over time. The possibility to re-instantiate
                    all editing states of the corpora also allows for the extraction of this
                    information for corpus versions for which the number of applied fixes was not
                    tracked yet.<note xml:id="ftn4" place="foot" n="5"> As the quality assurance
                        mechanisms were not present in their entirety in the early stages of the
                        project, the possibility to derive implicitly present information about
                        error and correction rates is important for the re-evaluation of older
                        editing stages of the corpora. This is possible by individually checking out
                        the respective corpus versions and creating current-state reports including
                        all necessary information.</note> While it is possible to re-generate the
                    needed numbers retroactively, it is obviously easier to collect them directly
                    during the fixing operations. </p>
                <p>Thus the quality assurance measures were complemented by saving counts of
                    automatic fixes done to the corpus data in a structured format<note
                        xml:id="ftn5" place="foot" n="6"> A JSON format was used that can be
                        directly imported into an ElasticSearch instance. The automatic creation of
                        those files after applying fixes was added as an extension to ‘Corpus
                        Services’ and is adaptable by other projects in Ferger et al. 2020. </note>
                    that can directly be used for further analysis. The collected data contains the
                    types of fixes that caused changes to the data, the date of their application
                    and the number of the corrections performed. For the exploration of the
                    quantitative and temporal information about the quality assurance and further
                    analysis an interactive dashboard was implemented that allows for creating
                    different visualizations. Overviews of checks and fixes can be displayed
                    chronologically, or sorted by type. This way it is possible to see if certain
                    types of errors are typical for certain corpus creation phases.</p>
                
                    <figure>
                        <graphic url="Pictures/b7f6dddc3d709e54f75d546689e9aaca.png"/>                     
                    </figure>
                <p>Dashboard for quality assurance evaluation</p>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head> Summary and outlook </head>
                <p>This contribution shows an approach to evaluate specific measures for quality
                    assurance in the creation of complex linguistic corpora for the documentation of
                    indigenous languages. Besides giving a basic idea about the applied methods, a
                    quantification mechanism is presented that complements the positive intuition of
                    the significance of the efforts taken to optimize the overall quality of the
                    created corpora. By tracking information about error and automatic correction
                    rates and importing it into a search engine, it is possible to get clearer
                    insights into the development of the corpus quality and the actual effect of the
                    quality assurance methods. </p>
                <p>After demonstrating the concept of building a quantification mechanism for enhancements in quality control measures, the approach should now be widened to take into account not only directly accessible information from existing error reports, but also previous corpus versions or even other corpora which were not yet covered by systematic reporting.</p>
                <p>Further operationalization should take place in future work on this topic. This can be as simple as correlating the size of a corpus to the error and correction rates, or more complex like taking into account temporal, infrastructural, or personal factors for the creation of corpora.</p>
            </div>
        </body>
        <back>
            <div type="bibliogr">
                <listBibl>
                    <head>Bibliography</head>
                    <bibl>
                        <hi rend="bold">Arkhangelskiy, Timofey</hi> / <hi rend="bold">Hedeland,
                            Hanna</hi> / <hi rend="bold">Riaposov, Aleksandr</hi> (2020):
                        "Evaluating and Assuring Research Data Quality for Audiovisual Annotated
                        Language Data", in: CLARIN (ed.): <hi rend="italic">Proceedings of CLARIN
                            Annual Conference 2020</hi> . Online Edition. October 2020, 131–135. </bibl>
                    <bibl>
                        <hi rend="bold">Arkhipov, Alexandre</hi> / <hi rend="bold">Däbritz, Chris
                            L.</hi> / <hi rend="bold">Gusev, Valentin</hi> (2019): <hi rend="italic"
                            >INEL Kamas corpus</hi>. User documentation &lt;<ref
                            target="https://corpora.uni-hamburg.de/repository/file:kamas-1.0_INEL_Kamas_Corpus_1.0_User_Documentation/PDF/INEL_Kamas_Corpus.pdf"
                            >https://corpora.uni-hamburg.de/repository/file:kamas-1.0_INEL_Kamas_Corpus_1.0_User_Documentation/PDF/INEL_Kamas_Corpus.pdf</ref>&gt;
                        [09.06.2021]. </bibl>
                    <bibl>
                        <hi rend="bold">Ferger, Anne</hi> / <hi rend="bold">Hedeland, Hanna</hi> /
                            <hi rend="bold">Jettka, Daniel</hi> / <hi rend="bold">Pirinen,
                            Tommi</hi> (2020): <hi rend="italic">Corpus Services</hi>. Zenodo DOI:
                        10.5281/zenodo.4725655. </bibl>
                    <bibl>
                        <hi rend="bold">Hedeland, Hanna</hi> (2020): "Towards Comprehensive
                        Definitions of Data Quality for Audiovisual Annotated Language Resources",
                        in: CLARIN (ed.): <hi rend="italic">Proceedings of CLARIN Annual Conference
                            2020</hi>. Online Edition. October 2020, 136–140. </bibl>
                    <bibl>
                        <hi rend="bold">Hedeland, Hanna</hi> / <hi rend="bold">Ferger, Anne</hi>
                        (2020): "Towards Continuous Quality Control for Spoken Language Corpora",
                        in: <hi rend="italic">International Journal of Digital Curation</hi> 15, 1:
                        1-13 DOI: 10.2218/ijdc.v15i1.601. </bibl>
                    <bibl>
                        <hi rend="bold">Hermann, Fabian</hi> / <hi rend="bold">Pietsch,
                            Christian</hi> / <hi rend="bold">Cimiano, Philipp</hi> (2021):
                        "Conquaire Infrastructure for Continuous Quality Control", in: Cimiano,
                        Philipp / Pietsch, Christian / Wiljes, Cord (eds.): <hi rend="italic"
                            >Studies in Analytical Reproducibility: the Conquaire Project</hi>.
                        Bielefeld 17-27 DOI: 10.4119/unibi/2942780. </bibl>
                </listBibl>
            </div>
        </back>
    </text>
</TEI>
