<?xml version="1.0" encoding="UTF-8"?>
<TEI xmlns="http://www.tei-c.org/ns/1.0">
    <teiHeader>
        <fileDesc>
            <titleStmt>
                <title type="full">
                    <title type="main">The Corpus of German Song Lyrics: Recent Developments and Interdisciplinary Potential</title>
                    <title type="sub"/>
                </title>
                <author>
                    <persName>
                        <surname>Schneider</surname>
                        <forename>Roman</forename>
                    </persName>
                    <affiliation>Leibniz-Institut für Deutsche Sprache (IDS), Deutschland</affiliation>
                    <email>schneider@ids-mannheim.de</email>
                </author>
            </titleStmt>
            <editionStmt>
                <edition>
                    <date>2021-06-06T12:24:24.470690997</date>
                </edition>
            </editionStmt>
            <publicationStmt>
                <publisher>Elisabeth Burr, University of Leipzig</publisher>
                <address>
                    <addrLine>Beethovenstr. 15</addrLine>
                    <addrLine>04107 Leipzig</addrLine>
                    <addrLine>Germany</addrLine>
                    <addrLine>Elisabeth Burr</addrLine>
                </address>
            </publicationStmt>
            <sourceDesc>
                <p>Converted from an OASIS Open Document</p>
            </sourceDesc>
        </fileDesc>
        <encodingDesc>
            <appInfo>
                <application ident="DHCONVALIDATOR" version="1.22">
                    <label>DHConvalidator</label>
                </application>
            </appInfo>
        </encodingDesc>
        <profileDesc>
            <textClass>
                <keywords scheme="ConfTool" n="category">
                    <term>Paper</term>
                </keywords>
                <keywords scheme="ConfTool" n="subcategory">
                    <term>Poster Presentation</term>
                </keywords>
                <keywords scheme="ConfTool" n="keywords">
                    <term>Song Lyrics</term>
                    <term>Empirical Studies</term>
                </keywords>
                <keywords scheme="ConfTool" n="topics">
                    <term>Gathering</term>
                    <term>Programming</term>
                    <term>Annotating</term>
                    <term>Content Analysis</term>
                    <term>Stylistic Analysis</term>
                    <term>Visualization</term>
                    <term>Contextualizing</term>
                    <term>Archiving</term>
                    <term>Publishing</term>
                    <term>Sharing</term>
                    <term>Meta: Teaching/Learning</term>
                    <term>DigitalHumanities</term>
                    <term>Text</term>
                    <term>Language</term>
                    <term>Research</term>
                    <term>Data</term>
                    <term>English</term>
                </keywords>
            </textClass>
        </profileDesc>
    </teiHeader>
    <text>
        <body>
            <p>Popular music and its lyrics have established themselves
                as integral parts of modern everyday culture. Given the varied situations in which
                we are surrounded by them - and the mass media presence of its prominent actors -
                lyrics certainly have an impact on the use of language that should not be
                underestimated. For example, the hip-hop genre with its largely melody-free rap
                chants has evolved from an originally subcultural phenomenon to a global form of
                articulation for (mostly) younger people (Androutsopoulos 2003), and therefore has
                become a legitimate subject of contemporary digital humanities studies, linguistics,
                and even language education (Werner / Tegge 2020). For demonstrable reasons related
                to the limited quantitative data available, extensive empirical investigations that
                rely on large textual databases have so far focused primarily on Anglo-American
                characteristics.</p>
            <p>The multilayer-annotated corpus of German lyrics (Schneider 2020) contributes to
                closing this data gap in the elusive continuum between standard and non-standard
                varieties as well as between written and spoken language. It allows multivariate,
                statistically based, and reproducible analyses of lexical, pragmatic or
                morphological phenomena, and comprises both thematic and author-specific archives.
                These include collected works of artists, covering a wide range from hip-hop to
                rock-pop to political singer-songwriters (Schneider et al. 2021), as well as the
                most successful German-language chart songs over the past five decades. In
                compliance with legal requirements, selected XML-TEI-annotated archives are freely
                available for academic use. They offer machine-driven annotations of constituent
                structures and manual annotations of lemmata, part-of-speech tags, named entities,
                neologisms, and rhyme types. A dedicated website ( <hi rend="italic"
                    >www.songkorpus.de</hi>) features a corpus front end with a wide range of search
                functions, exploration functions on various description levels, and live
                visualizations; cf. Figure 1. </p>
                <figure>
                    <graphic url="Pictures/d7804764ccade017eb433a61299133cd.jpg"/>
                </figure>
            <p>Figure 1: The Songkorpus Website</p>
            
            <p>We introduce substantial new corpus features, e.g. additional archives - increasing
                the number of corpus tokens to more than three million - and the calculation of word
                embeddings, using the GloVe unsupervised learning algorithm (Pennington et al.
                2014). A downloadable comprehensive collocation dataset now contains all bi-, tri-,
                tetra-, penta- and hexagrams of the corpus lyrics, with nine measures of association
                strength each, including widespread parameters like pointwise mutual information or
                log-likelihood as well as innovative approaches like lexical gravity or cost
                reduction (cf. Gries 2015). This is intended as a starting point for some methodical
                evaluation of the reliability and meaningfulness of statistical measures with
                respect to rather rare collocations, that are nevertheless worth investigating
                because of their unusal or creative usage. Moreover, a measure of context
                similarity, based on five-word context windows and multi-dimensional vector space
                models, allows for the examination of multiword expressions (MWE) that do not quite
                fit into their contexts. Typical examples are idioms or metaphors; both seem
                prominent in song lyrics (cf. Werner 2012). This new measure comes in two variants:
                one variant (CO_VEC_LEX) only takes content words (nouns, main verbs, adjectives and
                adverbs) into account, the second variant (CO_VEC) also counts in function words.
                All measures can be fruitfully applied for empirical tasks, e.g. the data-driven
                identification of idiomatic multiword-expressions (cf. Amin et al. 2021).</p>
            <p>We expect the corpus to provide a wide range of additional starting points for future transdisciplinary projects. Besides (computational) linguistics and literary studies, benefiting research areas can be located in the broad spectrum of cultural studies, language didactics, or media sciences.</p>
        </body>
        <back>
            <div type="bibliogr">
                <listBibl>
                    <head>Bibliographie</head>
                    <bibl>
                        <hi rend="bold">Amin, Miriam</hi> / <hi rend="bold">Fankhauser, Peter</hi> /
                            <hi rend="bold">Kupietz, Marc</hi> / <hi rend="bold">Schneider,
                            Roman</hi> (2021): "Data-driven Identification of Idioms in Song
                        Lyrics", in: Association for Computational Linguistics (ed.): <hi
                            rend="italic">Proceedings of the 17th Workshop on Multiword Expressions
                            (MWE 2021)</hi>. Bangkok, Thailand (online), August 6, 2021 13-22
                        &lt;<ref target="https://aclanthology.org/2021.mwe-1.3.pdf">https://aclanthology.org/2021.mwe-1.3.pdf</ref>&gt; [21.08.2021]. </bibl>
                    <bibl>
                        <hi rend="bold">Androutsopoulos, Jannis</hi> (ed.) (2003): <hi rend="italic"
                            >HipHop: Globale Kultur – lokale Praktiken</hi>. Bielefeld: transcript. </bibl>
                    <bibl>
                        <hi rend="bold">Gries, Stefan Th.</hi> (2015): "Quantitative designs and
                        statistical techniques", in: Biber, Douglas / Reppen, Randi (eds.): <hi
                            rend="italic">The Cambridge Handbook of English Corpus Linguistics</hi>.
                        Cambridge: University Press 50-71.</bibl>
                    <bibl>
                        <hi rend="bold">Pennington, Jeffrey</hi> / <hi rend="bold">Socher,
                            Richard</hi> / <hi rend="bold">Manning, Christopher D.</hi> (2014):
                        "GloVe: Global Vectors for Word Representation", in: Association for
                        Computational Linguistics (ed.): <hi rend="italic">Proceedings of the 2014
                            Conference on Empirical Methods in Natural Language Processing
                            (EMNLP)</hi>, Doha, Qatar, October 2014 1532-1543
                        &lt;<ref target="https://aclanthology.org/D14-1162.pdf">https://aclanthology.org/D14-1162.pdf</ref>&gt; [21.08.2021]. </bibl>
                    <bibl>
                        <hi rend="bold">Schneider, Roman</hi> (2020): "A Corpus Linguistic
                        Perspective on Contemporary German Pop Lyrics with the Multi-Layer Annotated
                        Songkorpus", in: ELRA (ed.): <hi rend="italic">Proceedings of The 12th
                            Language Resources and Evaluation Conference (LREC)</hi>. Marseille,
                        France, May 2020 835-841
                        &lt;<ref target="https://aclanthology.org/2020.lrec-1.105.pdf">https://aclanthology.org/2020.lrec-1.105.pdf</ref>&gt; [21.08.2021]. </bibl>
                    <bibl>
                        <hi rend="bold">Schneider, Roman</hi> / <hi rend="bold">Hansen, Sandra</hi>
                        / <hi rend="bold">Lang, Christian</hi> (2021): "Das Vokabular von Songtexten
                        im gesellschaftlichen Kontext – ein diachron-empirischer Beitrag", in: <hi
                            rend="italic">Sprache in Politik und Gesellschaft: Perspektiven und
                            Zugänge.</hi> Jahrbuch 2021 des Instituts für Deutsche Sprache. Berlin /
                        Boston: De Gruyter 295-304.</bibl>
                    <bibl>
                        <hi rend="bold">Werner, Valentin</hi> (2012): "Love is all around: A
                        corpus-based study of pop lyrics", in: <hi rend="italic">Corpora</hi> 7, 1:
                        19-50.</bibl>
                    <bibl><hi rend="bold">Werner, Valentin</hi> / <hi rend="bold">Tegge,
                            Friederike</hi> (eds.) (2020): <hi rend="italic">Pop Culture in Language
                            Education</hi>. Theory, Research, Practice. London: Routledge. </bibl>
                </listBibl>
            </div>
        </back>
    </text>
</TEI>
