<?xml version="1.0" encoding="UTF-8"?>
<TEI xmlns="http://www.tei-c.org/ns/1.0">
    <teiHeader>
        <fileDesc>
            <titleStmt>
                <title>The Corpus for Idiolectal Research (CIDRE)</title>
                <author>
                    <persName>
                        <surname>Seminck</surname>
                        <forename>Olga</forename>
                    </persName>
                    <affiliation>Centre national de la recherche scientifique; Lattice (Langues, Textes, Traitements informatiques, Cognition) - UMR 8094</affiliation>
                    <email>olga.seminck@cri-paris.org</email>
                </author>
                <author>
                    <persName>
                        <surname>Gambette</surname>
                        <forename>Philippe</forename>
                    </persName>
                    <affiliation>Université Gustave Eiffel; Lattice (Langues, Textes, Traitements informatiques, Cognition) - UMR 8094</affiliation>
                    <email>philippe.gambette@univ-eiffel.fr</email>
                </author>
                <author>
                    <persName>
                        <surname>Legallois</surname>
                        <forename>Dominique</forename>
                    </persName>
                    <affiliation>Université Sorbonne nouvelle; Lattice (Langues, Textes, Traitements informatiques, Cognition) - UMR 8094</affiliation>
                    <email>dominique.legallois@sorbonne-nouvelle.fr</email>
                </author>
                <author>
                    <persName>
                        <surname>Poibeau</surname>
                        <forename>Thierry</forename>
                    </persName>
                    <affiliation>Centre national de la recherche scientifique; Lattice (Langues, Textes, Traitements informatiques, Cognition) - UMR 8094</affiliation>
                    <email>thierry.poibeau@ens.psl.eu</email>
                </author>
            </titleStmt>
            <editionStmt>
                <edition>
                    <date>2021-06-15T13:47:00Z</date>
                </edition>
            </editionStmt>
            <publicationStmt>
                <publisher>Elisabeth Burr, University of Leipzig</publisher>
                <address>
                    <addrLine>Beethovenstr. 15</addrLine>
                    <addrLine>04107 Leipzig</addrLine>
                    <addrLine>Germany</addrLine>
                    <addrLine>Elisabeth Burr</addrLine>
                </address>
            </publicationStmt>
            <sourceDesc>
                <p>Converted from a Word document</p>
            </sourceDesc>
        </fileDesc>
        <encodingDesc>
            <appInfo>
                <application ident="DHCONVALIDATOR" version="1.22">
                    <label>DHConvalidator</label>
                </application>
            </appInfo>
        </encodingDesc>
        <profileDesc>
            <textClass>
                <keywords scheme="ConfTool" n="category">
                    <term>Paper</term>
                </keywords>
                <keywords scheme="ConfTool" n="subcategory">
                    <term>Poster Presentation</term>
                </keywords>
                <keywords scheme="ConfTool" n="keywords">
                    <term>Idiolect</term>
                    <term>Corpus Linguistics</term>
                    <term>French</term>
                    <term>Literature</term>
                    <term>Stylometry</term>
                </keywords>
                <keywords scheme="ConfTool" n="topics">
                    <term>Gathering</term>
                    <term>Programming</term>
                    <term>Annotating</term>
                    <term>Cleanup</term>
                    <term>Structural Analysis</term>
                    <term>Stylistic Analysis</term>
                    <term>Modeling</term>
                    <term>Theorizing</term>
                    <term>Archiving</term>
                    <term>Organizing</term>
                    <term>Publishing</term>
                    <term>Sharing</term>
                    <term>Meta: Assessing</term>
                    <term>DigitalHumanities</term>
                    <term>Language</term>
                    <term>Metadata</term>
                    <term>Literature</term>
                    <term>Research</term>
                    <term>Data</term>
                    <term>English</term>
                </keywords>
            </textClass>
        </profileDesc>
    </teiHeader>
    <text>
        <body>
            <div type="div1" rend="DH-Heading1">
                <head>Abstract</head>
                <p>It is well known that the idiolect (the language of an individual) evolves over time. However, there is a lack of quantitative studies on this topic, due to the lack of large corpora (but see Barlow 2013; Mollin 2009; Petré et al. 2019 for a few examples). To study what is specific in an idiolect and how it evolves over a lifetime, we assembled, cleaned and dated the fiction works of 11 very prolific 19th and early 20th century French writers. This resulted in the CIDRE corpus counting 37 million words and over 400 books.</p>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head>Motivation</head>
                <p>We want to assemble a longitudinal corpus for stylistics studies, that is:</p>
                <list type="unordered">
                    <item>Homogeneous (only fiction work)</item>
                    <item>Open (contain only works in the public domain)</item>
                    <item>Dense covering the 19th and early 20th century with overlaps between the different authors</item>
                </list>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head>Corpus Assembling</head>
                <p>Criteria to select relevant authors:</p>
                <list type="unordered">
                    <item>Large oeuvre of fiction novels</item>
                    <item>E-books available as high quality e-pub files</item>
                    <item>No collaborative works (e.g. A. Dumas)</item>
                </list>
                <p>Programming Scripts:</p>
                <list type="unordered">
                    <item>Automatic download of e-pub files from open source library projects (e.g. Wikisource, etc.)</item>
                    <item>Automatic cleaning of e-pub files (remove forewords, image captions, etc.)</item>
                    <item>Manual correction</item>
                </list>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head>Data</head>
                <figure>
                    <graphic n="1001" width="16.002cm" height="9.883069444444445cm" url="Pictures/c3e1730f8097b4d88891b56964de3191.jpg" rend="inline"/>
                </figure>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head>Dating of Works</head>
                <p>Important contribution of ours : Annotation of Books with year of writing</p>
                <list type="unordered">
                    <item>Crucial for a diachronical study of the idiolect</item>
                    <item>Might differ largely from:
                        <list type="unordered">
                            <item>First year of publication</item>
                            <item>Year of printing </item>
                        </list>
                    </item>
                    <item>Provided in a metadata file with the corpus</item>
                </list>
                <figure>
                    <graphic n="1002" width="16.002cm" height="6.552847222222222cm" url="Pictures/41b18953eda74cbd206068cc8a1b3a99.png" rend="inline"/>
                </figure>
                <p>Table 1: Some examples from the Gréville Corpus</p>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head>Dates of Works</head>
                <figure>
                    <graphic n="1003" width="16.002cm" height="9.883069444444445cm" url="Pictures/1cece7e711889a9537168b1b85cc4d1b.jpg" rend="inline"/>
                </figure>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head>Availability and Licenses</head>
                <list type="unordered">
                    <item>Download: 
                        <ref target="https://github.com/oseminck/cidre/tree/v2.0">https://github.com/oseminck/cidre/tree/v2.0</ref>
                    </item>
                    <item>Texts: public domain</item>
                    <item>Metadata: Creative Commons - Attribution-ShareAlike 4.0</item>
                    <item>Processing scripts: GPLv3 License</item>
                </list>
                <figure>
                    <graphic n="1004" width="3.92cm" height="4.73cm" url="Pictures/256791f507d2ecd6ffc3f41f36623dda.png" rend="inline"/>
                </figure>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head>Acknowledgements</head>
                <p style="text-align: left;">This work was funded in part by the French government under the management of the Agence Nationale de la Recherche as part of the "Investissements d’avenir" program, reference ANR-19-P3IA-0001 (PRAIRIE 3IA Institute).</p>
            </div>
        </body>
        <back>
            <div type="bibliogr">
                <listBibl>
                    <head>Bibliography</head>
                    <bibl>
                        <hi rend="bold">Barlow, Michael</hi> (2013): "Individual differences and
                        usage-based grammar", in: <hi rend="italic">International Journal of Corpus
                            Linguistics</hi> 18, 4: 443–478. </bibl>
                    <bibl>
                        <hi rend="bold">Mollin, Sandra</hi> (2009): "'I entirely understand' is a
                        Blairism: The methodology of identifying idiolectal collocations", in: <hi
                            rend="italic">International Journal of Corpus Linguistics</hi> 14, 3:
                        367–392. </bibl>
                    <bibl>
                        <hi rend="bold">Petré, Peter</hi> /<hi rend="bold"> Anthonissen, Lynn</hi> /
                            <hi rend="bold">Budts, Sara</hi> / <hi rend="bold">Manjavacas,
                            Enrique</hi> / <hi rend="bold">Silva, Emma-Louise</hi> / <hi rend="bold"
                            >Standing, William</hi> / <hi rend="bold">Strik, Odile A.O.</hi> (2019):
                        "Early Modern Multiloquent Authors (EMMA): Designing a large-scale corpus of
                        individuals’ languages", in: <hi rend="italic">ICAME Journal</hi> 43, 1:
                        83–122. </bibl>
                </listBibl>
            </div>
        </back>
    </text>
</TEI>
