<?xml version="1.0" encoding="UTF-8"?>
<TEI xmlns="http://www.tei-c.org/ns/1.0">
    <teiHeader>
        <fileDesc>
            <titleStmt>
                <title>Character Networks in a Collection of 19th Century German Novellas</title>
                <author>
                    <persName>
                        <surname>Päpcke</surname>
                        <forename>Simon</forename>
                    </persName>
                    <affiliation>ETH Zürich, Schweiz</affiliation>
                    <email>simon.paepcke@gess.ethz.ch</email>
                </author>
                <author>
                    <persName>
                        <surname>Brandes</surname>
                        <forename>Ulrik</forename>
                    </persName>
                    <affiliation>ETH Zürich, Schweiz</affiliation>
                    <email>ulrik.brandes@gess.ethz.ch</email>
                </author>
            </titleStmt>
            <editionStmt>
                <edition>
                    <date>2021-09-17T13:46:00Z</date>
                </edition>
            </editionStmt>
            <publicationStmt>
                <publisher>Elisabeth Burr, University of Leipzig</publisher>
                <address>
                    <addrLine>Beethovenstr. 15</addrLine>
                    <addrLine>04107 Leipzig</addrLine>
                    <addrLine>Germany</addrLine>
                    <addrLine>Elisabeth Burr</addrLine>
                </address>
            </publicationStmt>
            <sourceDesc>
                <p>Converted from a Word document</p>
            </sourceDesc>
        </fileDesc>
        <encodingDesc>
            <appInfo>
                <application ident="DHCONVALIDATOR" version="1.22">
                    <label>DHConvalidator</label>
                </application>
            </appInfo>
        </encodingDesc>
        <profileDesc>
            <textClass>
                <keywords scheme="ConfTool" n="category">
                    <term>Paper</term>
                </keywords>
                <keywords scheme="ConfTool" n="subcategory">
                    <term>Long paper</term>
                </keywords>
                <keywords scheme="ConfTool" n="keywords">
                    <term>character networks</term>
                    <term>spectral graph distance</term>
                    <term>network ensemble</term>
                </keywords>
                <keywords scheme="ConfTool" n="topics">
                    <term>Designing</term>
                    <term>Programming</term>
                    <term>Network Analysis</term>
                    <term>Stylistic Analysis</term>
                    <term>Modeling</term>
                    <term>DigitalHumanities</term>
                    <term>NamedEntities</term>
                    <term>Methods</term>
                    <term>Literature</term>
                    <term>not applicable</term>
                    <term>not applicable</term>
                    <term>not applicable</term>
                    <term>not applicable</term>
                    <term>not applicable</term>
                    <term>English</term>
                </keywords>
            </textClass>
        </profileDesc>
    </teiHeader>
    <text>
        <body>
            <div type="div1" rend="DH-Heading1">
                <head>Introduction</head>
                <p>In recent years, network analysis has become a frequently applied method in the digital humanities. As a common practice, literary scholars employ tools from network science, a discipline originating from mathematical graph theory, to investigate character networks in literature (Moretti 2011; Trilcke 2013; Piper / Algee-Hewitt 2014; Erlin / Tatlock 2014). Approaches use among others centrality measures, network motifs or community detection and apply these to static or dynamic character networks (Kydros / Anastasiadis 2015; Fischer et al. 2017). With this, one either verifies known results from literary studies or gains a deeper understanding of the characters roles.</p>
                <p>A large part of these analyses deals with the network in a singular work (Rochat
                    2014) or looks at individual outcomes for the networks of a larger corpus
                    (Jannidis et al. 2016; Isasi 2017). In the present work, we want to contribute
                    to this research by analyzing and comparing a corpus as an entire network
                    ensemble. Moreover, we compare different prose fiction corpora, namely novellas
                    to novels.</p>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head>Motivation</head>
                <p>We consider a 19th century corpus of novellas and analyze whether their character constellation networks have common structural properties. The text collection was composed with the editors’ intention to be a paradigmatic sample of the novella style and a strict formal criterion was given to distinguish novellas from novels. Therefore, the editors established the phrase strong silhouette ("starke Silhouette") and claimed it to be their guiding principle in the selection of texts. As such, a text does not demand to have a certain text length to be rated as a novella but instead needs to stay focused on a single topic that then can be executed thoroughly. Hence, we hypothesize that the novellas in the corpus have a similar character constellation network, which further motivates our research question whether the restricted form of a novella gives preference to a specific character constellation.</p>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head>Corpus</head>
                <p>We illustrate our methods on the
                    <hi rend="italic"> Deutsche Novellenschatz</hi>, a corpus
                    of 19th century German language novellas published by Paul Heyse and Hermann
                    Kurz between 1871 and 1876. <note place="foot" xml:id="ftn1" n="1">
                        <p rend="footnote text"> available via
                            http://www.deutschestextarchiv.de/doku/textquellen\#novellenschatz
                            (Weitin 2016).</p>
                    </note> It contains 86 novellas of 82 different authors (11 female, 71 male)
                    that are of various length. The texts that have been originally published
                    between 1811 and 1875 cover the epochs of German romanticism, Biedermeier, Young
                    Germany, Vormärz and literary realism. Main topics are love, wedding and
                    marriage as well as village life, art and justice. Moreover, the corpus itself
                    contains a long introduction in which the editors state their intent to build a
                    canonical collection of 19th century German language novellas. </p>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head>Character networks </head>
                <p>To test our hypothesis, we use methods from natural language processing and network analysis to generate and compare the required character networks. We will consider two different network building approaches. </p>
                <p>First, we construct undirected co-occurrence networks where the nodes represent main characters that are derived from the texts by a manually improved named entity recognition. More concretely, we take into account all entities that are marked as persons by either 
                    <hi rend="italic">spacy</hi>'s German model parser
                    <note place="foot" xml:id="ftn2" n="2">
                        <p rend="footnote text"> spacy.io</p>
                    </note> or Akers part-of-speech tagger
                    <note place="foot" xml:id="ftn3" n="3">
                        <p rend="footnote text"> available via http://staffwww.dcs.shef.ac.uk/people/A.Aker/activityNLPProjects.html</p>
                    </note> or are present in a list of German noble titles and remove all names that only occur once in the text. We then remove false positives manually and do a simple matching for same names with wrong lemmatization (e.g. 
                    <hi rend="bold">Rosalie</hi> vs. 
                    <hi rend="bold">Rosalien</hi>). This method is used due to the lack of a larger training set to identify characters in prose literature as done by Jannidis et al. (2015). In the given networks, a link between two nodes is present if they co-occur in the same sentence. Furthermore, there are two link weights attached to each link reflecting its strength and overall sentiment between the characters involved. Therefore, we use the sentiment analysis implemented in Textblob
                    <note place="foot" xml:id="ftn4" n="4">
                        <p rend="footnote text"> https://textblob.readthedocs.io/en/dev/index.html</p>
                    </note> on the sentences where characters co-occur. 
                </p>
                <p>Second, we also derive character networks by syntactic structures and make use of case grammar networks as proposed by Franzosi (2004: 29-108). While the nodes are again the main characters, the links are now directed from subjects to objects that are connected via an action (verb). These syntactic relations are deduced from the dependency parser implemented in 
                    <hi rend="italic">spacy</hi>.
                </p>
                <p>
                    <hi rend="italic">Example:</hi>
                </p>
                <quote>“Im Garten saß (
                    <hi rend="bold">action</hi>) nun Basset (
                    <hi rend="bold">subject</hi>) dem Francoeur (
                    <hi rend="bold">object</hi>) gegenüber und sah ihn stillschweigend an [...]“ (
                    <hi rend="bold">sentiment = -0.7</hi>)
                </quote>
                <quote style="text-align: right;">Arnim, 
                    <hi rend="italic">Der tolle Invalide auf dem Fort Ratonneau</hi>
                </quote>
                <p>In Figure 1 we see an example of such character networks. By looking at the network statistics (Figure 2), we can observe differences of the two network types.</p>
                <p>While the number of characters in each of the networks does not differ, the average degree decreases for the directed networks even though it is the sum of in- and out-degree. The graph centralization w.r.t. degree as defined by Freeman (1978) indicates how ‘star’-like a network is. While some case grammar networks do have a higher centralization (in the example of Schücking 0.43 for the co-occurrence network vs. 0.91 for the case grammar network), others become too sparse, resulting in a less centralized structure. </p>
                <figure>
                    <graphic n="1001" width="14.208125cm" height="7.739944444444444cm" url="Pictures/3fd3eb625036c01f8d3e0f75e4664e01.jpeg" rend="block"/>
                </figure>
                <p>Figure 1: Undirected co-occurrence network and directed case grammar network for Schücking's novella 
                    <hi rend="italic">Die Schwester</hi>. Color represents sentiment and link width the number of interactions
                </p>
                <figure>
                    <graphic n="1002" width="16.002cm" height="2.7252083333333332cm" url="Pictures/62ccfff75d79771cd90b4faa7fbd79d7.jpeg" rend="block"/>
                </figure>
                <figure>
                    <graphic n="1003" width="16.002cm" height="2.7252083333333332cm" url="Pictures/6662a653700dcec33c8a97fa857f1648.jpeg" rend="block"/>
                </figure>
                <p>Figure 2: Histograms for number of nodes, average degree, centralization and average sentiment for the co-occurrence (top) and case grammar (bottom) networks</p>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head>Network ensemble</head>
                <p>To answer our research question we regard these networks as instances in a metric space. Therefore, we use a spectral graph distance (Nagel 2011: 65-94), a similarity measure that compares networks by considering the spectrum transformation costs. Thus, its invariance under network size and automorphisms make it applicable for our purpose of comparing 86 different networks of diverse size with unrelated characters.</p>
                <p>The resulting similarities are represented in an adjacency matrix of pairwise distances between the 86 networks. For visualization, we use the stress variant of multidimensional scaling (MDS, Borg / Groenen 2005), a technique that generates two-dimensional scatterplots approximately representing the input distances and thus preserving some clustering structure (Figures 3 &amp; 4).</p>
                <figure>
                    <graphic n="1004" width="16.002cm" height="6.38175cm" url="Pictures/2dc64175f7a23594a886f8bdabacedbb.jpeg" rend="block"/>
                </figure>
                <p>Figure 3: MDS scatterplot of all texts from the three corpora (left) and a magnification of the densest area (right). </p>
                <figure>
                    <graphic n="1005" width="16.002cm" height="5.288138888888889cm" url="Pictures/a5b1cd9c6a2644446abda6a608170417.jpeg" rend="inline"/>
                </figure>
                <p>Figure 4: Co-occurrence (σ
                    <hi rend="subscript">1 </hi>= 0.0089) 
                    <note place="foot" xml:id="ftn5" n="5">
                        <p rend="footnote text"> The stress σ
                            <hi rend="subscript">1 </hi>is an indicator of the goodness-of-fit for the MDS. Kruskal (1964) defines a stress of $&lt;0.025$ as excellent
                        </p>
                    </note> and case grammar networks (σ
                    <hi rend="subscript">1 </hi>=0.0119) in the 
                    <hi rend="italic">Novellenschatz. </hi>The area of each circle is proportional to the text length and the color indicates the centralization (in %).
                </p>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head>Findings</head>
                <p>We hypothesized that the novellas will have a common character constellation. In contrast, we observed in Figure 2 that the different novellas do have a variety of present characters, diverse network densities, as well as some differences in centralization. However, most of these statistics are in strong correlation to text length and hence, are less insightful for deeper structural comparisons. Instead, we used a spectral graph distance and indeed observed a short distance for a large share of the corpus elements. As a reference, we compare this corpus together with its subsequent corpus 
                    <hi rend="italic">Neuer deutscher Novellenschatz</hi> to a corpus of 305 texts of general German prose published between 1655 and 1881 with a focus on texts released between 1770 and 1850 (Figure 3). 
                </p>
                <p>Moreover, we want to emphasize the positions of two exemplary novellas. Auerbach's 
                    <hi rend="italic">Die Geschichte des Diethelm zu Buchenberg</hi> is the longest novella in the corpus and is seen as a novel instead of a novella by many literary scholars. This could be a possible explanation for its outlier position at the top of both plots in Figure 4. In addition, for the co-occurrence network Heyse's 
                    <hi rend="italic">Der Weinhüter von Meran</hi> is the novella closest to the centriod, an interesting finding, especially if we recall that Heyse is one of the editors of the 
                    <hi rend="italic">Novellenschatz</hi> and was called the “Virtuose des Durchschnitts" (Jeziorkowski 1987) by other scholars. 
                </p>
                <p>If we compare the two different network types, we first want to point out that the case grammar networks can be viewed as subgraphs of the co-occurrence networks and indeed only 38% of the links remain notwithstanding that there are two possible directions for each link in the co-occurrence network.</p>
                <p>The result that many of the texts do have a short spectral graph distance is even more pronounced in the case grammar networks.</p>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head>Conclusion </head>
                <p>We used a spectral graph distance measure to analyze character constellations in
                    a corpus of novellas, and found that, outliers notwithstanding, high similarity
                    overall. Additionally, the approach can be understood as a guiding principle for
                    the more general comparison of network ensembles in the digital humanities
                    beyond considering the descriptive statistics. We showed that the method yields
                    to interpretable results for directed (case grammar) and undirected
                    (co-occurrence) networks as well as for signed and weighted ones (sentiment).
                    One could imagine to cluster character networks from movies to investigate
                    whether there is a genre specific character constellation, compare networks in
                    diverse archeological settings with each other (e.g. trade networks of different
                    cultures), or analyze the style of correspondence networks for authors.</p>
            </div>
        </body>
        <back>
            <div type="bibliogr">
                <listBibl>
                    <head>Bibliography</head>
                    <bibl>
                        <hi rend="bold">Borg, Ingwer</hi> / <hi rend="bold">Groenen, Patrick J. F.
                        </hi>(<hi rend="superscript">2</hi>2005): <hi rend="italic">Modern
                            Multidimensional Scaling: Theory and Applications</hi> (= Springer
                        series in statistics). New York: Springer. </bibl>
                    <bibl>
                        <hi rend="bold">Erlin, Matt</hi> (2014): “The Location of Literary History:
                        Topic Modeling, Network Analysis, and the German Novel, 1731–1864”, in:
                        Erlin, Matt / Tatlock, Lynne (eds.): <hi rend="italic">Distant Readings:
                            Topologies of German Culture in the Long Nineteenth Century</hi>.
                        Cambridge: Cambridge University Press 55-90. </bibl>
                    <bibl>
                        <hi rend="bold">Fischer, Frank</hi> / <hi rend="bold">Göbel, Mathias</hi> /
                            <hi rend="bold">Kampkaspar Dario</hi> / <hi rend="bold">Kittel,
                            Christopher</hi> / <hi rend="bold">Trilcke, Peer</hi> (2017): “Network
                        dynamics, plot analysis: Approaching the progressive structuration of
                        literary texts”, in: <hi rend="italic">Digital Humanities 2017 (Montréal,
                            8-11 August 2017)</hi>. Book of Abstracts.
                        McGill University &lt;<ref target="https://dh2017.adho.org/abstracts/071/071.pdf">https://dh2017.adho.org/abstracts/071/071.pdf
                        </ref>&gt; [15.06.2021]. </bibl>
                    <bibl>
                        <hi rend="bold">Franzosi, Roberto</hi> (2004): <hi rend="italic">From Words
                            to Numbers: Narrative, Data, and Social Science.</hi> Cambridge:
                        Cambridge University Press. </bibl>
                    <bibl>
                        <hi rend="bold">Freeman, Linton C.</hi> (1978): “Centrality in social
                        networks conceptual clarification”, in: <hi rend="italic">Social
                            Networks</hi> 1, 3: 215–239. </bibl>
                    <bibl>
                        <hi rend="bold">Isasi, Jennifer</hi> (2017): <hi rend="italic">Posibilidades
                            de la mineria de datos digital para el analisis del personaje literario
                            en la novela española: El caso de Galdos y los “Episodios
                            Nacionales”.</hi> Ph.D. thesis, University of Nebraska - Lincoln.
                        &lt;<ref target="http://digitalcommons.unl.edu/dissertations/AAI10682923">http://digitalcommons.unl.edu/dissertations/AAI10682923</ref>&gt;
                        [15.06.2021]. </bibl>
                    <bibl>
                        <hi rend="bold">Jannidis, Fotis</hi> / <hi rend="bold">Krug, Markus</hi>
                            /<hi rend="bold"> Toepfer, Martin</hi> / <hi rend="bold">Puppe,
                            Frank</hi> / <hi rend="bold">Reger, Isabella</hi> / <hi rend="bold"
                            >Weimer, Lukas</hi> (2015): “Automatische Erkennung von Figuren in
                        deutschsprachigen Romanen”, in: <hi rend="italic">Digital Humanities im
                            deutschsprachigen Raum</hi>, Graz. </bibl>
                    <bibl>
                        <hi rend="bold">Jannidis, Fotis</hi> /<hi rend="bold"> Reger, Isabella</hi>
                        / <hi rend="bold">Krug, Markus</hi> / <hi rend="bold">Weimer, Lukas</hi> /
                            <hi rend="bold">Macharowsky, Luisa</hi> / <hi rend="bold">Puppe,
                            Frank</hi> (2016): “Comparison of Methods for the Identification of Main
                        Characters in German Novels”, in: <hi rend="italic">Digital Humanities 2016</hi>.
                            Conference Abstracts. Jagiellonian University &amp; Pedagogical
                            University, Kraków 578–582. &lt;<ref target="http://dh2016.adho.org/abstracts/297">http://dh2016.adho.org/abstracts/297</ref>&gt;
                        [15.06.2021]. </bibl>
                    <bibl>
                        <hi rend="bold">Jannidis, Fotis</hi> (2017): “Perspektiven quantitativer
                        Untersuchungen des Novellenschatzes”, in: <hi rend="italic">Zeitschrift für
                            Literaturwissenschaft und Linguistik</hi> 47, 1: 7-27. </bibl>
                    <bibl>
                        <hi rend="bold">Jeziorkowski, Klaus</hi> (1987): “Der Virtuose des
                        Durchschnitts. Der Salonautor in der deutschen Literatur des 19.
                        Jahrhunderts, dargestellt am Beispiel Paul Heyse”, in: Jeziorkowski, Klaus
                        (ed.): <hi rend="italic">Eine Iphigenie rauchend</hi>. Aufsätze und
                        Feuilletons zur deutschen Tradition <hi rend="italic">.</hi> Frankfurt a.M.:
                        Suhrkamp 114-129. </bibl>
                    <bibl>
                        <hi rend="bold">Kruskal, Joseph</hi> (1964): “Multidimensional scaling by
                        optimizing goodness of fit to a nonmetric hypothesis”, in: <hi rend="italic"
                            >Psychometrika</hi> 29, 1: 1-27. </bibl>
                    <bibl>
                        <hi rend="bold">Kydros, Dimitrios</hi> /<hi rend="bold"> Anastasiadis,
                            Anastasios</hi> (2015): “Social network analysis in literature. The case
                        of The Great Eastern by A. Embirikos”, in: <hi rend="italic">Proceedings of
                            the 5th European Congress of Modern Greek Studies of the European
                            Society of Modern Greek Studies</hi> 4: 681-702. </bibl>
                    <bibl>
                        <hi rend="bold">Moretti, Franco</hi> (2011): “Network theory, plot
                        analysis”, in: <hi rend="italic">New Left Review</hi> 68: 80-102. </bibl>
                    <bibl>
                        <hi rend="bold">Nagel, Uwe</hi> (2011): <hi rend="italic">Analysis of
                            Network Ensembles</hi>. Ph.D. thesis, University of Konstanz
                        &lt;<ref target="http://nbn-resolving.de/urn:nbn:de:bsz:352-212891">http://nbn-resolving.de/urn:nbn:de:bsz:352-212891</ref>&gt; [15.06.2021]. </bibl>
                    <bibl>
                        <hi rend="bold">Piper, Andrew</hi> / <hi rend="bold">Algee-Hewitt, Mark</hi>
                        (2014): “The Werther Effect I: Goethe, objecthood, and the handling of
                        knowledge”, in: Erlin, Matt / Tatlock, Lynne (eds.): <hi rend="italic"
                            >Distant Readings: Topologies of German Culture in the Long Nineteenth
                            Century</hi>. Cambridge: Cambridge University Press 155-184. </bibl>
                    <bibl>
                        <hi rend="bold">Rochat, Yannick</hi> (2014): <hi rend="italic">Character
                            networks and centrality</hi>. Ph.D. thesis, University of Lausanne. </bibl>
                    <bibl>
                        <hi rend="bold">Trilcke, Peer</hi> (2013): “Social Network Analysis (SNA)
                        als Methode einer textempirischen Literaturwissenschaft”, in: Ajouri, Philip
                        / Mellmann, Katja / Rauen, Christoph (eds.): <hi rend="italic">Empirie in
                            der Literaturwissenschaft</hi>. Münster: mentis 201–247. </bibl>
                    <bibl>
                        <hi rend="bold">Weitin, Thomas</hi> (2016): <hi rend="italic">Fully
                            digitized XML corpus. Der Deutsche Novellenschatz. Published by Paul
                            Heyse, Hermann Kurz. 24 vols. 1871-1876</hi>. Darmstadt / Konstanz &lt;<ref target="http://www.deutschestextarchiv.de/doku/textquellen\#novellenschatz">http://www.deutschestextarchiv.de/doku/textquellen\#novellenschatz</ref>&gt; 
                        [15.06.2021]. </bibl>
                </listBibl>
            </div>
        </back>
    </text>
</TEI>
