Towards the Analysis of Fan Fictions in German Language: Exploration of a Corpus from the Platform Archive of Our Own

Schmidt, Thomas
Media Informatics Group, University of Regensburg, Germany
thomas.schmidt@ur.de

Grünler, Johanna
Media Informatics Group, University of Regensburg, Germany
Johanna.Gruenler@stud.uni-regensburg.de

Schönwerth, Nicole
Media Informatics Group, University of Regensburg, Germany
Nicole.Schoenwerth@stud.uni-regensburg.de

Wolff, Christian
Media Informatics Group, University of Regensburg, Germany
christian.wolff@ur.de

Table of contents

1. Introduction

Online media and content have gained a lot of interest in Digital Humanities (DH) in recent years (e.g. Moßburger et al. 2020; Schmidt et al. 2020a; Schmidt et al. 2020c). In the context of literary studies, the analysis of online creative writing platforms has gained more and more popularity (Hellekson / Busse 2006; Jamison 2013). While some platforms focus on the creation of original content, other platforms like Archive of our Own (AO3)1 and Fanfiction.net2 focus on the specific genre of fan fiction. Fan fictions are fan-created works using already existing characters and plot elements of existing famous media like literature, movies or games to write new stories based on those characters (Dym et al. 2018). Scholars have analyzed the history and cultural influence of this text genre (Cuntz-Leng / Meintzinger 2015; Hellekson / Busse 2006; Jamison 2013; Thomas 2011; Van Steenhuyse 2011). Hellekson and Busse (2006) highlight the striking dominance of slash fan fiction (stories focused on male-male romantic and erotic relationships) in the fan fiction community. Researchers in Natural Language Processing (NLP) make use of the online availability of these large bodies of narrative texts with rich metadata to explore and evaluate new methods (Liu et al. 2019; Muttenthaler et al. 2019; Vilares / Gómez-Rodríguez 2019; Zhang et al. 2019). However, fan fictions themselves have also been subject to computational research. Among other, researchers examine the metadata of fan fictions (Milli / Bamman 2016; Yin et al. 2017; Schmidt / Kleindienst 2020), gender and stereotypes (Fast et al. 2016) and the role and content of user feedback (Frens et al. 2018; Pianzola et al. 2020; Rebora / Pianzola 2018).

Overall, the focus of research is currently on English fan fictions. However, researchers in humanities argue that style, content and progression of fan fictions differ with respect to different regional cultures (Cutz-Leng / Meintzinger 2015). We propose that country-specific features, which are of interest for DH as well as cultural studies, might be expressed in such corpora and should be analyzed to verify such assumptions. Therefore, we want to investigate the benefits of the corpus analysis of non-English fan fiction for the example of German. For the preliminary analysis presented in this abstract, we focus on metadata of fan fiction and how it reflects national-specific features. In future studies we plan to compare the content of multiple languages to each other to identify country specific differences in content and style.

2. Corpus

We have chosen AO3 as the source for creating our preliminary corpus. AO3 describes itself as a non-commercial archive for transformative fan fiction.

We have created a scraper to gather every chapter of every German text on AO3 (more precisely texts marked as German by the creator) and the corresponding metadata by using the language-based search function of AO3. AO3 explicitly allows the scraping of their content in their terms of use. Filtering AO3 for languages shows that 93% of all texts are marked as English, while only 7% are non-English. Overall, the German texts account for 0.2% of all AO3 material only. We acquired the German texts in September 2019. We filtered out any non-German text as well as pages containing solely links, pictures or text pages that were empty. This reduced the overall number of writings to 9,6403.

Next to the text, AO3 offers a rich set of metadata which is currently in the focus of our analysis. Table 1 summarizes the attributes of the items of the corpus. Table 2 illustrates the basic statistics of the corpus. Tokenization was performed via the NLTK standard tokenizer4 and sentence splitting via NLTK and the Punkt sentence splitter5.

Table 1. Structure of a corpus item.

Table 2. Token and sentence statistics.

3. Metadata analysis

We present and discuss results about metadata analysis that show specific expressions of German culture and therefore allow us to further investigate differences and features in online writings and fan culture. To identify nation-specific differences we compared results of our corpus to research on English-dominated corpora (Milli / Bamman 2016; Yin et al. 2017) as well as on fan-based analysis on AO3 in general6.

We found that many of the most popular fandoms are indeed specific expressions of German culture (table 3). The most popular one being Tatort: a German Sunday evening police procedural television series. Fan fictions based on real persons from popular sports in Germany like soccer and ski jumping are also rather popular in Germany. Taking the most popular relationships into account, we also found that stories about the two famous German poets Schiller and Goethe are quite frequent (table 6). Other than that, fandom distributions are similar to other research (Milli / Bamman 2016; Yin et al. 2017) with Harry Potter, Supernatural and Sherlock being among the most popular fandoms. It is often argued that the rise of fan fictions is strongly intertwined with Anime in Germany (Cuntz-Leng / Meintzinger 2015); however, in our corpus the most popular Anime-fandom is Naruto with only 97 stories showing that Anime is not as popular as in general on AO3.

Table 3. Distribution of fandoms.

One of the most striking attributes of the corpus is the dominance of male-male relationships (table 4) and male characters in general as shown by the analysis of most popular characters (table 5), which are predominantly male, and the most popular relationships which are all male (table 6). The popularity of this type of stories is a well-documented attribute of fan fictions (cf. Hellekson / Busse 2006). Please note that this content does not have to be erotic or sexualized but is mostly focused on romance and friendship as can be seen in the analysis of additional tags (table 7). While research in the humanities focuses on explaining this popularity via gender and political discourse (Duggan 2017; Hellekson / Busse 2006; Tosenberger 2008), we also plan to support this research with computational methods.

Table 4. Distribution of relationship categories.

Table 5. Distribution of the most frequent character tags.

Table 6. Distribution of the 10 most frequent character relationships.

Table 7. Distribution of the 10 most frequent additional tags.

While we focus on metadata analysis in this paper, we also have performed some basic text analyses. Table 8 illustrates the most frequent words of the entire corpus after stop word removal. Striking is the rather frequent usage of terms describing physical attributes (augen, hand, kopf, gesicht, stimme). Since the data considering the metadata show that most stories are relationship- and romance-driven, we assume that those terms point to romantic and erotic descriptions and actions of the characters.

Table 8. Most frequent Tokens.

4. Future work

We were able to gain insights about expressions of German culture in fan fictions showing that it is reasonable not to focus on English language corpora alone. We are currently planning to continue our work in various ways: We want to investigate the phenomenon of Tatort-fanfictions more precisely applying methods of distant and close reading. Indeed, there is a long tradition of German literary and media studies for the analysis of this TV show (Buhl 2013). We plan to work closely together with media studies scholars to investigate this special part of German culture in more detail. We also plan to extend the corpus by exploring larger and more popular fan fiction platforms in Germany like FanFiktion.de. Furthermore, we see potential in applying methods that have proven to be beneficial in other DH-contexts like sentiment analysis (Sprugnoli et al. 2016; Schmidt / Burghardt 2018a; Schmidt / Burghardt 2018b; Schmidt et al. 2019b), collocation analysis (Schmidt et al. 2020d), text visualization (Schmidt et al. 2019a), network analysis (Painter et al. 2019) and topic modeling (Schmidt et al. 2020b; Schöch 2021).

Appendix A

Bibliography
  1. Buhl, Hendrik (2013): Tatort: gesellschaftspolitische Themen in der Krimireihe. Konstanz: UVK.
  2. Cuntz-Leng, Vera / Meintzinger, Jacqueline (2015): "A brief history of fan fiction in Germany", in: Transformative Works and Cultures 19 DOI: < https://doi.org/10.3983/twc.2015.0630>.
  3. Duggan, Jennifer (2017): "Revising hegemonic masculinity: Homosexuality, masculinity, and youth-authored Harry Potter fan fiction", in: Bookbird A Journal of International Children's Literature 55, 2: 38-45.
  4. Dym, Brianna / Aragon, Cecilia / Bullard, Julia / Davis, Ruby / Fiesler, Casey (2018): "Online Fandom: Boldly Going Where Few CSCW Researchers Have Gone Before", in: Companion of the 2018 ACM Conference on Computer Supported Cooperative Work and Social Computing. November 3–7, 2018, Jersey City, NJ, USA. ACM 121-124.
  5. Fast, Ethan / Vachovsky, Tina / Bernstein, Michael S. (2016): "Shirtless and dangerous: Quantifying linguistic signals of gender bias in an online fiction writing community", in: Tenth International AAAI Conference on Web and Social Media <https://hci.stanford.edu/publications/2016/ethan/gender.pdf> [02.09.2021].
  6. Frens, John / Davis, Ruby / Lee, Jihyun / Zhang, Diana / Aragon, Cecilia (2018): "Reviews Matter: How Distributed Mentoring Predicts Lexical Diversity on Fan fiction", preprint in: arXiv.org <https://arxiv.org/abs/1809.10268> [02.09.2021].
  7. Hellekson, Karen / Busse, Kristina (2006): Fan Fiction and Fan Communities in the Age of the Internet. New Essays. Jefferson, NC: McFarland.
  8. Jamison, Anne (2013): Fic: Why fan fiction is taking over the world. Dallas, Texas: BenBella Books.
  9. Liu, Chen / Osama, Muhammad / De Andrade, Anderson (2019): "DENS: A Dataset for Multi-class Emotion Analysis", preprint in: arXiv.org <https://arxiv.org/abs/1910.11769>.
  10. Milli, Smitha / Bamman, David (2016): "Beyond canonical texts: A computational analysis of fan fiction", in: Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing 2048-2053.
  11. Moßburger, Luis / Wende, Felix / Brinkmann, Kay / Schmidt, Thomas (2020): "Exploring Online Depression Forums via Text Mining: A Comparison of Reddit and a Curated Online Forum", in: Proceedings of the Fifth Social Media Mining for Health Applications Workshop & Shared Task 70-81.
  12. Muttenthaler, Lukas / Lucas, Gordon / Amann, Janek (2019): "Authorship Attribution in Fan-Fictional Texts given variable length Character and Word N-Grams", in: Working Notes of CLEF 2019 Conference and Labs of the Evaluation Forum.
  13. Painter, Deryc T. / Daniels, Bryan C. / Jost, Jürgen (2019): "Network analysis for the digital humanities: principles, problems, extensions", in: Isis 110, 3: 538-554.
  14. Pianzola, Federico / Rebora, Simone / Lauer, Gerhard (2020): "Wattpad as a resource for literary studies. Quantitative and qualitative examples of the importance of digital social reading and readers’ comments in the margins", in: PloS one 15, 1: e0226708.
  15. Rebora, Simone / Pianzola, Federico (2018): "A New Research Programme for Reading Research: Analysing Comments in the Margins on Wattpad", in: DigitCult-Scientific Journal on Digital Cultures 3, 2: 19-36.
  16. Schmidt, Thomas / Burghardt, Manuel (2018a): "An Evaluation of Lexicon-based Sentiment Analysis Techniques for the Plays of Gotthold Ephraim Lessing", in: Association for Computational Linguistics (ed.): Proceedings of the Second Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature. Santa Fe, New Mexico 139-149.
  17. Schmidt, Thomas / Burghardt, Manuel (2018b): "Toward a Tool for Sentiment Analysis for German Historic Plays", in: Piotrowski, Michael (ed.): COMHUM 2018: Book of Abstracts for the Workshop on Computational Methods in the Humanities 2018. Lausanne, Switzerland: Laboratoire laussannois d'informatique et statistique textuelle 46-48.
  18. Schmidt, Thomas / Kleindienst, Nina (2020): "Investigating the Transformation of Original Work by the Online Fan Fiction Community. A Case Study for Supernatural", in: Digital Practices. Reading, Writing and Evaluation on the Web. November 23–25, 2020, University of Basel, Switzerland.
  19. Schmidt, Thomas / Burghardt, Manuel / Dennerlein, Katrin / Wolff, Christian (2019a): "Katharsis. A Tool for Computational Drametrics", in: Digital Humanities Conference 2019 (DH 2019). Book of Abstracts. Utrecht, Netherlands.
  20. Schmidt, Thomas / Burghardt, Manuel / Wolff, Christian (2019b): "Towards Multimodal Sentiment Analysis of Historic Plays: A Case Study with Text and Audio for Lessing’s Emilia Galotti", in: Proceedings of the DHN (DH in the Nordic Countries) Conference. Copenhagen, Denmark 405-414.
  21. Schmidt, Thomas / Hartl, Philipp / Ramsauer, Dominik / Fischer, Thomas / Hilzenthaler, Andreas / Wolff, Christian (2020a): "Acquisition and Analysis of a Meme Corpus to Investigate Web Culture", in: Digital Humanities Conference 2020 (DH 2020). Virtual Conference.
  22. Schmidt, Thomas / Bauer, Marlene / Habler, Florian / Heuberger, Hannes / Pilsl, Florian / Wolff, Christian (2020b): "Der Einsatz von Distant Reading auf einem Korpus deutschsprachiger Songtexte", in DHd 2020. Book of abstracts 296-299.
  23. Schmidt, Thomas / Kaindl, Florian / Wolff, Christian (2020c): "Distant Reading of Religious Online Communities: A Case Study for Three Religious Forums on Reddit", in: Proceedings of the Digital Humanities in the Nordic Countries 5th Conference (DHN 2020). Riga, Latvia.
  24. Schmidt, Thomas / Kaindl, Florian / Wolff, Christian (2020d): "Visualizing Collocations in Religious Online Forums", in: Digital Humanities Conference 2020 (DH 2020). Virtual Conference.
  25. Schöch, Christof (2021): "Topic modeling genre: an exploration of french classical and enlightenment drama", preprint in: arXiv.org <https://arxiv.org/abs/2103.13019v1> [02.09.2021].
  26. Sprugnoli, Rachele / Tonelli, Sara / Marchetti, Alessandro / Moretti, Giovanni (2016): "Towards sentiment analysis for historical texts", in:  Digital Scholarship in the Humanities 31, 4: 762-772.
  27. Thomas, Bronwen (2011): "What Is Fan fiction and Why Are People Saying Such Nice Things about It?" In: Storyworlds A Journal of Narrative Studies 3: 1-24.
  28. Tosenberger, Catherine (2008): "Homosexuality at the online Hogwarts: Harry Potter slash fan fiction", in: Children's Literature 36, 1: 185-207.
  29. Van Steenhuyse, Veerle (2011): "The writing and reading of fan fiction and transformation theory", in: CLCWeb: Comparative Literature and Culture 13, 4 <https://docs.lib.purdue.edu/cgi/viewcontent.cgi?article=1691&context=clcweb> [02.09.2021].
  30. Vilares, David / Gómez-Rodríguez, Carlos (2019): "Harry Potter and the Action Prediction Challenge from Natural Language", preprint in: arXiv.org <https://arxiv.org/abs/1905.11037>.
  31. Yin, Kodlee / Aragon, Cecilia / Evans, Sarah / Davis, Katie (2017): "Where No One Has Gone Before: A Meta-Dataset of the World's Largest Fan fiction Repository", in: Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems. ACM 6106-6110.
  32. Zhang, Weiwei / Cheung, Jacki C. K. / Oren, Joel (2019): "Generating Character Descriptions for Automatic Summarization of Fiction", in: Proceedings of the AAAI Conference on Artificial Intelligence 33: 7476-7483.
Notes
1.
2.
3.

Due to legal issues the corpus is currently only available upon request via mail (thomas.schmidt@ur.de). We will publish parts of the corpus via the following GitHub repository: https://github.com/lauchblatt/German_Fan_Fictions.

4.
5.
6.