<?xml version="1.0" encoding="UTF-8"?>
<TEI xmlns="http://www.tei-c.org/ns/1.0">
    <teiHeader>
        <fileDesc>
            <titleStmt>
                <title>Talking about Europe: exploring 70 years of news archives</title>
                <author>
                    <persName>
                        <surname>Bergamini</surname>
                        <forename>Enrico</forename>
                    </persName>
                    <affiliation>University of Turin, Italy</affiliation>
                    <email>enrico.bergamini@unito.it</email>
                </author>
                <author>
                    <persName>
                        <surname>Mourlon-Druol</surname>
                        <forename>Emmanuel</forename>
                    </persName>
                    <affiliation>University of Glasgow, UK</affiliation>
                    <email>Emmanuel.Mourlon-Druol@glasgow.ac.uk</email>
                </author>
            </titleStmt>
            <editionStmt>
                <edition>
                    <date>2021-09-14T15:30:00Z</date>
                </edition>
            </editionStmt>
            <publicationStmt>
                <publisher>Elisabeth Burr, University of Leipzig</publisher>
                <address>
                    <addrLine>Beethovenstr. 15</addrLine>
                    <addrLine>04107 Leipzig</addrLine>
                    <addrLine>Germany</addrLine>
                    <addrLine>Elisabeth Burr</addrLine>
                </address>
            </publicationStmt>
            <sourceDesc>
                <p>Converted from a Word document</p>
            </sourceDesc>
        </fileDesc>
        <encodingDesc>
            <appInfo>
                <application ident="DHCONVALIDATOR" version="1.22">
                    <label>DHConvalidator</label>
                </application>
            </appInfo>
        </encodingDesc>
        <profileDesc>
            <textClass>
                <keywords scheme="ConfTool" n="category">
                    <term>Paper</term>
                </keywords>
                <keywords scheme="ConfTool" n="subcategory">
                    <term>Long paper</term>
                </keywords>
                <keywords scheme="ConfTool" n="keywords">
                    <term>European public opinion</term>
                    <term>big data</term>
                    <term>machine learning</term>
                    <term>LDA</term>
                    <term>digital humanities</term>
                    <term>digital history</term>
                    <term>topic modelling</term>
                    <term>media analysis</term>
                    <term>European history</term>
                    <term>natural language processing</term>
                    <term>large text analysis</term>
                    <term>distant reading</term>
                </keywords>
                <keywords scheme="ConfTool" n="topics">
                    <term>Discovering</term>
                    <term>Gathering</term>
                    <term>Programming</term>
                    <term>Writing</term>
                    <term>Annotating</term>
                    <term>Cleanup</term>
                    <term>Content Analysis</term>
                    <term>Structural Analysis</term>
                    <term>Modeling</term>
                    <term>Archiving</term>
                    <term>Identifying</term>
                    <term>Communicating</term>
                    <term>Publishing</term>
                    <term>DigitalHumanities</term>
                    <term>Text</term>
                    <term>Language</term>
                    <term>Software</term>
                    <term>ResearchProcess</term>
                    <term>Data</term>
                    <term>not applicable</term>
                    <term>English</term>
                </keywords>
            </textClass>
        </profileDesc>
    </teiHeader>
    <text>
        <body>
            <p>This paper quantitatively explores news coverage on ‘Europe’ through three newspapers, respectively in France, Italy and Germany, by means of a novel big dataset of 13 million full-text daily articles, spanning from 1945 to 2019. The question of how frequently the media have talked about Europe throughout the course of European integration is one of evident relevance. </p>
            <p>Analysis of the public debate around given issues in European integration has long
                been of interest to European studies’ scholars, in particular historians. Some have
                studied the press based on a specific case study, such as Mathias Haeussler who
                analysed the debate about Europe in the British popular press through an analysis of
                the <hi rend="italic">Daily Express</hi> and <hi rend="italic">Daily Mirror</hi> of
                the early 1960s (Haeussler 2014). Haeussler stresses the opposing narratives of the
                two newspapers: the <hi rend="italic">Express</hi> opposed European integration,
                while the <hi rend="italic">Mirror</hi> favoured UK’s entry to the EEC. </p>
            <p>Diez Medrano used the ‘quality press’ in Germany, Spain, and the UK to shed light on
                these countries’ different attitudes to European integration. The methodology is
                based on a close-reading of “a sample of newspaper editorials and opinion pieces
                published in British, German, and Spanish quality newspapers between 1946 and 1997,
                the year I ended my fieldwork” (Díez Medrano 2003: 16). In the case of German press,
                Diez Medrano analysed even years only for the <hi rend="italic">Frankfurter
                    Allgemeine Zeitung</hi>, and odd years only for <hi rend="italic">Die Zeit</hi>,
                and concentrated on op-ed articles only (Díez Medrano 2003: 267–269). Reading these
                articles allowed him to code them into several categories (attitude to European
                integration and to transfer of sovereignty being two of the most relevant). He used
                a sample of 90 articles in the <hi rend="italic">Frankfurter Rundschau</hi> between
                1950 and 1995 to verify his findings, as this journal is more overtly leftist and
                regionalist (<hi rend="italic">FAZ</hi> being conservative-liberal and <hi
                    rend="italic">Die Zeit</hi> liberal). The analysis allows Diez Medrano to draw
                conclusions on the negative and positive comments about European integration (Díez
                Medrano 2003: 106–156). </p>
            <p>Media attention plays an important role in forming public opinion and allowing
                citizens to form their ideas. Especially following the 2010s, when arguably the
                European unification process suffered a setback compared to the acceleration of the
                1990s, public discourse about Europe, in national media outlets, became even more
                relevant. But how can we measure the media interest in European affairs? The
                digitalisation of public archives has opened doors to the application of statistical
                analysis and natural language processing. Statistical and machine-learning based
                techniques now make it possible to analyse large amounts of historical data. The
                combination of digitised databases and computational techniques, at the intersection
                between humanities and computer science, has helped answer questions from different
                fields: history, journalism, public opinion and economics (Broersma / Harbers 2018). </p>
            <p>Our paper builds in this quantitative direction and focuses on the history of the
                coverage of the European Union in the news. In a similar context, Müller et al.
                (2018), have explored the dynamics of the “blame game”, looking at which countries
                were blamed for the financial crisis. In the aftermath of the financial crisis of
                2008, different media in different countries reported different reactions around
                blame. </p>
            <p> Vliegenthart et al. (2008) have investigated the effects of European media presence
                at the aggregate level, in the context of the European unification process. They
                explored the role of framing news in terms of benefits or conflict and how this
                framing affects public opinion and support for the European project. They found that
                media coverage, especially if framed in terms of conflict, has a significant effect
                on public opinion. Public opinion was, in this case, proxied by survey data on trust
                in European institutions and perceived benefits from EU’s membership. </p>
            <p>We aim to contribute to the understanding of Europe across European media. Our computational analysis makes use of longer-spanning newspapers archives than previous exercises in the literature, ranging from the end of the Second World War to the end of the 2010s. We collect and organise large web-scraped datasets covering the period from 1945 to 2019 for three relevant newspapers. This large-scale and unprecedented ‘distant reading’ analysis of the digitalised archives of three European newspapers over more than 70 years and 13 million articles allows to identify the overall rising share of European news in printed media. We detail the procedure involved in obtaining and systemising these archives. The archives consist of daily (or weekly in the case of 
                <hi rend="italic">Der Spiegel</hi>), full-text articles, collecting all the material present in the web-archives, which has previously been transformed into machine-readable text via OCR. They are intended to proxy for the media narrative, and how we made it as consistent across time and as comparable across countries as possible. We selected three outlets of three of the founding members of the European Union: 
                <hi rend="italic">La Stampa</hi> (Italy), 
                <hi rend="italic">Der Spiegel</hi> (Germany) and 
                <hi rend="italic">Le Monde</hi> (France). 
            </p>
            <p>The magnitude of these news archives obviously makes a large-scale manual labelling of the single articles too resource intensive to be possible. We propose a new methodology, after ensuring the quality of the archives, for identifying articles referring to “European” news while leaving aside national and other non-European news, based on a mix of keywords matching, large-scale natural language processing and topic identification on the full text news articles. For example, the archives contain information about sports, which may explicitly refer to Europe, but are outside the scope of our analysis, which considers the “European” topic in its commonly understood institutional, political, and economic sense.</p>
            <p>In a first stage, we develop a three-language taxonomy of words that captures very
                broadly this concept of Europe. We firstly identify the set of articles in the
                corpus, which match the keywords. In a second stage we identify the topics within
                it, by performing large-scale Latent Dirichlet Allocation (Blei et al. 2003). By
                deciding which topics identified are potentially in line with our definition of
                “Europe”, we further filter down the subset of “European” articles. We propose a
                method for flagging only the articles more likely to be “European”, based on both
                the cumulative score coming from the keywords match, and the degree of certainty we
                have over each topic within the subset. We estimate the topic models for each of the
                three archives separately: our analysis is language dependent, and it requires
                building three different models. One of the key parameters to tune when fitting an
                LDA model is the number of topics potentially present in the corpus. A priori, this
                number is unknown, particularly when dealing with unstructured data. We run the same
                model while varying the number of topics and use a common evaluation metric, the
                coherence score (CV metric), to choose the best performing one. We rely on the
                MALLET framework (McCallum 2002). </p>
            <p>To the state of our knowledge, we believe that our methodological contribution can be
                of help to scholars aiming at extracting an a-priori defined topic from large and
                unstructured datasets. Conceptually, we can define a measure of frequency of
                European news pieces as E/T, where E is the number of European news, while T is the
                total number of news. We can investigate how this variable evolved over time and
                across countries. Once articles are classified and the datasets labelled, we perform
                a time-series analysis, detect salient events of European history, across France,
                Germany and Italy. We implement a simple z-score algorithm in order to flag events
                in the three daily time series (Brakel 2020). We analyse these events in light of
                the evolution of European cooperation and integration since 1945. We find that the
                most important events in post-war European history are easily identifiable in the
                archives and that European issues gather substantially greater attention starting in
                the early 1990s. </p>
            <figure>
                <graphic n="1001" width="16.51cm" height="7.877527777777778cm" url="Pictures/4efbc11ce0e8d8afaca767b5711e6327.png" rend="inline"/>
            </figure>
            <p>Our study thus contributes to considerably widen and systematise the scope of previous studies on the press, thanks to the use of a novel quantitative methodology. This allows us to observe the growth of coverage of European news over national ones, with an acceleration from the mid 1990s. Our study not only confirms that European elections and summits are traditionally high points of media coverage since the end of the Second World War, it also identifies some key milestones in European integration, from the rejection of the European Defence Community in 1954 to the euro crisis summits of the 2010s.</p>
        </body>
        <back>
            <div type="bibliogr">
                <listBibl>
                    <head>Bibliography</head>
                    <bibl rend="Bibliography" style="text-align: left; ">
                        <hi rend="bold">Blei, David M.</hi> / <hi rend="bold">Ng, Andrew Y.</hi> /
                            <hi rend="bold">Jordan, Michael I.</hi> (2003): "Latent Dirichlet
                        Allocation", in: <hi rend="italic">The Journal of Machine Learning
                            Research</hi> 3: 993–1022. </bibl>
                    <bibl rend="Bibliography" style="text-align: left; ">
                        <hi rend="bold">Brakel, Jean-Paul</hi> (2020): <hi rend="italic">Peak Signal
                            Detection in Realtime Timeseries Data - Smoothed Z-Score Algorithm (Peak
                            Detection with Robust Threshold)</hi>. Stack Overflow &lt;<ref
                            target="https://stackoverflow.com/questions/22583391/peak-signal-detection-in-realtime-timeseries-data"
                            >https://stackoverflow.com/questions/22583391/peak-signal-detection-in-realtime-timeseries-data</ref>>
                        [14.09.2021). </bibl>
                    <bibl rend="Bibliography" style="text-align: left; ">
                        <hi rend="bold">Broersma, Marcel</hi> / <hi rend="bold">Harbers, Frank</hi>
                        (2018): "Exploring Machine Learning to Study the Long-Term Transformation of
                        News: Digital Newspaper Archives, Journalism History, and Algorithmic
                        Transparency", in: <hi rend="italic">Digital Journalism</hi> 6, 9: 1150–1164
                        DOI: 10.1080/21670811.2018.1513337. </bibl>
                    <bibl rend="Bibliography" style="text-align: left; ">
                        <hi rend="bold">Díez Medrano, Juan</hi> (2003):<hi rend="italic">Framing
                            Europe: Attitudes to European Integration in Germany, Spain, and the
                            United Kingdom</hi> (nachdr. Princeton). New Jork: Princeton Univ.
                        Press. </bibl>
                    <bibl rend="Bibliography" style="text-align: left; ">
                        <hi rend="bold">Haeussler, Mathias</hi> (2014): "The Popular Press and Ideas
                        of Europe: The Daily Mirror, the Daily Express, and Britain’s First
                        Application to Join the EEC, 1961-63", in: <hi rend="italic">Twentieth
                            Century British History</hi> 25, 1: 108–131 DOI: 10.1093/tcbh/hws050. </bibl>
                    <bibl rend="Bibliography" style="text-align: left; ">
                        <hi rend="bold">McCallum, Andrew Kachites</hi> (2002): <hi rend="italic"
                            >MALLET: A Machine Learning for Language Toolkit</hi>. </bibl>
                    <bibl rend="Bibliography" style="text-align: left; ">
                        <hi rend="bold">Müller, Henrik</hi> / <hi rend="bold">Porcaro, Giuseppe</hi>
                        / <hi rend="bold">von Nordheim, Gerret</hi> (2018): <hi rend="italic">Tales
                            from a Crisis: Diverging Narratives of the Euro Area</hi>. Bruegel
                        Policy Contribution Issue N˚03 | February 2018 &lt;<ref
                            target="http://bruegel.org/2018/02/tales-from-a-crisis-diverging-narratives-of-the-euro-area/"
                            >http://bruegel.org/2018/02/tales-from-a-crisis-diverging-narratives-of-the-euro-area/</ref>>
                        [14.09.2021). </bibl>
                    <bibl rend="Bibliography" style="text-align: left; ">
                        <hi rend="bold">Vliegenthart, Rens et al.</hi> (2008): "News Coverage and
                        Support for European Integration, 1990-2006", in: <hi rend="italic"
                            >International Journal of Public Opinion Research</hi> 20, 4: 415–439
                        DOI: 10.1093/ijpor/edn044. </bibl>
                </listBibl>
            </div>
        </back>
    </text>
</TEI>
