Schmidt, Thomas
Media Informatics Group, University of Regensburg, Germany
thomas.schmidt@ur.de
El-Keilany, Alina
Media Informatics Group, University of Regensburg, Germany
Alina.El-Keilany@stud.uni-regensburg.de
Eger, Johannes
Media Informatics Group, University of Regensburg, Germany
johannes.eger@stud.uni-regensburg.de
Kurek, Sarah
Media Informatics Group, University of Regensburg, Germany
sarah.kurek@stud.uni-regensburg.de
Quantitative methods have a long tradition in film analysis going back to the predigital era (Salt 1974; Vonderau 2020). Nowadays, multiple projects explore movies via computational methods to investigate colors (Burghardt et al. 2016, 2018; Flueckiger 2017; Kurzhals et al. 2016; Masson et al. 2020; Pause / Walkowski 2018), shot lengths (Baxter et al. 2017; DeLong 2015) or annotation possibilities (Halter et al. 2019; Kuhn et al. 2015; Schmidt / Halbhuber 2020; Schmidt et al. 2020a). Recent research has also led to the definition of the term Distant Viewing (Arnold / Tilton 2019) to describe large-scale digital movie analysis. A lot of the current research is focused on the analysis of text via scripts or subtitles (Byszuk 2020; Holobut et al. 2016; Holubut / Rybicki 2020; Hoyt et al. 2014). However, developments in computer vision have led to novel methods for the image channel of movies and are already applied in computer science to develop recommender systems (Deldioo et al. 2016; Wei et al. 2004) but also in Digital Humanities (DH) to analyze movies (Howanitz et al. 2019; Pustu-Iren et al. 2020; Zaharieva et al. 2012) and other visual media (Schmidt et al. 2020e). We argue that these methods are beneficial for digital film studies and give new perspectives.
We present an exploratory study for the methods: Object detection, emotion recognition, gender- and age-prediction. We apply state-of-the-art models on a subset of frames of five different movies of varied decades and genres. We apply the exploratory research approach defined by Wulff (1998) for traditional film analysis in this study for computational approaches. Our goals are (1) to inspect the benefits and problems of the methods, (2) explore if the methods uncover specific characteristics of the movies and (3) what research questions seem promising to follow in further large-scale studies.
We limited the analysis on five movies. Table 1 presents the movies and metadata. For all movies except Avengers, we use a digitally restored version. All movies have a 720x576 resolution, 25 frames per second and 32 bits per sample. We focus on canonical work and Hollywood productions.

Table 1. Movies and metadata.
All analysis was performed in Python 3. We extracted the frames of every movie since all of the applied methods are image-based. However, we take one frame per second of a movie and regard this as the sample of a movie. We decided to employ this approach because using all frames makes the data processing very performance/resource-intensive and we argue that one frame per second offers sufficient information for our first explorations.
To perform the object detection, we use Detectron2 (Wu et al. 2019) which offers state-of-the-art object detection models by Facebook AI Research1. We use a pretrained masked RCCN-model trained on the well-known COCO-Dataset (Lin et al. 2015), which can predict 80 object classes including vehicles, animals, and sports objects. Applying this prediction model on an image, we receive the number of predicted objects, the locations, and the prediction confidence (0-100%). As threshold for the detection, we select 50% which is usually very low but fits our exploratory approach.
Emotion recognition is a sub-field of affective computing (cf. Halbhuber et al. 2019; Hartl et al. 2019; Ortloff et al. 2019; Schmidt et al. 2020c) and is often applied in DH to predict sentiment and emotions from written text (Moßburger et al. 2020; Schmidt / Burghardt 2018; Schmidt, 2019; Schmidt et al. 2019a; Schmidt et al. 2020b). We focus on the image channel of movies and for the emotion prediction we use the Python module FER2 (Goodfellow et al. 2013). The module first performs face detection via a MTCNN Face Detector3 (Zhang et al. 2016) and then predicts the emotion via a convolutional neural network (CNN) trained on over 35,000 images. The model predicts the seven classes anger,disgust,fear,happiness,sadness,surprise and neutral on a scale from 0 to 1. All values sum up to 1 for one face.
We perform gender- and age-prediction via the module py-agender4 which is also a CNN trained on the IMDB-Wiki dataset (Rothe et al. 2018) consisting of over 500,000 faces. The model achieves a mean average error of 4.08 on standardized datasets (Agustsson et al. 2017). For the gender prediction the model produces a value between 0 and 1, with values below 0.5 being male and above being female faces.
We summarize the results of the object detection by looking at the 10 most frequent objects overall and per movie. Table 2 and 3 show the objects starting with the most frequent per unit. Freq is the absolute number of detected instances while % is the percentage of frames at least one of the specific objects was detected.

Table 2. Detected objects per movie and overall (part 1).

Table 3. Detected objects per movie and overall (part 2).
Persons are the most frequently detected “objects” (figure 1). Other frequent objects are mostly furniture (book, chair), clothes (tie, handbag) and drinking objects (cup, wine glass).

Figure 1. Frame with the most detected persons ( Metropolis).
Comparing the movies, we identified that movies below 90% of frames with persons are indeed the more action-oriented movies ( Avengers,Metropolis) or include fantasy/animal-like characters ( Wizard of Oz). Many modern objects (e.g cell phones and airplanes) are more frequent in the contemporary movie Avengers (figure 2). One outlier we identified is the clock-object in Metropolis, which is not a frequent object in the other movies but represents a well-studied reoccurring motif of this specific movie (figure 3; cf. Cowan 2007).

Figure 2. Detected airplanes in Avengers.

Figure 3. Clocks as a reoccurring motif in Metropolis.
While we did not perform a systematic evaluation, but we identified a lot of mistakes in the prediction e.g. guns were predicted as handbags or the character “Cowardly Lion” in Wizard of Oz was oftentimes predicted as dog (figure 4).

Figure 4. The “Cowardly Lion” in Wizard of Oz detected as „dog“.
Nevertheless, we see potential in the method of object detection to explore specifics of the mise-en-scène as well as motif-like reoccurring objects in movies (Zaharieva / Breiteneder 2012). Furthermore, as object classes of the COCO dataset are not necessarily fitting for movies, we recommend exploring the possibilities of post-training via Detectron to analyze objects that are not part of the pretrained models.
For the emotion recognition we decided to create an average for a frame if multiple faces are detected. If no face is detected, we mark the frame with missing values. Table 4 summarizes the results. Maximums and minimums are marked in bold.

Table 4. Emotion values per movie and overall (M=mean, Max=maximum, Sd=standard deviation).
Overall, highest averages for emotions are the neutral (M=0.24) and the sad class (M=0.29). Surprise (M=0.11) and disgust (M=0.00) are rather rare among the movies. The two comedies in the movie corpus (Wizard of Oz, Some Like it Hot) do indeed have the highest happy-averages (M=0.13) (figure 5).

Figure 5. Frame with maximum happy value (Some Like it Hot).
However, the results are rather inconsistent since Wizard of Oz has also the highest sad- and angry-averages and therefore is the movie with generally the strongest emotional expressions. Breakfast at Tiffany’s on the contrast is the most neutral movie (M=0.37; figure 6).

Figure 6. Frame with highest neutrality value in the corpus (Breakfast at Tiffany’s).
Additionally, we performed a Welch-ANOVA to investigate if the movies differ to each other significantly (all requirements for the test are met according to Field (2009)). Indeed, we do find significant differences (p<0.05) for all emotion categories but rather small effects according to Cohen (1988) defining η²<0.01 as weak, <0.06 as moderate and <.14 as strong effect. We report the p-, F- and η²-value (table 5).

Table 5. Results of Welch-ANOVA-Tests for all emotion categories
The strongest effect can be seen for neutral. Performing post-hoc tests and inspecting a box-plots graph (figure 7) we identified Breakfast at Tiffany’s as interesting outlier. This might be due to the fact that the main characters of the movie try to stay rather “unaffected” up until the ending of the movie while Wizard of Oz, as a musical, consist of strong emotional outbursts.

Figure 7. Box-plots graph for the emotion class neutral
Table 6 illustrates the descriptive statistics for the gender- and age-detection.

Table 6. Descriptive statistics for age and average gender.
The average age is for most movies is around 40 which is a rather consistent over-estimation since most leading actors in the selected movies are around 30. Performing a Welch-ANOVA shows that the difference between the movies is significant (p<0.001, F=336.07, η²=0.09) with a moderate effect. The strongest outlier movie, as shown with post hoc tests, is Wizard of Oz with a child/teenager as leading actor that gets correctly detected as around 14-16 years old (figure 8).

Figure 8. Lowest age in the corpus (Wizard of Oz).
An average score for gender below 0.5 points to more male detections and it is striking that all movies point below 0.5, thus a more frequent representation of males which is in line with the reality of the movies. There is a significant difference considering gender but with a smaller effect compared to age (p<0.001, F=251.36, η²=0.06) and with the strongest differences concerning Wizard of Oz. The differences become apparent regarding the distribution of gender-classes (table 7). We assigned every frame with male if average gender > 0.6 and female if <0.4. We decided to include a class androgynous for in-between-values pointing to either multiple genders on one screen or uncertainty by the model.

Table 7. Frequency distributions of gender classes.
Wizard of Oz has the most frames classified as androgynous. In general, this means that female and male characters are equally on the frame but in this case the classification is due to the high number of human-like fantasy creatures for which the model is unsure to pick a gender (figure 9).

Figure 9. An “androgynous” face (Wizard of Oz).
While this study was rather small and exploratory in the approach, we did gain important first insights for our future research. Overall, we find it promising that we were able to find significant results, even for this small set of movies. For object detection we see the most potential in adjusting pretrained models to objects that are of interest for a specific research question. We see a lot of potential for interesting diachronic but also genre-based emotion and gender analysis with larger corpora. For this case study, we did not find striking differences of method performance considering technical differences between the movies. We are planning systematic evaluations on a cross section of movies of different decades to get a better understanding on the performance of the methods before we move on to explore more concrete research questions. Modern cultural artefacts have shown to be of interest for gender studies in the DH context (Schmidt et al. 2020d). We see potential concerning research on the intercourse of gender and film studies. We plan to explore the relationship of gender representations with expressed emotions throughout the time to explore how the representation of gender roles developed. Furthermore, we want to also explore multimodal approaches combining the various modality channels of movies (similar to Schmidt et al. 2019b).
More Information: https://github.com/facebookresearch/detectron2
More Information: https://pypi.org/project/fer/
More Information: https://github.com/ipazc/mtcnn
More Information: https://github.com/yu4u/age-gender-estimation