<?xml version="1.0" encoding="UTF-8"?>
<TEI xmlns="http://www.tei-c.org/ns/1.0">
    <teiHeader>
        <fileDesc>
            <titleStmt>
                <title>Evaluation of Face detection Algorithms for of a Corpus of Early Modern Portraits </title>
                <author>
                    <persName>
                        <surname>Diem</surname>
                        <forename>Sebastian</forename>
                    </persName>
                    <affiliation>Stiftung Universität Hildesheim, Germany</affiliation>
                    <email>diem@uni-hildesheim.de</email>
                </author>
                <author>
                    <persName>
                        <surname>Üresin</surname>
                        <forename>Esra</forename>
                    </persName>
                    <affiliation>Stiftung Universität Hildesheim, Germany</affiliation>
                    <email>ueresin@uni-hildesheim.de</email>
                </author>
                <author>
                    <persName>
                        <surname>Mandl</surname>
                        <forename>Thomas</forename>
                    </persName>
                    <affiliation>Stiftung Universität Hildesheim, Germany</affiliation>
                    <email>mandl@uni-hildesheim.de</email>
                </author>
                <author>
                    <persName>
                        <surname>Beyer</surname>
                        <forename>Hartmut</forename>
                    </persName>
                    <affiliation>Herzog August Bibliothek Wolfenbüttel, Germany</affiliation>
                    <email>beyer@hab.de</email>
                </author>
                <author>
                    <persName>
                        <surname>Niedermeier</surname>
                        <forename>Nina</forename>
                    </persName>
                    <affiliation>Herzog August Bibliothek Wolfenbüttel, Germany</affiliation>
                    <email>niedermeier@hab.de</email>
                </author>
                <author>
                    <persName>
                        <surname>Rößler</surname>
                        <forename>Hole</forename>
                    </persName>
                    <affiliation>Herzog August Bibliothek Wolfenbüttel, Germany</affiliation>
                    <email>roessler@hab.de</email>
                </author>
            </titleStmt>
            <editionStmt>
                <edition>
                    <date>2021-06-14T17:23:00Z</date>
                </edition>
            </editionStmt>
            <publicationStmt>
                <publisher>Elisabeth Burr, University of Leipzig</publisher>
                <address>
                    <addrLine>Beethovenstr. 15</addrLine>
                    <addrLine>04107 Leipzig</addrLine>
                    <addrLine>Germany</addrLine>
                    <addrLine>Elisabeth Burr</addrLine>
                </address>
            </publicationStmt>
            <sourceDesc>
                <p>Converted from a Word document</p>
            </sourceDesc>
        </fileDesc>
        <encodingDesc>
            <appInfo>
                <application ident="DHCONVALIDATOR" version="1.22">
                    <label>DHConvalidator</label>
                </application>
            </appInfo>
        </encodingDesc>
        <profileDesc>
            <textClass>
                <keywords scheme="ConfTool" n="category">
                    <term>Paper</term>
                </keywords>
                <keywords scheme="ConfTool" n="subcategory">
                    <term>Long paper</term>
                </keywords>
                <keywords scheme="ConfTool" n="keywords">
                    <term>Face Detection</term>
                    <term>Portraits</term>
                </keywords>
                <keywords scheme="ConfTool" n="topics">
                    <term>Discovering</term>
                    <term>Imaging</term>
                    <term>Content Analysis</term>
                    <term>Stylistic Analysis</term>
                    <term>Visualization</term>
                    <term>Meta: GiveOverview</term>
                    <term>Images</term>
                    <term>Metadata</term>
                    <term>ResearchProcess</term>
                    <term>not applicable</term>
                    <term>not applicable</term>
                    <term>not applicable</term>
                    <term>not applicable</term>
                    <term>not applicable</term>
                    <term>English</term>
                </keywords>
            </textClass>
        </profileDesc>
    </teiHeader>
    <text>
        <body>
            <div type="div1" rend="DH-Heading1">
                <head> Introduction early modern portraits prints </head>
                <p>The study of portraits opens opportunities to study cultural, social and artistic dimensions of the early modern period. The use of modern image processing technologies allows to enter new directions in research e.g. the analysis of social and artistic traditions. </p>
                <p>This article briefly reports on experiments with image processing systems for a well curated collection of 35,000 portraits mainly from the 15
                    <hi rend="superscript">th</hi> to the 19
                    <hi rend="superscript">th</hi> century. 
                </p>
                <p>The study of portraits has been used for studying art history as well as social and cultural history. The focus of research has been mainly on the individual rather than on its historical and economic context, the media environment and its materiality. Portrait research in art history has often focused on paintings and thus neglected the large quantities of traditions of printed portraits. Portraits can show how individuals wanted to be depicted and which visual traditions and social norms they followed. </p>
                <p>Face detection and face recognition have become established fields in computer vision. A variety of applications are based on large pre-trained models consisting of millions of realistic images. In the experiments reported here, we use a state-of-the-art face detection model called OpenFace (Baltrusaitis et al. 2018) to analyse the applicability of those models on printed portraits from the early modern period until the 19
                    <hi rend="superscript">th</hi> century. 
                </p>
                <p>First, an overview for face detection projects will be given. A short introduction to the portrait collection will follow. In the next part, we describe the setup of the experiment including the pre-processing. The results part first discusses the applicability of the OpenFace model based on different observations. We show notable influences of the printing technique on the face detection rate and observe the model’s confidence based on head size and the number of faces in the image. In the second part, we report trends and correlations in the portrait collection like the gaze direction compared to the body’s orientation. In the conclusion, we discuss the applicability of this approach and the validity of the observed trends. </p>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head>Related work</head>
                <p>OpenFace 2.0 implements multiple tasks from the basic face detection itself to eye and head pose estimation and even the interpretation of facial expressions. State of the art landmark detection models are based on cascading regression models in combination with convolutional neural networks which got enhanced over the last few years (Valle et al. 2019). While OpenFace’s gaze estimation had a mean angular error of 7.9° (Wood et al. 2015) modern approaches like FewShot are able to get as low as 3.14° angular error (Park et al. 2019). </p>
                <p>Faces in art have been automatically analysed before. The beauty of faces
                    depicted between the 13 <hi rend="superscript">th</hi> and 18 <hi
                        rend="superscript">th</hi> century was studied by de la Rosa and colleagues.
                    For face detection analysis, the authors used 120,000 portraits and received
                    25,000 portraits with 47,000 faces detected. They observed that the symmetry of
                    faces is the highest between the 15 <hi rend="superscript">th</hi> and 18 <hi
                        rend="superscript">th</hi> century. (De la Rosa et al. 2015) This is in line
                    with the FACES study which used a 15 <hi rend="superscript">th</hi> to 18 <hi
                        rend="superscript">th</hi> century portrait corpus to compare two different
                    feature detection approaches, due to the more naturalistic representations. The
                    authors also mention that the detection rate depends on the artistic style and
                    artistic quality (Rudolph et al. 2017).</p>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head>Portraits and image analysis</head>
                <p>The Herzog August Library owns a collection of over 35,000 portraits which are used as a foundation for the creation of an image recognition model. The goals of ongoing research is to find and use similarities to identify origins of portraits, their composition and other important visual features. We hope to find similarities between collections of portraits, conventions of representations, group specific symbolism and submittals for derived works and copies. We use OpenFace which was previously tested in a pre-study (Üresin 2020). With this tool we hope to gain first insights which might be useful in later applications. </p>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head> Experiment setup </head>
                <p>Before the analysis, portraits needed to be cropped, so that unnecessary elements in the digital image do not distort the evaluation. Cropped portraits were extracted based on the most outer portrait frame. A non-pretrained DLA-34 Neural Network (Yu et al. 2018) was used for this process. Due to the variety in the dataset about 20,000 portraits from about 27,000 have been successfully cropped and were used further for the face detection. </p>
                <p>OpenFace offers different combinations of face and landmark detectors where the recommended detectors have been used based to the most promising results. As a result, OpenFace created a csv file containing the confidence of the model, eye position, gaze direction, eye landmarks, face landmarks, face position and face rotation for every found face. The metadata of the portraits provides useful information regarding the year the portrait was created, the body orientation of the person in the portrait and the printing technique if available. </p>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head> Face detection results </head>
                <p>From the 19,913 portraits OpenFace detected 6,267 potential faces. For 518 portraits multiple faces were detected. This results in a detection rate of about 28.9% which is in line with previous studies (De la Rosa et al. 2015) and an evaluation of 250 protraits in a pre-study (Üresin 2020).</p>
                <p>The detection rate can further be dissected using insights of the different printing techniques from the metadata. Of the 26 printing techniques present inside the corpus, only 18 techniques led to successful identifications of faces. Table 1 shows the detection rate per technique for techniques over 100 examples. It is clear, that Steel engraving has the highest detection rate with 59.3% whereas Lithography and Mezzotint have a detection rate above 40%. Copper Engraving, which represents by far the biggest portion in the dataset, has a detection rate of 28%. It is also important to notice that Wood Cut portraits have a detection rate of only 3.4%. This indicates that there are substantial differences between the detectability of different techniques by a modern computer vision approach.</p>
                <figure>
                    <graphic n="1001" width="15.980833333333333cm" height="2.4553333333333334cm" url="Pictures/063b568395207686a4584b2fb17c5960.png" rend="inline"/>
                </figure>
                <p>Table 1</p>
                <p>In a next step, the certainty of the model is measured with a confidence value. The confidence over all detected faces measured 88.6%.</p>
                <figure>
                    <graphic n="1002" width="9.359194444444444cm" height="11.789833333333334cm" url="Pictures/111c9805b0a82b559d897b318a759be8.jpeg" rend="inline"/>
                </figure>
                <p>Figure 1</p>
                <p>91.6% of all found faces have a confidence above 90% and 6.7% below 10% confidence which indicates that the model is either very confident or extremely unconfident. Comparing the confidence with the head size a correlation can be observed where the smallest head sizes have the lowest confidence and contain the most false-positives, like Figure 1 (based on manual evaluation). We can conclude that the confidence value is helpful for identifying errors. </p>
                <p>For multiple faces per portrait the biggest head size is the first recorded detection. The confidence for the first detection has an average of about 90.1% confidence whereas for every next detection it drops to 71.1%. Based on these results we filtered out detections with a confidence below 40% due to their unreliability. Using this approach, we filtered out 443 portraits. </p>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head>Face Orientation</head>
                <p>The orientation of a portrait is one of the artist’s decisions and an important
                    feature for analysing artistic traditions. Besides the position of the eyes in
                    the portrait OpenFace also gives the gaze direction for each eye, visually
                    indicated by a gaze vector. This vector has x,y,z coordinates to indicate the
                    direction the person is looking outgoing from the eye of the portraited person.
                    In x-dimension 3448 (59,1%) persons look to the left and 2384 to the right based
                    on the observer’s point of view. Most of the gazes are around -15° to +15°. In
                    y-dimension 95,7% look down with the majority between -30° and -15°. The
                    horizontal head rotation is measured by the Pose_Ry (yaw) where 3431 (58,8%)
                    heads are oriented to the right as seen in Figure 2. For both orientations, the
                    focus lies between 5° and 20°.</p>
                <figure>
                    <graphic n="1003" width="16.002cm" height="9.144cm" url="Pictures/7da837f04ac2220d196afa31bdb37d6e.jpeg" rend="inline"/>
                </figure>
                <p>Figure 2</p>
                <p>Looking at the head pose distribution over the years from about the 15
                    <hi rend="superscript">th</hi> to the 19
                    <hi rend="superscript">th</hi> century no visible trend is observable (Figure 3). The general head pose distribution of 60% to 40% is still visible based on the opacity of the point clouds.
                </p>
                <figure>
                    <graphic n="1004" width="16.4465cm" height="7.6793861111111115cm" url="Pictures/46ab3d7eb9fd4481bab27a64d668d1bc.jpeg" rend="inline"/>
                </figure>
                <p>Figure 3 </p>
                <p>Lastly, a correlation between the eye, face and body orientation can be seen.
                    Comparing the eye to the face orientation a strong negative correlation can be
                    calculated, which indicates that the gaze direction is opposite to the face
                    orientation. Comparing the body orientation to the gaze direction (Figure 4a)
                    this opposing trend can also be observed, while the face orientation is equal to
                    the body orientation as seen in Figure 4b. This implies that the body and the
                    face are oriented in one direction and the eyes in the opposite direction, which
                    would make sense of the eyes were to look at the viewers point of view. This
                    correlates with the observations made like in Figure 2.  </p>
                <figure>
                    <graphic n="1005" width="15.980833333333333cm" height="7.0485cm" url="Pictures/b13075e4883b5dcacadb6aa1cb110452.jpeg" rend="inline"/>
                </figure>
                <p>Figure 4</p>
            </div>
            <div type="div1" rend="DH-Heading1">
                <head>Conclusion</head>
                <p>Our analysis showed that state-of-the-art approaches like OpenFace can be applied on very different kinds of dataset like printed portraits out of the box. Even though the overall detection rate is very low with only every fourth portrait detected, the detected faces have a low error rate and basic features of faces can be correctly identified and used for mass processing and trend analysis. It also showed that there are substantial differences regarding the detection rate based on the printing technique. Therefore, further explorations in this field could be useful. </p>
            </div>
        </body>
        <back>
            <div type="bibliogr">
                <listBibl>
                    <head>Bibliography</head>
                    <bibl style="text-align: left;">  
                        <hi rend="bold">Baltrusaitis, Tadas / Zadeh, Amir / Lim, Yao Chong / Morency, Louis-Philippe</hi> (2018): "OpenFace 2.0: Facial Behavior Analysis Toolkit", in: 
                            <hi rend="italic">IEEE International Conference on Automatic Face and Gesture Recognition.</hi>
                       </bibl>
                    <bibl style="text-align: left;">
                        <hi rend="bold">De la Rosa, Javier / Suárez, Juan-Luis</hi>(2015): "A Quantitative Approach to Beauty. Perceived Attractiveness of Human Faces in World Painting", in: 
                        <hi rend="italic">International Journal for Digital Art History.</hi>
                    </bibl>
                    <bibl style="text-align: left;">
                        <hi rend="bold">Park, Seonwook / De Mello, Shalini / Molchano, Pavlo / Iqba, Umar / Hilliges, Ottmar / Kautz, Kautz</hi> (2019). "Few-Shot Adaptive Gaze Estimation", 
                        in: <hi rend="italic">arXiv: 1905.01941v2.</hi>
                    </bibl>
                    <bibl style="text-align: left;">
                        <hi rend="bold">Kohl, Jeanette / Rudolph, Conrad / Srinivasan, Ramya / Roy-Chowdhuy, Amit</hi> (2017): "FACES: Faces, Art, and Computerized Evaluation Systems - A Feasibility Study of the Application of Face Recognition Technology to Works of Portrait Art", in:
                        <hi rend="italic">artibus et historiae</hi> 75, XXXVIII.
                    </bibl>
                    <bibl style="text-align: left;">
                        <hi rend="bold">Üresin, Esra</hi> 2020: "Digital Humanities: Evaluation of CNN-based face detection systems for portraits", Stiftung Universität Hildesheim, Fachbereich 3.
                    </bibl>
                    <bibl style="text-align: left;">
                        <hi rend="bold">Valle, Roberto / Buenaposada, José M. / Valdés, Antonio / Baumela, Luis</hi> (2019): "Face Alignment using a 3D Deeply-initialized Ensemble of Regression Trees", in: 
                        <hi rend="italic">arXiv:1902.01831v2.</hi>
                    </bibl>
                    <bibl style="text-align: left;">
                        <hi rend="bold">Wood, Erroll / Baltrusaitis, Tadas / Zhang, Xucong / Sugano, Yusuke / Robinson, Peter / Bulling, Andreas</hi> (2015): "Rendering of Eyes for Eye-Shape Registration and Gaze Estimation", in:
                        <hi rend="italic">IEEE International Conference on Computer Vision</hi> (ICCV).
                    </bibl>
                    <bibl style="text-align: left;">
                        <hi rend="bold">Yu, Fisher / Wang, Dequan / Shelhamer, Evan / Darrell, Trevor</hi> (2018): "Deep Layer Aggregation", in: 
                        <hi rend="italic">Computer Vision and Pattern Recognition</hi> (CVPR).
                    </bibl>
                    <bibl style="text-align: left;">
                        <hi rend="bold">Zadeh, Amir / Baltrusaitis, Tadas / Morency, Louis-Philippe</hi> (2017): "Convolutional Experts Network for Facial Landmark Detection", in: 
                        <hi rend="italic">Computer Vision and Pattern Recognition Workshops</hi>.
                    </bibl>
                </listBibl>
            </div>
        </back>
    </text>
</TEI>
