Pages of Early Soviet Performance: Transforming images of Soviet performing arts periodicals into data for computational analysis

Ermolaev, Natalia
Princeton University, United States of America
nataliae@princeton.edu

Puchkovskaia, Antonina
ITMO University, Russian Federation
artonina@gmail.com

Reischl, Katherine
Princeton University, United States of America
kreischl@princeton.edu

Keenan, Thomas
Princeton University, United States of America
tkeenan@princeton.edu

Janco, Andrew
Haverford College, United States of America
ajanco@haverford.edu

Jacobson, Alexander
Princeton University, United States of America
alexander.jacobson@princeton.edu

Kudryashov, Alexander
ITMO University, Russian Federation
alexndr.kudryashov@gmail.com

Our project has created a dataset of rare, early-Soviet illustrated periodicals related to the performing arts, including Rabis, Rabochii teatr, Zreslishcha, 30 dnei and Ermitazh. Through utilizing machine learning techniques we aim to better understand this rich cultural material and to facilitate new avenues of research about Soviet culture during the first decades after the October Revolution (1917-1932). The poster will outline our workflow from scanned images to computer vision models to data for analysis. We used transfer learning to add new labels to a Yolo v5 computer vision model. For this task, we created annotation data using makesense.ai (Skalski 2019-) and a custom annotation tool called Mayakovsky. After initial training on 100 annotations, we further refined the model using 400 annotations to increase precision and the model’s ability to distinguish between text, titles, images, and mixed text categories. Using the trained Yolo model we were able to identify images in the collection and to create a separate collection of images. These files were then labeled with Google Vision and described by a text generation model from IBM. The resulting files and metadata can be viewed and researched using PixPlot from the Yale DH Lab. This is an iterative process where domain experts identify relevant objects, we annotate those objects in the images, then a model is trained, and we assess what the model has learned and interpret its results. EADH participants will gain an introduction to our project and its outcomes as well as a detailed discussion of our process and design choices which can be applied to comparable digitization and digital humanities projects that seek to transform collections into data

Appendix A

Bibliography
  1. Skalski, Piotr (2019-): makesense.ai. Free to use online tool for labelling photos <https://www.makesense.ai/> [21.08.2021].