Pascucci, Antonio
University of Naples L'Orientale, Italia
apascucci@unior.it
Punzi Zarino, Wanda
University of Naples L'Orientale, Italia
wzarino@unior.it
Manna, Raffaele
University of Naples L'Orientale, Italia
rmanna@unior.it
Simoniello, Vincenzo
University of Naples L'Orientale, Italia
vsimoniello@unior.it
Magliacane, Annarita
University of Liverpool
a.magliacane@liverpool.ac.uk
Monti, Johanna
University of Naples L'Orientale, Italia
jmonti@unior.it
Microblogging services have proven to be extremely valuable resources for linguistic data collection because of the variety and the amount of linguistic data that are shared on a daily basis by their users. Previous research (Imran et al. 2015) has also shown that social media platforms, and Twitter in particular, are also commonly used by ordinary people to communicate during emergency crisis, especially natural disasters, and have often provided precious information to first responders. Against this backdrop, we aim to investigate the ability to leverage collective intelligence (Büscher et al. 2014) as a resource for monitoring crimes against the environment through linguistic data derived from social media platforms.
In particular, this paper aims to address the following research questions: i) Can Natural Language Processing (NLP) techniques be used to detect environmental crimes on Twitter? ii) What kind of linguistic features is it possible to identify in tweets reporting such crimes? iii) Is it possible to use such features in order to automatically filter alert tweets?
The research is carried out within the framework of the project Crowd for the Environment (in short C4E), a project which aims at developing innovative linguistic methods for the detection and monitoring of illegal spills, such as illegal landfills, micro-dumps or illegal releases in surface waters. The project is conducted with a view to aiding and enhancing the on-site action of local organizations responsible for environmental and territorial protection with the application of NLP techniques to linguistic data taken from social media platforms. In our research special attention was dedicated in the preliminary phases of our project to the issue of La Terra dei Fuochi (literally the Land of Fires) (Peluso 2015), a sadly famous large area located between Naples and Caserta, in the south of Italy, victim of illegal toxic waste which has been dumped by criminal organizations for about fifty years and which is currently routinely burned to make space for new toxic one. The focus of the research was then extended geographically to the entire Italian peninsula and was thematically expanded to include more and different types of environmental crimes and not exclusively the fires of toxic waste in this specific geographical area.
In order to achieve our goals, two linguistics resources have been compiled: the C4E - Environmental Crimes Glossary of Terms (available to download from the project website) and the Unior Eye Corpus. Initially, a glossary of terms, consisting of 43 terms related to environmental crimes, was created. To compile the glossary, we consulted and merged domain-specific glossaries on environmental issues containing definitions of the terms in English and Italian as well as we referred to legal documents about sentences for environmental crimes. The compilation of this linguistic resource represented a crucial starting point for our research project as it allowed us to build the corpus on the grounds of thematic relevance (i.e., the presence of keywords and hashtags). Subsequently, we compiled the corpus by downloading from Twitter all tweets containing the terms from our glossary, preceded by hashtags. Hashtags are a form of social tagging that allows microbloggers to embed metadata in social media posts. They are marked with a # symbol and may include a word, initialism, or even an entire clause. Moreover, they generally perform different types of meanings or linguistic functions. For example, they can indicate the core semantic domain of the tweet or they can also be used to link the tweet/post to the account of another user as well as expressing a comment or opinion about the theme of the tweet/post (Zappavigna 2015). In this study, we will mainly focus on their use as thematic aggregators because their use in the tweets allowed us to gather linguistic data about environmental crimes and, in turn, to develop methods which allowed us to detect reports of such crimes by Twitter users.
The resulting corpus is made up of 228,412 tweets, 22,780,746 tokens and 569,905 types with a type/token ratio (TTR) of 0.025. It is further divided into four semantic sections (and their relative subsections), each focusing on a specific environmental crime:
However, due to the abundance of information that may circulate on social media on environmental issues (i.e. sharing news, opinions, personal reflections), the first step of our analysis was to identify the features of tweets aiming at reporting a specific crime from all other tweets. These reporting tweets have been referred to as 'alerts' and have been distinguished from all other tweets (non-alerts) dealing with environmental issues but performing different functions.
Broadly speaking, it is possible to define an alert tweet as a tweet which contains enough details (i.e. geographical location, time reference, type of environmental crimes) which may be of help in environmental crime detection. In other words, it can be mentioned that alert tweets have a higher informative value and their level of informativeness can vary according to the amount and the quality of the details provided in the tweet. Therefore, in our analysis, an alert tweet is a post where the platform user (be it the private citizen or an association/organization) provides accurate information on the environmental crime with the intention of reporting such an environmental crime and helping on-site action on the part of first responders or environmental organizations. Conversely, non-alert tweets do not provide such information and details and include all tweets which do not have the specific aim of reporting a crime. They generally take the form of ironic messages, aggressive messages, comments on municipal/regional/national administrations, messages of open protest, opinions around the crimes under consideration, hypothesis of users, news about sanctions already carried out against the crime. Consequently, although non-alerts tweets may be still be relevant to increase the awareness of the public domain toward environmental issues, due to the fact that they do not help detect onsite crimes, a more fine-grained analysis of the big category 'non-alert' tweets has not been considered in the analysis as it was beyond the scope of the paper, i.e. the use of NLP techniques to detect environmental crimes.
After establishing the distinction between alert and non-alert tweets, we focused on the textual features of the ‘alerts’ in order to further analyze how their use can enhance prompt on-site action and, in turn, help to monitor environmental crimes. For this reason we implemented a further step in the tagging of the alerts by using named-entity recognition (NER). Being these tweets written by private individuals or organizations, the level of informativeness and details (the informativeness value) they provide can extremely vary, i.e. some tweets can contain extremely detailed information which may help to detect a specific crime while others may not be that informative and the lack of details may hinder a quick intervention on site. Therefore, by using a corpus-driven and bottom up approach on a specific section of the corpus (La Terra dei Fuochi) then applied to other sections of the corpus, we identified a number of textual features that are present in alert tweets commonly written by Twitter users (i.e. time or geographical reference) and we are developing a model that can help to (semi)automatically detect such features with the use of NLP techniques.
In conclusion, linguistic data collected from Twitter have proven to be extremely useful in emergency situations of natural disasters and more specifically they represent a way to detect and monitor environmental crimes. However, the amount and the quality of the features possessed by an alert tweet are directly proportional to its informativeness. In other words, a tweet in which the type of environmental crime, the time reference and the exact geographical location are included will certainly be more informative than a tweet in which the social media user merely reports the type of crime or does not provide exact details regarding the location or the time reference. Consequently, the use of NLP techniques on alert tweets about environmental crimes can help to detect such information and can be used to safeguard and protect the environment from human-made disasters. Furthermore, we plan to enlarge our glossary and create an ontology to offer a formal knowledge organization related to environmental issues. Finally, having discriminated between alert and non-alert tweets, future work involves the use of computational stylometry on the dataset consisting of non-alert tweets to identify hate speech and fake news.