To read this content please select one of the options below:

Towards corpora creation from social web in Brazilian Portuguese to support public security analyses and decisions

Victor Diogho Heuer de Carvalho (Eixo das Tecnologias, Universidade Federal de Alagoas, Campus do Sertão, Delmiro Gouveia, Brazil)
Ana Paula Cabral Seixas Costa (Departamento de Engenharia de Produção, Centro de Tecnologia e Geociências, Universidade Federal de Pernambuco, Recife, Brazil)

Library Hi Tech

ISSN: 0737-8831

Article publication date: 25 October 2022

Issue publication date: 23 July 2024

198

Abstract

Purpose

This article presents two Brazilian Portuguese corpora collected from different media concerning public security issues in a specific location. The primary motivation is supporting analyses, so security authorities can make appropriate decisions about their actions.

Design/methodology/approach

The corpora were obtained through web scraping from a newspaper's website and tweets from a Brazilian metropolitan region. Natural language processing was applied considering: text cleaning, lemmatization, summarization, part-of-speech and dependencies parsing, named entities recognition, and topic modeling.

Findings

Several results were obtained based on the methodology used, highlighting some: an example of a summarization using an automated process; dependency parsing; the most common topics in each corpus; the forty named entities and the most common slogans were extracted, highlighting those linked to public security.

Research limitations/implications

Some critical tasks were identified for the research perspective, related to the applied methodology: the treatment of noise from obtaining news on their source websites, passing through textual elements quite present in social network posts such as abbreviations, emojis/emoticons, and even writing errors; the treatment of subjectivity, to eliminate noise from irony and sarcasm; the search for authentic news of issues within the target domain. All these tasks aim to improve the process to enable interested authorities to perform accurate analyses.

Practical implications

The corpora dedicated to the public security domain enable several analyses, such as mining public opinion on security actions in a given location; understanding criminals' behaviors reported in the news or even on social networks and drawing their attitudes timeline; detecting movements that may cause damage to public property and people welfare through texts from social networks; extracting the history and repercussions of police actions, crossing news with records on social networks; among many other possibilities.

Originality/value

The work on behalf of the corpora reported in this text represents one of the first initiatives to create textual bases in Portuguese, dedicated to Brazil's specific public security domain.

Keywords

Acknowledgements

This work was financed in part by the Coordenação de Aperfeiçoamento de Pessoal de Nível Superior–Brasil (CAPES) - Finance Code 001, by the Conselho Nacional de Desenvolvimento Científico e Tecnológico – Brasil (CNPq), and by the Universidade Federal de Alagoas–Brasil (UFAL).

Citation

de Carvalho, V.D.H. and Costa, A.P.C.S. (2024), "Towards corpora creation from social web in Brazilian Portuguese to support public security analyses and decisions", Library Hi Tech, Vol. 42 No. 4, pp. 1080-1115. https://doi.org/10.1108/LHT-08-2022-0401

Publisher

:

Emerald Publishing Limited

Copyright © 2022, Emerald Publishing Limited

Related articles