This paper presents the methodology developed for the MultiPRIDE shared task by the Hate Busters team, which focuses on automatically identifying reappropriative intent in multilingual social media messages containing LGBTQ+ related slurs. Reappropriation is a nuanced sociolinguistic phenomenon in which historically derogatory terms are deliberately reclaimed with positive or community-affirming meanings. To model this behavior, we work with the multilingual dataset provided for the MultiPRIDE task, including texts in Italian, Spanish, and English. Our approach integrates a comprehensive methodological pipeline, encompassing preprocessing of noisy social media text, the construction of traditional sparse lexical representations such as Bag-of-Words and TF–IDF, and contextual Transformer-based embeddings from XLM-RoBERTa, followed by training linear classifiers and fine-tuned Transformer models. The goal is to systematically compare how different modelling paradigms capture the subtle and context-dependent characteristics of reappropriative language.
Ciminelli, A., Corvino, G., Gentili, C., Viviani, M. (2026). The Hate Busters at MultiPRIDE: Automatic Identification of Reappropriated Slurs in Multilingual LGBTQ+ Discourse. In Proceedings of the 9th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA 2026) (pp.1-17). CEUR-WS.
The Hate Busters at MultiPRIDE: Automatic Identification of Reappropriated Slurs in Multilingual LGBTQ+ Discourse
Viviani M.
2026
Abstract
This paper presents the methodology developed for the MultiPRIDE shared task by the Hate Busters team, which focuses on automatically identifying reappropriative intent in multilingual social media messages containing LGBTQ+ related slurs. Reappropriation is a nuanced sociolinguistic phenomenon in which historically derogatory terms are deliberately reclaimed with positive or community-affirming meanings. To model this behavior, we work with the multilingual dataset provided for the MultiPRIDE task, including texts in Italian, Spanish, and English. Our approach integrates a comprehensive methodological pipeline, encompassing preprocessing of noisy social media text, the construction of traditional sparse lexical representations such as Bag-of-Words and TF–IDF, and contextual Transformer-based embeddings from XLM-RoBERTa, followed by training linear classifiers and fine-tuned Transformer models. The goal is to systematically compare how different modelling paradigms capture the subtle and context-dependent characteristics of reappropriative language.| File | Dimensione | Formato | |
|---|---|---|---|
|
Ciminelli et al-2026-EVALITA-CEUR-VoR.pdf
accesso aperto
Tipologia di allegato:
Publisher’s Version (Version of Record, VoR)
Licenza:
Creative Commons
Dimensione
1.93 MB
Formato
Adobe PDF
|
1.93 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


