Hate Speech Detection in Turkish and Arabic: A Comprehensive Study
About
Online hate speech has been linked to a global rise in violence against minorities, including incidents such as mass shootings, lynchings, and ethnic cleansing. Societies grappling with this issue, particularly when hate speech targets specific groups based on religion, race, ethnicity, culture, nationality, or migration status, face the challenge of balancing freedom of expression with the need for effective content moderation on widely used online platforms. In response to this challenge, we introduce a comprehensive hate speech dataset covering five distinct topics in Turkish: refugees, the Israel-Palestine conflict, anti-Greek sentiment in Turkey, ethnic or religious communities (Alevis, Armenians, Arabs, Jews, and Kurds), and LGBTI+, alongside one topic in Arabic (refugees). In addition, we develop state-of-the-art BERT-based models to address multiple dimensions of hate speech analysis, including hate category classification, hate intensity prediction, target identification, and hate speech span detection, enabling a comprehensive understanding of hateful content in online discourse.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Multi-label Target Classification | Turkish Dataset | Precision80 | 8 | |
| Multi-label Target Classification | Arabic Dataset | Precision86 | 8 | |
| Binary Span Detection | Turkish Dataset | Precision0.57 | 1 | |
| Span Categorization (Multi-class) | Turkish Dataset | Precision29 | 1 | |
| Hate Intensity Prediction | Turkish Dataset | -- | 1 | |
| Hate Intensity Prediction | Arabic Dataset | -- | 1 |