AfroScope: A Framework for Studying the Linguistic Landscape of Africa
About
Language Identification (LID), the task of determining the language of a given text, is a fundamental preprocessing step that shapes the reliability of downstream NLP applications. While recent work has expanded African LID, existing systems remain limited in both language coverage and fine-grained discrimination among closely related languages and varieties. We introduce AfroScope, a unified framework for African LID that includes AfroScope-Data, a dataset covering 640 languages, and AfroScope-Models, a suite of strong LID models with broad African language coverage. To address persistent confusions among closely related languages, we propose a hierarchical classification approach that leverages AfroScope-Mirror, a specialized embedding model for targeted disambiguation, improving macro-F1 by 1.57 points on the confusable subset compared to our best base model. We further analyze cross-lingual transfer and domain effects, showing how language-family structure, script compatibility, and domain coverage shape LID performance. We position African LID as an enabling technology for large-scale measurement of Africa's linguistic landscape in digital text, and release AfroScope-Data and AfroScope-Models online.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Hierarchical classification | AfroScope-Data confusable (test) | Macro F198 | 36 | |
| Language Identification | AfroScope High resource | Macro-F1100 | 16 | |
| Language Identification | AfroScope-Data Mid resource | Macro F198.67 | 8 | |
| Language Identification | Afroscope | Macro-F197.83 | 5 | |
| Language Identification | BLOOM | Macro F195.76 | 5 | |
| Language Identification | FineWeb2 | Macro F194.52 | 5 | |
| Language Identification | Mafand | Macro F193.54 | 5 | |
| Language Identification | MCS-350 | Macro F1 Score70.38 | 5 | |
| Language Identification | Smol | Macro F190.02 | 5 | |
| Language Identification | UDHR | Macro F189.68 | 5 |