This AI research company wants to put 1,000 African languages into AI
{
"title": "The 100 Billion Token Gambit: African Languages Lab Confronts AI's Lingua Franca Bias",
"article": "For years, the promise of artificial intelligence has often overlooked the vast linguistic diversity of Africa. The fundamental problem, as identified by African Languages Lab founder Sheriff Issaka, was a severe scarcity of adequate data required to train AI systems for African languages.\n\nAfrican Languages Lab, established by Issaka in 2020 as an AI research and deployment company, initially faced a critical bottleneck. Instead of relying solely on insufficient existing datasets, the company pivoted to actively collect the necessary data. This strategic shift has culminated in what the company claims is now the largest collection of African-language data, encompassing over 70 languages, notably Amharic, Hausa, Zulu, Twi, Igbo, and Yoruba.\n\nOn Tuesday, this monumental data curation effort bore fruit with the launch of Mansa, a multilingual and multimodal AI platform. Mansa has put approximately 30 of these languages into production, making them accessible via web, mobile, and through APIs for developers and businesses. The scale of this undertaking is underscored by the company’s assertion that its datasets comprise more than 100 billion curated tokens—the textual units used for AI model training—alongside over 19,000 hours of speech recordings, all rigorously reviewed and validated by language experts.\n\nWhile Africa boasts over 2,000 languages, a stark reality remains: fewer than 5% possess the digital resources essential for natural language processing. This severe underrepresentation in digital datasets directly impedes AI systems' ability to comprehend these languages effectively. Issaka himself minced no words, stating, “The models are just very bad at understanding our languages.” He further noted the impracticality of having a "full-blown conversation with an LLM in most African languages” due to this limitation.\n\nBeyond functional inadequacy, this data deficit carries significant economic and safety implications. Generating tokens in African languages can be disproportionately expensive; Chioma Agwuedo, Executive Director of TechHerNG, reported in 2025 that producing a token in Yoruba could cost four times as much as one in English. This cost barrier could hinder local innovation and adoption. More critically, Issaka highlighted a profound safety gap: models trained with less data exhibit weaker safety tuning. He found that harmful prompts in certain African languages were at least 10 times more likely to elicit erroneous responses that would typically be blocked in better-resourced languages, posing a serious ethical challenge.\n\nAfrican Languages Lab's initiative is not isolated but part of a growing continental effort to bridge this AI language divide. Global players, such as Google, have also expanded AI Search to include African languages and launched WAXAL, an open-source speech dataset covering 21 Sub-Saharan African languages. This collective push signals a critical inflection point, moving beyond passive observation of the linguistic data gap to active investment in foundational resources crucial for equitable AI development in African markets.\n\nThe launch of Mansa and the underpinning data collection by African Languages Lab represent a tangible step towards digital equity. By addressing the critical scarcity of African language data, the platform has the potential to mitigate the technical, economic, and safety disadvantages currently faced by speakers of these languages in the AI ecosystem. The challenge remains vast, given Africa's linguistic tapestry, but the 100 billion token gambit is a formidable starting point.",
"tweet": "AI's language gap is no joke. African Languages Lab just launched Mansa, pushing 30 languages into production after collecting 100B tokens & 19K hrs of speech. Meanwhile, some African languages still cost 4x more & are 10x less safe in AI. The digital divide speaks many tongues. #AfricanAI #Mansa #Tech",
"excerpt": "For years, AI's promised revolution sidestepped the vast linguistic landscape of Africa. With over 2,000 languages, less than 5% possess the digital resources needed for natural language processing, leaving models 'very bad' at understanding the continent's diverse voices. Now, African Languages Lab is challenging this disparity head-on, launching Mansa, a multilingual AI platform, after a monumental effort to curate over 100 billion tokens and 19,000 hours of speech data.",
"keywords": "African Languages Lab, Mansa, AI, African languages, data scarcity, natural language processing, Sheriff Issaka, digital equity, language technology, TechHerNG, Google WAXAL"
}