bli-asr-1 · Lingala ASR
Second iteration of our Lingala speech-recognition model — Whisper large-v3 adapted with LoRA, trained on Waxal and our own LRSC corpus. Open weights on Hugging Face.
Open datasets, NLP/ASR/TTS models and language tooling for the 500+ Bantu languages spoken across Central, East and Southern Africa — starting with Lingala.
Hundreds of millions of speakers. Almost no representation in major AI systems. Both numbers are real.
The largest language models in the world are trained on data that's overwhelmingly English, Mandarin, and a handful of European languages. When 350 million people who speak a Bantu language try to use a voice assistant, a chatbot, a search engine — the systems often don't even pretend to listen.
It's not a quirk of the technology. It's a question of which voices were judged worth recording. We exist to change that answer — not by hand-waving about “AI for Africa”, but by shipping the dull, exact, foundational work: clean datasets, reproducible pipelines, public benchmarks.
If you can't train on it, evaluate on it, and improve it in the open — it doesn't exist for AI. We're changing that.
Lingala is the lingua franca of Kinshasa — a city of fifteen million and one of the fastest-growing in the world — and of much of the Congo River basin. It is sung across Africa, broadcast on dozens of radio stations, and used daily in markets, offices, schools and ministries.
It is also functionally invisible to the AI systems that increasingly mediate banking, education and healthcare. Starting here is a choice with weight: a major African language, a rich oral and written tradition, and a measurable gap we can close.
“Mbote na yo, ndenge nini ozali?”
Hello, how are you? · [mbote na jo, ndenge nini oˈzali]
Second iteration of our Lingala speech-recognition model — Whisper large-v3 adapted with LoRA, trained on Waxal and our own LRSC corpus. Open weights on Hugging Face.
Congolese Speech Radio Corpus — unlabelled radio archives in Lingala, Kikongo and Tshiluba, for self-supervised pretraining and acoustic adaptation.
Lingala Read Speech Corpus — the labelled Lingala subset, repackaged to load directly with the Hugging Face datasets library.
Our first speech-recognition model for Lingala — Whisper large-v3 adapted with LoRA. Open weights on Hugging Face.