CSRC
● liveCongolese Speech Radio Corpus — unlabelled radio archives for self-supervised pretraining and acoustic adaptation.
Every release lives on Hugging Face under BantuLanguagesInitiative. Source code and pipelines on GitHub. Standards documented and enforced.
Congolese Speech Radio Corpus — unlabelled radio archives for self-supervised pretraining and acoustic adaptation.
Lingala Read Speech Corpus — the labelled Lingala read-speech subset, packaged to load directly with the Hugging Face datasets library.
We'd rather have one well-documented dataset than ten unusable scrapes. Every contribution is held to the same bar.
Every dataset ships with a data sheet: provenance, collection method, consent, known biases.
Cleaning and processing scripts live in the public repo. Anyone can rerun the pipeline.
Default to CC-BY-4.0. If a more restrictive license is needed, we explain why on the dataset card.
Contributors are named on the dataset page. Communities providing data are consulted and acknowledged.