bli://opendatasetsv0.1
/ Open datasets

The data layer. Public, documented, reusable.

Every release lives on Hugging Face under BantuLanguagesInitiative. Source code and pipelines on GitHub. Standards documented and enforced.

Hosted on
01
Hugging Face
Default license
02
CC-BY-4.0
Code
03
GitHub · Open
First release
04
2026
/ 01 · Catalogue

Published datasets.

View on Hugging Face ↗

CSRC

● live

Congolese Speech Radio Corpus — unlabelled radio archives for self-supervised pretraining and acoustic adaptation.

Language
Lingala · Kikongo · Tshiluba
Type · Size
audio · radio · 10K–100K samples
Open ↗

LRSC

● live

Lingala Read Speech Corpus — the labelled Lingala read-speech subset, packaged to load directly with the Hugging Face datasets library.

Language
Lingala
Type · Size
audio + transcript · 1K–10K samples
Open ↗
/ 02 · How to contribute

Quality standards.

We'd rather have one well-documented dataset than ten unusable scrapes. Every contribution is held to the same bar.

01

Document the source.

Every dataset ships with a data sheet: provenance, collection method, consent, known biases.

02

Make it reproducible.

Cleaning and processing scripts live in the public repo. Anyone can rerun the pipeline.

03

License for reuse.

Default to CC-BY-4.0. If a more restrictive license is needed, we explain why on the dataset card.

04

Credit contributors.

Contributors are named on the dataset page. Communities providing data are consulted and acknowledged.