🌍 ThendoLabs
ASR for 10 African languages · TTS for isiZulu, Setswana, and Yoruba
| Text | Language |
|---|
About these models
ASR — w2v-bert-2.0 fine-tunes: a 9-language polyglot head (Dagbani, Igbo, Hausa,
Ibibio, Yoruba, Nigerian Pidgin, Setswana, Adamawa Fulfulde, Twi/Akan) and a dedicated
isiZulu head. Decoding uses the production per-language KenLM beam search (tuned per
language); isiZulu decodes greedily (no KenLM built for it yet).
TTS — StyleTTS 2 fine-tunes for isiZulu (NCHLT corpus, CC-BY 3.0; explicit click phonemes ǀ ǃ ǁ via a rule-based G2P) and Setswana (SLR32; the BR3 champion, with English code-switch routing), plus a Yoruba VITS voice (OpenSLR-129). Voices derive from corpus speakers; native-listener evaluation gated every release.
Limitations — lexical tone is not modeled (no digital tone lexicons exist for these languages); typing English into the isiZulu box will produce Zulu-rule pronunciations (including clicks on the letter c); single-listener validation per language so far.
Attribution — NCHLT isiZulu (SADiLaR, CC-BY 3.0) · OpenSLR SLR32 / SLR129 ·
LibriTTS (CC-BY 4.0) · facebook/w2v-bert-2.0 (MIT) · StyleTTS 2 (MIT) ·
Coqui VITS (MPL-2.0). The MMS-300m adversarial critic (CC-BY-NC) was used at
training time only and is not part of, nor distributed with, these models.