SAUTI — Swahili Voice AI Platform
Built by FiniFlow Labs | GitHub
Interactive demo of 5 Swahili voice AI products. All models run on this HuggingFace Space. First use may take 30-60s to load models.
Speak Swahili to an AI assistant
Record a message in Swahili (or type below). The AI understands and responds in Swahili with voice.
Pipeline: Your voice → ASR (Whisper) → LLM (Llama 3.1) → TTS (MMS) → AI voice
Or type in Swahili:
Swahili Speech Recognition
Upload or record Swahili audio. Our fine-tuned Whisper model transcribes it with 13.5% WER — trained on 89 hours of Swahili speech.
Model: Finiflowlabs/sauti-asr-v1
Clone Any Voice
Upload 6-30 seconds of someone's voice, then type text to hear the AI speak in that voice. Uses XTTS v2 zero-shot cloning.
Note: Voice cloning requires GPU hardware for acceptable speed. On CPU, synthesis takes ~30-60s per utterance. XTTS v2 works best with English text; Swahili uses cross-lingual transfer.
Swahili Text-to-Speech
Type Swahili text and hear it spoken using Meta's MMS-TTS model.
Model: facebook/mms-tts-swh
English ↔ Kiswahili Translation
Speak or type in one language, hear and read the translation in the other.
Pipeline: Your voice → ASR (Whisper) → MT (NLLB-200) → TTS (MMS) → Translated voice
Note: First use loads the translation model (~60s on CPU, ~10s on GPU).
Or type text to translate:
Built by FiniFlow Labs | ASR Model | Contact: hello@finiflowlabs.com