Speech Recognition & Voice AI for Local Languages

Duration
8 weeks
Investment
UGX 700,000
Certificate
Included
Teaching
Live online
Program Introduction
This advanced speech recognition and Voice AI course is designed for developers, data scientists, researchers and digital-product teams in Uganda, East Africa and across Africa. Learners work with audio data, speech-to-text, text-to-speech, responsible local-language datasets, accent adaptation and inclusive IVR/USSD design, then build a small farming-advice prototype that responds to spoken Luganda questions.
Key Features & Benefits
• Uganda and East Africa context • Hands-on speech-to-text and text-to-speech • Responsible local-language dataset preparation • Local-accent and low-resource model adaptation • Inclusive IVR and USSD voice-interface design • Luganda farming-advice capstone
Real-World Applications
• Voice-enabled agricultural advisory and farmer helplines • Local-language customer support and call-centre transcription • Accessible IVR and USSD services for low-literacy users • Radio and media transcription or monitoring • Voice interfaces for education, public services and community information • African-language voice assistants and smart-device controls
Course outline and learning expectations
This is a tutor-led course. The outline shows what your tutor will cover; teaching materials and examinations are provided directly to enrolled students.
Live online
English (Uganda)
University · General Public
What you will learn
- Explain how speech recognition systems process audio data
- Use speech-to-text tools and APIs to transcribe recorded or live speech
- Build and evaluate a basic text-to-speech workflow
- Assess speech-model performance across local languages, accents and recording conditions
- Plan and prepare a responsibly collected local-language audio dataset
- Adapt or fine-tune a speech model for a selected local language or accent
- Design an inclusive voice interface that combines IVR, USSD and speech
- Build and demonstrate a small voice-enabled tool that responds to spoken Luganda questions
Modules
- 1
Speech Recognition Foundations
Understand audio as data and the main stages of automatic speech recognition
WaveformsSampling rate and bit depthSpectrograms and acoustic featuresASR pipelineAudio quality and noise - 2
Speech-to-Text Tools and APIs
Use cloud and open-source speech-to-text tools and compare their output on African accents
API authentication and requestsAudio formatsBatch and streaming transcriptionConfidence scoresWord error rateTesting Ugandan and East African accents - 3
Text-to-Speech Basics
Build and evaluate simple text-to-speech workflows for local-language applications
Text normalizationPronunciation and phonemesVoice synthesis pipelineSSML basicsNaturalness and intelligibility testing - 4
Local Languages and Accent Challenges
Analyse why existing speech models may underperform on low-resource African languages and diverse accents
Data scarcityCode-switchingDialect and accent variationDomain mismatchNoise and recording conditionsFairness and bias - 5
Responsible Local-Language Audio Datasets
Plan, collect, label and document speech data with consent and community safeguards
Use-case definitionSpeaker consentSampling and representationRecording protocolsTranscription and annotationQuality controlPrivacy and dataset documentation - 6
Adapting Speech Models for Local Accents
Adapt or fine-tune a speech model and measure results on a selected Ugandan or African language
Transfer learningData preparationTrain-validation-test splitsFeature extractionFine-tuning workflowError analysisModel evaluation - 7
Voice Interfaces for Low-Literacy Users
Design inclusive voice services that combine IVR, USSD and speech for practical African contexts
Conversational flow designPrompt writingFallbacks and confirmationsIVR integrationUSSD hand-offAccessibilityOffline and low-bandwidth considerations - 8
Capstone Voice-Enabled Tool
Build a small voice-based farming-advice IVR prototype that responds to spoken Luganda questions
Problem definitionLuganda question setSpeech-to-text integrationIntent or response logicText-to-speech outputUser testingDocumentation and demonstration
Before you enroll
- Having Successfully learnt Course:NLP for African Languages
What you need
- Computer with microphone
- Headphones
- Stable internet connection
- Python 3
- Jupyter Notebook or Google Colab
- Code editor such as Visual Studio Code
- Audio editor such as Audacity
- Access to a speech-to-text and text-to-speech API or open-source speech model
Frequently asked questions
More in Machine Learning

Introduction to Computer Vision

Big Data Fundamentals for AI

MLOps & AI Deployment at Scale

Ensemble Learning & Advanced Model Techniques

Version Control & Collaborative Coding with Git & GitHub

Unsupervised Learning & Clustering
More for University · General Public

Dart Programming for Beginners in Uganda & East Africa

Advanced C Systems Programming Course – Africa

Advanced C# Programming in Uganda: Async, Generics & Performance

Advanced C++ Course Uganda – Performance and Concurrency

Advanced Computer Vision & Image Recognition

Advanced Dart Programming & Concurrency for Africa
Quick Actions
Need Help?
Have questions about this program? Our team is here to help!