Generative AI — Image, Audio & Video Models

Duration
6 weeks
Investment
UGX 650,000
Certificate
Included
Teaching
Live online
Program Introduction
Generative AI — Image, Audio & Video Models is an advanced, practical program for developers, designers, content teams and digital innovators in Uganda, East Africa and across Africa. Learners examine how generative models differ from predictive systems, how diffusion-based image generation works, and how modern tools create and adapt images, audio and video. The program combines technical experimentation with responsible use, covering copyright, licensing, consent, data protection, bias, cultural representation and disclosure. Practical activities lead to a capstone: an AI-assisted image and caption generator for a small Ugandan business.
Key Features & Benefits
• Advanced multimodal generative AI workflows • Hands-on image generation and lightweight model adaptation • Audio and video prototyping • Responsible AI, privacy and copyright practice • Uganda and Africa-focused business examples • Capstone project for a small Ugandan business
Real-World Applications
• Create campaign images and captions for Ugandan small businesses • Prototype product visuals and design mockups for e-commerce and retail • Develop voiceover, music and sound concepts for digital campaigns • Produce short video concepts for tourism, education and social enterprises • Build responsible creative workflows for agencies, startups and in-house marketing teams
Course outline and learning expectations
This is a tutor-led course. The outline shows what your tutor will cover; teaching materials and examinations are provided directly to enrolled students.
Live online
English (Uganda)
University · General Public
What you will learn
- Distinguish generative models from predictive and discriminative models
- Explain diffusion model training, denoising and inference in clear technical terms
- Design and evaluate text-to-image, image-to-image, inpainting and controlled-generation workflows
- Adapt an image model using LoRA or another parameter-efficient method while managing data and compute constraints
- Build and assess introductory generative audio workflows for speech, music and sound
- Prototype generative video workflows and explain temporal consistency, cost, safety and quality limitations
- Apply copyright, licensing, consent, privacy, provenance, bias and disclosure practices relevant to Uganda and Africa
- Develop, test and present an AI-assisted marketing content application for a small Ugandan business
Modules
- 1
Generative media foundations
Differentiate generative and predictive modelling and map multimodal workflows
Generative versus predictive modelsDiscriminative and generative objectivesLatent representationsTransformers, GANs, VAEs and diffusion overviewModel selection and evaluation - 2
Diffusion models for image generation
Explain how text-to-image diffusion works from noise to image
Forward and reverse diffusionDenoising networksLatent diffusionText conditioning and guidanceSchedulers and samplingQuality, diversity and common failure modes - 3
Image generation and model adaptation
Build and evaluate image workflows and perform lightweight fine-tuning
Prompt designNegative prompts and generation parametersImage-to-imageInpainting and outpaintingControl methodsDataset preparationLoRA and parameter-efficient fine-tuningEvaluation and documentation - 4
Generative audio systems
Create and assess introductory music, speech and sound-generation workflows
Audio representationsText-to-speech and voice synthesis conceptsMusic and sound generationVoice consent and impersonation risksAudio editing and quality checksAfrican-language support and limitations - 5
Generative video systems
Prototype short video workflows while accounting for technical limitations
Text-to-video and image-to-videoShot and storyboard planningMotion and temporal consistencyLip-sync and voice integrationEditing and exportCompute, cost and current limitations - 6
Responsible AI, copyright and data protection
Apply ethical, legal and governance controls to generative media projects
Uganda copyright and licensing basicsConsent and personal dataDeepfakes and disclosureBias and cultural representationDataset provenanceWatermarking and content credentialsPlatform safety policies - 7
Africa-focused business applications
Design practical creative workflows for Uganda, East Africa and African markets
Small-business marketing contentTourism and hospitality promotionE-commerce product visualsEducation and public-information mediaDesign mockupsLocalization and accessibilityHuman review and approval - 8
Capstone application
Build and present an AI-assisted marketing content generator for a small Ugandan business
Problem definitionUser and brand requirementsImage and caption workflowResponsible-use checklistPrototype implementationTesting with realistic promptsEvaluation and iterationDemo and documentation
Before you enroll
- Successful completion of Course:Deep Learning with TensorFlow/PyTorch
- Working knowledge of Python and machine learning
- Understanding of neural networks and model training concepts
- Basic experience using Jupyter Notebook or Google Colab
- Ability to read technical documentation in English
What you need
- Computer with at least 8 GB RAM
- Reliable internet connection
- Modern web browser
- Python 3.10 or later
- Jupyter Notebook or Google Colab
- Git and GitHub
- PyTorch
- Hugging Face Diffusers
- FFmpeg
- Audacity or similar audio editor
- Access to a GPU-enabled cloud notebook or compatible local GPU
- Access to browser-based image, audio and video generation tools
Frequently asked questions
More in Machine Learning

Introduction to Computer Vision

Speech Recognition & Voice AI for Local Languages

Big Data Fundamentals for AI

MLOps & AI Deployment at Scale

Ensemble Learning & Advanced Model Techniques

Version Control & Collaborative Coding with Git & GitHub
More for University · General Public

Dart Programming for Beginners in Uganda & East Africa

Advanced C Systems Programming Course – Africa

Advanced C# Programming in Uganda: Async, Generics & Performance

Advanced C++ Course Uganda – Performance and Concurrency

Advanced Computer Vision & Image Recognition

Advanced Dart Programming & Concurrency for Africa
Quick Actions
Need Help?
Have questions about this program? Our team is here to help!