Reinforcement Learning Fundamentals

Duration
6 weeks
Investment
UGX 650,000
Certificate
Included
Teaching
Live online
Program Introduction
Reinforcement Learning Fundamentals is an advanced practical course for learners in Uganda, East Africa and across Africa who already have the required background from Course 10 and Course 14. The program explains how intelligent agents learn through interaction, rewards and feedback, then guides learners from Markov Decision Processes and tabular Q-learning to the intuition behind Deep Q-Networks. Using Python, Gymnasium and PyTorch, learners build and evaluate agents in controlled simulations, including a retail restocking capstone. The course also examines where reinforcement learning can support routing, resource allocation and pricing research, and where simpler machine learning or optimisation methods are more appropriate. Emphasis is placed on responsible experimentation, clear evaluation and practical limitations rather than exaggerated claims about production-ready automation.
Key Features & Benefits
• Agent-environment and MDP foundations • Hands-on Q-learning implementation • Simple Gymnasium agent project • DQN intuition with PyTorch • Africa-relevant case studies • Responsible RL evaluation and limitations • Retail restocking capstone
Real-World Applications
• Prototype inventory and restocking policies for retail simulations • Explore routing and dispatch decisions in logistics simulations • Test resource allocation strategies for constrained systems • Study dynamic pricing in safe simulated markets • Develop game and control agents for learning and research • Evaluate whether RL is suitable for an African business or public-service problem
Course outline and learning expectations
This is a tutor-led course. The outline shows what your tutor will cover; teaching materials and examinations are provided directly to enrolled students.
Live online
English (Uganda)
Beginner
University · General Public
What you will learn
- Explain the reinforcement learning agent-environment framework and core terminology
- Represent sequential decision problems as Markov Decision Processes
- Implement and tune tabular Q-learning with exploration strategies
- Build and evaluate a simple agent in a Gymnasium environment
- Explain the main components and training logic of Deep Q-Networks
- Compare reinforcement learning with supervised learning, optimisation and rule-based approaches
- Identify practical, ethical and operational limitations of reinforcement learning
- Develop a capstone agent for simulated retail restocking decisions
Modules
- 1
Reinforcement learning foundations
Understand how agents learn through interaction and rewards
Agent-environment loopStates and observationsActionsRewards and returnsPoliciesEpisodesRL compared with supervised learning - 2
Markov Decision Processes
Model sequential decisions with states, actions, transitions and rewards
Markov propertyTransition dynamicsReward functionsDiscount factorValue functionsBellman intuition - 3
Q-learning basics
Implement value-based learning for discrete environments
Q-tablesTemporal-difference updatesLearning rateDiscount factorEpsilon-greedy explorationTraining loopsConvergence intuition - 4
Building a simple RL agent
Create, train and evaluate an agent in a game-style environment
Gymnasium APIReset and step cycleEnvironment wrappersLogging rewardsEvaluation episodesReproducibility and random seeds - 5
Deep Q-Networks intuition
Understand how neural networks extend Q-learning to larger state spaces
Function approximationReplay bufferTarget networkLoss calculationExploration schedulingTraining stabilityPyTorch DQN workflow - 6
Real-world RL use cases
Assess RL applications through Africa-relevant examples and simulations
Routing and dispatchInventory and restockingResource allocationDynamic pricingEnergy and network optimisationSimulation requirements - 7
Limitations and responsible use
Decide when RL is useful, risky or unnecessary
Sample inefficiencyReward designSafety and testingDistribution shiftCompute and data constraintsEthical considerationsSimpler alternatives - 8
Capstone retail restocking agent
Design and evaluate an agent for a simulated small-shop inventory problem
Problem definitionState and action designReward functionDemand simulationBaseline policyTraining and evaluationResults presentation
Before you enroll
- Completion of the course:Supervised Learning in Depth
- Completion of the course:Introduction to Deep Learning & Neural Networks
- Comfort with Python programming
- Basic probability and linear algebra
- Basic machine learning concepts
What you need
- Computer
- Reliable internet connection
- Python 3
- Jupyter Notebook or Google Colab
- Visual Studio Code
- NumPy
- Matplotlib
- Gymnasium
- PyTorch
- Git
Frequently asked questions
More in Machine Learning

Introduction to Computer Vision

Speech Recognition & Voice AI for Local Languages

Big Data Fundamentals for AI

MLOps & AI Deployment at Scale

Ensemble Learning & Advanced Model Techniques

Version Control & Collaborative Coding with Git & GitHub
More for University · General Public

Dart Programming for Beginners in Uganda & East Africa

Advanced C Systems Programming Course – Africa

Advanced C# Programming in Uganda: Async, Generics & Performance

Advanced C++ Course Uganda – Performance and Concurrency

Advanced Computer Vision & Image Recognition

Advanced Dart Programming & Concurrency for Africa
Quick Actions
Need Help?
Have questions about this program? Our team is here to help!