Big Data Fundamentals for AI

Duration
8 weeks
Investment
UGX 750,000
Certificate
Included
Teaching
Live online
Program Introduction
Big Data Fundamentals for AI builds practical judgement and hands-on skills for processing tabular datasets that no longer fit comfortably in Excel or a single pandas workflow. Learners compare when ordinary tools are still sufficient, then use Dask, PySpark and Parquet to partition, transform, aggregate and store larger datasets. Examples and decisions are grounded in Uganda, East Africa and Africa, including public statistics, trade, agriculture, finance and service-delivery data, with attention to limited bandwidth, computing cost, data protection and responsible use.
Key Features & Benefits
• Decision framework for when big-data tools are justified • Hands-on Dask larger-than-memory workflows • PySpark DataFrame fundamentals and distributed-computing concepts • Parquet versus CSV storage and performance comparisons • Uganda, East Africa and Africa public-data examples • Cost-aware workflows for laptops, local servers and cloud notebooks • Data quality, privacy and governance practices • Capstone using a large public dataset
Real-World Applications
• Prepare large datasets for machine learning and AI workflows • Process UBOS census, household-survey or trade datasets • Analyse East African socioeconomic and trade data • Aggregate anonymised telecom, mobile-money or digital-service records • Clean logistics, agriculture, health and public-service datasets • Convert large CSV archives into partitioned Parquet datasets • Build scalable data-preparation and ETL prototypes for African organisations
Course outline and learning expectations
This is a tutor-led course. The outline shows what your tutor will cover; teaching materials and examinations are provided directly to enrolled students.
Live online
English (Uganda)
Beginner
University · General Public
What you will learn
- Explain big data using practical workload, scale and infrastructure considerations without treating every large file as a big-data problem
- Choose appropriately among Excel, pandas, SQL, Dask and Spark for a given dataset and task
- Identify pandas memory bottlenecks and apply basic optimisation before moving to distributed tools
- Explain partitions, lazy evaluation, task scheduling, shuffles and scaling up versus scaling out
- Build a Dask DataFrame workflow for a dataset that is larger than available memory
- Read, transform, join and aggregate structured data with PySpark DataFrames
- Compare CSV and Parquet, then create an efficient partitioned Parquet dataset
- Estimate important compute, storage and data-transfer trade-offs for Uganda and African operating environments
- Apply data minimisation, anonymisation, access control and responsible public-data practices
- Complete and present a capstone that processes a large public dataset efficiently and documents costs, assumptions and limitations
Modules
- 1
Big Data Decisions and African Context
Define big data practically and decide when specialised tools are justified
Volume, velocity, variety and veracityBig data versus simply large filesWorkload and latency requirementsExcel, pandas, SQL, Dask and Spark decision frameworkUganda, East Africa and Africa use casesAvoiding unnecessary complexity - 2
Pandas at Scale and Memory Limits
Identify why single-machine pandas workflows fail or become inefficient
RAM and dataset sizeData types and memory usageReading data in chunksProfiling bottlenecksVectorisation and efficient filteringWhen optimisation is enoughWhen to move beyond pandas - 3
Distributed Computing Foundations
Understand how large jobs are divided, scheduled and executed across resources
PartitionsTasks and directed acyclic graphsLazy evaluationWorkers and schedulersData localityShufflesFault-tolerance conceptsScaling up versus scaling out - 4
Larger-than-Memory Processing with Dask
Build partitioned data workflows that extend familiar pandas patterns
Dask DataFrameReading multiple filesPartitions and divisionsLazy transformationsAggregationsPersist and computeDiagnostics dashboardMemory-aware processing - 5
Apache Spark and PySpark Basics
Use PySpark DataFrames for distributed structured-data processing
Spark architecture overviewSparkSessionSchemas and data typesReading CSV and Parquetselect, filter and withColumngroupBy and aggregationsJoinsLazy execution and actionsAvoiding collect on large datasets - 6
Efficient Storage with Parquet
Compare storage formats and organise datasets for efficient analytics
CSV limitationsColumnar storageParquet filesCompression and encoding conceptsSchema and data typesPartitioned datasetsPredicate and column pruning conceptsConverting CSV to ParquetFile-size trade-offs - 7
Cost, Infrastructure and Responsible Data Practice
Design realistic workflows for Uganda, East Africa and Africa
Bandwidth and data-transfer costsLaptop, local-server and cloud choicesCompute and storage budgetingPower and connectivity constraintsOpen-source toolingData minimisationAnonymisation and access controlUganda data-protection responsibilitiesDocumenting limitations - 8
Capstone — Large Public Dataset Processing
Process and document a large public dataset efficiently using Dask or PySpark
Select a suitable UBOS, EAC or other African public datasetDefine the analytical questionInspect data quality and licensingBuild an ingestion and cleaning pipelineConvert or write ParquetAggregate and validate resultsMeasure runtime and resource useCompare with a pandas baseline where feasiblePresent findings, costs and limitations
Before you enroll
- Successful completion of Course Introduction to Cloud Computing & APIs
- Successful completion of Course 9
- Confident Python programming
- Working knowledge of pandas and tabular data analysis
- Basic understanding of databases, files and machine learning workflows
What you need
- Laptop or desktop computer
- Reliable internet connection
- Python 3
- Jupyter Notebook or Google Colab
- pandas
- Dask
- PySpark
- Apache Arrow or pyarrow
- Java runtime for local PySpark
- Modern web browser
Frequently asked questions
More in Machine Learning

Introduction to Computer Vision

Speech Recognition & Voice AI for Local Languages

MLOps & AI Deployment at Scale

Ensemble Learning & Advanced Model Techniques

Version Control & Collaborative Coding with Git & GitHub

Unsupervised Learning & Clustering
More for University · General Public

Dart Programming for Beginners in Uganda & East Africa

Advanced C Systems Programming Course – Africa

Advanced C# Programming in Uganda: Async, Generics & Performance

Advanced C++ Course Uganda – Performance and Concurrency

Advanced Computer Vision & Image Recognition

Advanced Dart Programming & Concurrency for Africa
Quick Actions
Need Help?
Have questions about this program? Our team is here to help!