ProgrammingUniversity · General PublicBig Data Fundamentals for AI
Machine Learning
University · General Public

Big Data Fundamentals for AI

Big Data Fundamentals for AI

Duration

8 weeks

Investment

UGX 750,000

Certificate

Included

Teaching

Live online

Program Introduction

Big Data Fundamentals for AI builds practical judgement and hands-on skills for processing tabular datasets that no longer fit comfortably in Excel or a single pandas workflow. Learners compare when ordinary tools are still sufficient, then use Dask, PySpark and Parquet to partition, transform, aggregate and store larger datasets. Examples and decisions are grounded in Uganda, East Africa and Africa, including public statistics, trade, agriculture, finance and service-delivery data, with attention to limited bandwidth, computing cost, data protection and responsible use.

Key Features & Benefits

• Decision framework for when big-data tools are justified • Hands-on Dask larger-than-memory workflows • PySpark DataFrame fundamentals and distributed-computing concepts • Parquet versus CSV storage and performance comparisons • Uganda, East Africa and Africa public-data examples • Cost-aware workflows for laptops, local servers and cloud notebooks • Data quality, privacy and governance practices • Capstone using a large public dataset

Real-World Applications

• Prepare large datasets for machine learning and AI workflows • Process UBOS census, household-survey or trade datasets • Analyse East African socioeconomic and trade data • Aggregate anonymised telecom, mobile-money or digital-service records • Clean logistics, agriculture, health and public-service datasets • Convert large CSV archives into partitioned Parquet datasets • Build scalable data-preparation and ETL prototypes for African organisations

Course outline and learning expectations

This is a tutor-led course. The outline shows what your tutor will cover; teaching materials and examinations are provided directly to enrolled students.

Big Data FundamentalsBig Data Course UgandaBig Data for AIDask Course UgandaPySpark Course UgandaApache Spark East AfricaParquetDistributed ComputingData Engineering UgandaData Science East AfricaAI Training AfricaLarge Dataset ProcessingUBOS DataAfrican Data Engineering
Teaching format

Live online

Teaching language

English (Uganda)

Course difficulty

Beginner

Intended learners

University · General Public

What you will learn

  • Explain big data using practical workload, scale and infrastructure considerations without treating every large file as a big-data problem
  • Choose appropriately among Excel, pandas, SQL, Dask and Spark for a given dataset and task
  • Identify pandas memory bottlenecks and apply basic optimisation before moving to distributed tools
  • Explain partitions, lazy evaluation, task scheduling, shuffles and scaling up versus scaling out
  • Build a Dask DataFrame workflow for a dataset that is larger than available memory
  • Read, transform, join and aggregate structured data with PySpark DataFrames
  • Compare CSV and Parquet, then create an efficient partitioned Parquet dataset
  • Estimate important compute, storage and data-transfer trade-offs for Uganda and African operating environments
  • Apply data minimisation, anonymisation, access control and responsible public-data practices
  • Complete and present a capstone that processes a large public dataset efficiently and documents costs, assumptions and limitations

Modules

  1. 1

    Big Data Decisions and African Context

    Define big data practically and decide when specialised tools are justified

    Volume, velocity, variety and veracityBig data versus simply large filesWorkload and latency requirementsExcel, pandas, SQL, Dask and Spark decision frameworkUganda, East Africa and Africa use casesAvoiding unnecessary complexity
  2. 2

    Pandas at Scale and Memory Limits

    Identify why single-machine pandas workflows fail or become inefficient

    RAM and dataset sizeData types and memory usageReading data in chunksProfiling bottlenecksVectorisation and efficient filteringWhen optimisation is enoughWhen to move beyond pandas
  3. 3

    Distributed Computing Foundations

    Understand how large jobs are divided, scheduled and executed across resources

    PartitionsTasks and directed acyclic graphsLazy evaluationWorkers and schedulersData localityShufflesFault-tolerance conceptsScaling up versus scaling out
  4. 4

    Larger-than-Memory Processing with Dask

    Build partitioned data workflows that extend familiar pandas patterns

    Dask DataFrameReading multiple filesPartitions and divisionsLazy transformationsAggregationsPersist and computeDiagnostics dashboardMemory-aware processing
  5. 5

    Apache Spark and PySpark Basics

    Use PySpark DataFrames for distributed structured-data processing

    Spark architecture overviewSparkSessionSchemas and data typesReading CSV and Parquetselect, filter and withColumngroupBy and aggregationsJoinsLazy execution and actionsAvoiding collect on large datasets
  6. 6

    Efficient Storage with Parquet

    Compare storage formats and organise datasets for efficient analytics

    CSV limitationsColumnar storageParquet filesCompression and encoding conceptsSchema and data typesPartitioned datasetsPredicate and column pruning conceptsConverting CSV to ParquetFile-size trade-offs
  7. 7

    Cost, Infrastructure and Responsible Data Practice

    Design realistic workflows for Uganda, East Africa and Africa

    Bandwidth and data-transfer costsLaptop, local-server and cloud choicesCompute and storage budgetingPower and connectivity constraintsOpen-source toolingData minimisationAnonymisation and access controlUganda data-protection responsibilitiesDocumenting limitations
  8. 8

    Capstone — Large Public Dataset Processing

    Process and document a large public dataset efficiently using Dask or PySpark

    Select a suitable UBOS, EAC or other African public datasetDefine the analytical questionInspect data quality and licensingBuild an ingestion and cleaning pipelineConvert or write ParquetAggregate and validate resultsMeasure runtime and resource useCompare with a pandas baseline where feasiblePresent findings, costs and limitations

Before you enroll

  • Successful completion of Course Introduction to Cloud Computing & APIs
  • Successful completion of Course 9
  • Confident Python programming
  • Working knowledge of pandas and tabular data analysis
  • Basic understanding of databases, files and machine learning workflows

What you need

  • Laptop or desktop computer
  • Reliable internet connection
  • Python 3
  • Jupyter Notebook or Google Colab
  • pandas
  • Dask
  • PySpark
  • Apache Arrow or pyarrow
  • Java runtime for local PySpark
  • Modern web browser

Frequently asked questions

More in Machine Learning

Introduction to Computer Vision
University · General Public

Introduction to Computer Vision

6 weeksUGX 600,000
Speech Recognition & Voice AI for Local Languages
University · General Public

Speech Recognition & Voice AI for Local Languages

8 weeksUGX 700,000
MLOps & AI Deployment at Scale
University

MLOps & AI Deployment at Scale

6 weeksUGX 650,000
Ensemble Learning & Advanced Model Techniques
University

Ensemble Learning & Advanced Model Techniques

6 weeksUGX 550,000
Version Control & Collaborative Coding with Git & GitHub
Upper Secondary · University · General Public

Version Control & Collaborative Coding with Git & GitHub

6 weeksUGX 450,000
Unsupervised Learning & Clustering
University · General Public

Unsupervised Learning & Clustering

6 weeksUGX 600,000

More for University · General Public

 Dart Programming for Beginners in Uganda & East Africa
Lower Secondary · Upper Secondary · University · General Public

Dart Programming for Beginners in Uganda & East Africa

4 weeksUGX 450,000
Advanced C Systems Programming Course – Africa
Upper Secondary · University · General Public

Advanced C Systems Programming Course – Africa

4 weeksUGX 500,000
Advanced C# Programming in Uganda: Async, Generics & Performance
Upper Secondary · University · General Public

Advanced C# Programming in Uganda: Async, Generics & Performance

4 weeksUGX 500,000
Advanced C++ Course Uganda – Performance and Concurrency
Upper Secondary · University · General Public

Advanced C++ Course Uganda – Performance and Concurrency

4 weeksUGX 500,000
Advanced Computer Vision & Image Recognition
Upper Secondary · University · General Public

Advanced Computer Vision & Image Recognition

6 weeksUGX 650,000
Advanced Dart Programming & Concurrency for Africa
Lower Secondary · Upper Secondary · University · General Public

Advanced Dart Programming & Concurrency for Africa

4 weeksUGX 500,000

Quick Actions

Enroll Now

Need Help?

Have questions about this program? Our team is here to help!