Program Curriculum

🔹 Module 1: Foundations of Big Data Engineering

  • Introduction to Big Data: Volume, Velocity, Variety
  • Evolution from RDBMS to NoSQL
  • Data Engineering vs Data Science
  • Data Ecosystem Overview
  • Linux and Shell Scripting Basics

🔹 Module 2: Programming for Data Engineering

  • Python for Data Engineering (Data types, Functions, File I/O, Pandas, NumPy)
  • API Calls, JSON, XML Handling
  • SQL for Data Analysis (Joins, Subqueries, Window Functions, Performance Tuning)

🔹 Module 3: Hadoop Ecosystem

  • Hadoop Architecture & HDFS
  • MapReduce Programming
  • YARN and Resource Management
  • Hive for Data Warehousing & Pig for Data Flow Scripts
  • HBase for NoSQL Storage
  • Sqoop and Flume for Data Ingestion

🔹 Module 4: Apache Spark

  • Spark Architecture and RDDs
  • DataFrames and Spark SQL
  • Spark Streaming for Real-Time Processing
  • PySpark Programming
  • Performance Tuning & Cluster Deployment

🔹 Module 5: Apache Kafka and Real-Time Data Pipelines

  • Kafka Architecture and Concepts (Producers, Consumers, Topics, Brokers)
  • Kafka Streams API
  • Kafka Integration with Spark & Flink
  • Real-time Event Streaming Use Cases

🔹 Module 6: NoSQL Databases

  • Introduction to NoSQL (CAP Theorem)
  • MongoDB: CRUD Operations, Aggregation Framework, Indexing
  • Cassandra: Data Modeling, CQL and Architecture

🔹 Module 7: Data Warehousing & ETL

  • Data Warehouse Architecture (OLAP vs OLTP)
  • Star & Snowflake Schema
  • ETL Tools: Apache NiFi, Talend (overview)
  • Cloud DWs: Snowflake, BigQuery, Redshift (Intro)

🔹 Module 8: Cloud & DevOps for Data Engineering

  • Cloud Platforms: AWS, GCP, Azure (basic services)
  • Infrastructure as Code: Terraform Basics
  • CI/CD for Data Pipelines
  • Docker & Kubernetes for Data Workloads

🔹 Module 9: Data Governance, Quality, and Security

  • Data Lineage and Metadata Management
  • Data Quality Checks
  • GDPR, HIPAA Compliance
  • Encryption and Access Controls

🔹 Module 10: Capstone Project

  • Real-world project to build an end-to-end big data pipeline from ingestion to deployment on the cloud (AWS/GCP).

🛠️ Tools & Technologies Covered

Languages:

Python, SQL, Shell

Big Data Tools:

Hadoop, Spark, Kafka, Hive, Sqoop, Flume

NoSQL:

MongoDB, Cassandra

Cloud:

AWS/GCP/Azure Basics

Workflow & Orchestration:

Apache Airflow, NiFi

DevOps:

Docker, Kubernetes, Git, Jenkins