Topic-wise study material for weeks 1–12 of the 12-week plan.
SELECT/WHERE/GROUP BY, JOINs, CTEs, subqueries, CASE, window functions, indexes, optimization, transactions, views, e-commerce project, NoSQL databases, SQL vs NoSQL.
Functions, OOP basics, exceptions, files, JSON/CSV, APIs, logging, venv, pandas, SQLAlchemy, psycopg2, API→Postgres project.
ETL vs ELT, batch processing, incremental loading, CDC, validation, dedup, idempotency, fact/dim tables, star schema, SCD Type 1/2.
IAM, S3 (main focus + data lake design), Glue, Lambda, CloudWatch, Redshift.
Core Linux commands, Git workflow (branch/merge/rebase), Docker concepts, Dockerfile, running Postgres+Python+Airflow locally with docker-compose.
DAG, Task, Operator, Scheduler, Dependency, Executors, Retry, XCom, Backfill, Monitoring, plus dbt sources, models, tests, macros, snapshots, docs & lineage.
Spark architecture, DataFrames, Spark SQL, transformations vs actions, partitions, shuffle, joins, broadcast joins, caching, window functions.
Producer, Consumer, Topic, Partition, Offset, Consumer Group, Replication — Application → Kafka → Spark → PostgreSQL project.
The AI-powered data platform, as a Mermaid diagram — sources → S3 → Airflow → PySpark/dbt → warehouse, forking into analytics and a RAG → LLM assistant.
What it is, when to use it, architecture, notebooks, clusters, Delta Lake, Autoloader, Databricks SQL, Jobs/Workflows, Unity Catalog, MLflow.