Data Engineering
Start from pipelines that stay safe when they fail, and learn to judge by the numbers the engines leave behind — Spark and Hadoop execution plans, Iceberg table snapshots, Flink state and checkpoints, and ClickHouse's storage layout. By the end, when someone reports that the data looks wrong, you can narrow down with numbers which stage went astray.
코스
- Data Pipelines — Fail, rerun, and the result must be the same
- Apache Spark — The answer to a slow job is in the plan and the event log — Measure shuffles, joins and skew in local-mode Spark with the event log
- Apache Hadoop — Stand up and run HDFS and YARN in one pod — Blocks, permissions, quotas and MapReduce by hand on pseudo-distributed HDFS and YARN
- Lakehouse Table Format — Understanding Apache Iceberg Through Its Metadata — Verify snapshots, schema and partition evolution, MERGE and maintenance in the metadata
- Apache Flink — Running Streams on a Real Engine — Verify watermarks, state and checkpoints from the engine's own output
- ClickHouse — A Columnar Analytics Database from the Inside — Read sort keys, parts and merges through the numbers in system tables