The students study the principles and tools of reproducible data engineering with a focus on energy data. They learn to manage code and data with version control (Git), to isolate and reproduce computational environments using virtual environments, containers (Docker), and purely functional package management (Nix), and to persist data in relational databases (SQLite, MySQL). The students study how to automate data pipelines and quality checks with continuous integration and delivery (CI/CD). The methods are applied to messy, real-world energy datasets to build fully reproducible, automatically tested data pipelines.
- verantwortliche Lehrperson: Jonathan Berrisch
ePortfolio: No