Workshop
digicode: PYTH04
Python Data Wrangling & Analytics Engineering
Course facts
Download as PDF- Optimizing Data for AI Applications
- Consolidating data from files, databases, and APIs in an optimal way
- Learning about Parquet and Apache Arrow as modern, column-oriented data formats
- Ensuring data quality with schemas, rules, and data contracts
- Validating transformations with automated tests
- Applying best practices in AI-powered applications
Day 1: Understanding Data and Importing It Reliably
- The Role of Data Wrangling and Analytics Engineering
- Building a Simple “Raw–Clean–Curated” Structure
- Pandas 3.x Refresh
- Importing Data from CSV, Excel, JSON, SQL, and REST APIs
- Parquet and Apache Arrow
- Data types, schemas, and metadata
- Data profiling and data dictionaries
- Missing values, duplicates, invalid categories, and outliers
- Timestamps, time zones, and calendar data
- AI-assisted interpretation of data profiles and formulation of quality questions
Day 2: Transform, Validate, and Scale
- Filtering, Grouping, Aggregating, and Reshaping
- Linking Tables and Systematically Checking Joins
- Vectorization Instead of Slow Python Loops
- Transformations with DuckDB and SQL
- Data Validation with Pandera
- Schema, Range, Uniqueness, and Relationship Checks
- Data Contracts and Guardrails
- Automated Tests for Transformations
- Comparison Lab: Pandas, Polars, and DuckDB
- AI for Generating Transformation Designs, Edge Cases, and Test Cases
- Deterministically Validating All AI Results
Day 3: Reproducible Analytics Pipeline
- Migrating notebook code into functions and modules
- Configuration and parameters
- Logging and traceable error handling
- Idempotent and repeatable processing
- Data provenance and transformation log
- Separating raw, clean, and curated data
- Storing results in Parquet
- Query validated data with DuckDB
- Generate synthetic sample data using AI and check it for data privacy and statistical quality
Open Data, Provenance & Reproducibility with a Final Project:
AI becomes particularly powerful when diverse data sets are linked, and open data gives you the opportunity to access and use a wide variety of data. A sensible best-practice approach to leveraging these data treasures for your own business purposes. In this module, you’ll learn:
- How to document the provenance and license of a dataset. Open data is not automatically free of:
- personal data
- usage restrictions
- quality issues
- bias
- Log the data source, retrieval time, and transformations
- Use open formats such as Parquet and Arrow
- Record the versions of the libraries used
- Create reproducible environments and data pipelines
- Do not automatically assume that AI-generated synthetic data is anonymous or unproblematic
- Workshop involving exercises and practical applications
- Participants will complete the exercises on their own laptops
- Before the course begins, participants will receive detailed information on how to prepare and install the relevant open-source Python software
This course is aimed at data analysts, business analysts, data engineers, BI specialists and database specialists.
This course is not a general introductory course on Python; corporate data in LLM and dashboard development are not covered in this course.
A basic understanding of data analysis and advanced Python skills, as taught in the following course, are required:
You have an account on ChatGPT, Claude, Copilot, Gemini, or Perplexity. The free versions are sufficient.
If you’re attending in person, we recommend bringing your own laptop if possible. That way, you can install the necessary tools from the class yourself and take the data you’ve worked on home with you.