🔐Specialization

PII & Data Governance

A “delete my data” request arrives and you realize one user's PII is copied across a dozen tables, three Parquet snapshots, and a Kafka topic with seven-day retention. Right-to-erasure is an engineering problem, not a legal checkbox.

This specialization is a lens across the whole stack: classifying PII, masking and tokenization at ingestion, de-identification, and what GDPR, CCPA, and HIPAA actually require of the systems you build.

What you'll learn

  • Classify PII and design masking or tokenization at ingestion
  • Map GDPR, CCPA, and HIPAA requirements to concrete data-engineering controls
  • Engineer right-to-erasure across tables, file formats, and streams
  • Apply de-identification techniques and reason about re-identification risk

Tracks & courses

Full navigation is in the sidebar. Here's what each track gives you and the courses inside it.

PII at Scale with Spark & Iceberg

Take PII protection from a single CSV to the whole warehouse: detect personal data across terabytes with PySpark and Presidio, then govern, retain, and erase it natively in Apache Iceberg. Ends in a full end-to-end pipeline.

Related topics

Start PII & Data Governance free

The first chapters of every course are free to read — no account needed.