PII & Data Governance
A “delete my data” request arrives and you realize one user's PII is copied across a dozen tables, three Parquet snapshots, and a Kafka topic with seven-day retention. Right-to-erasure is an engineering problem, not a legal checkbox.
This specialization is a lens across the whole stack: classifying PII, masking and tokenization at ingestion, de-identification, and what GDPR, CCPA, and HIPAA actually require of the systems you build.
What you'll learn
- Classify PII and design masking or tokenization at ingestion
- Map GDPR, CCPA, and HIPAA requirements to concrete data-engineering controls
- Engineer right-to-erasure across tables, file formats, and streams
- Apply de-identification techniques and reason about re-identification risk
Tracks & courses
Full navigation is in the sidebar. Here's what each track gives you and the courses inside it.
PII Foundations
From identifying personal data to navigating GDPR, CCPA, and HIPAA — the engineer's guide to data privacy.
PII Fundamentals
Identify, classify, and manage PII across your data stack — from direct identifiers to quasi-identifiers hiding in plain sight.
7 ch
1 freeRegulations & Compliance
GDPR, CCPA/CPRA, and HIPAA from an engineer's perspective: lawful basis, right to erasure, and pipeline obligations.
8 ch
1 freeMasking & Anonymization
Turn classified PII into protected data: hashing, tokenization, format-preserving encryption, dynamic masking, and the formal anonymization models that give mathematical guarantees instead of false comfort.
Masking Techniques
When to use hashing vs tokenization vs format-preserving encryption, and how to implement dynamic masking in modern data platforms.
7 ch · 2h 40m
1 freeAnonymization Deep Dive
k-anonymity, l-diversity, t-closeness, and differential privacy: formal models that give re-identification guarantees instead of just obscuring data.
7 ch · 2h 45m
1 freePII at Scale with Spark & Iceberg
Take PII protection from a single CSV to the whole warehouse: detect personal data across terabytes with PySpark and Presidio, then govern, retain, and erase it natively in Apache Iceberg. Ends in a full end-to-end pipeline.
PySpark PII Detection
Detect PII at scale with PySpark regex and Microsoft Presidio: patterns, UDF performance, confidence scoring, and a real scan job.
7 ch · 2h 50m
1 freeIceberg Data Governance for PII
Tag PII columns, execute GDPR deletes with row-level deletes, and prove erasure with time travel, all native to Iceberg.
7 ch · 2h 50m
1 freeIceberg PII Lifecycle
Retention policies, partition and snapshot expiration, orphan-file cleanup, and an automated, verifiable PII purge pipeline in Iceberg.
7 ch · 2h 50m
1 freeCapstone: End-to-End PII Pipeline
Ship the full pipeline: raw -> detect -> mask -> govern -> store in Iceberg, then handle a complete GDPR erasure cycle end to end.
8 ch · 3h 35m
1 freeRelated topics
Start PII & Data Governance free
The first chapters of every course are free to read — no account needed.