Semantic & Metrics Layer
Finance and growth report different revenue for the same month, because each team rebuilt “the” customer table its own way. Modeling is not academic — it is why your numbers do or don't reconcile.
This stratum covers dimensional modeling end to end: star schemas and the Kimball bus matrix, slowly changing dimensions (SCD Type 1 and 2), Data Vault 2.0, One Big Table and wide-table designs, conformed dimensions, and how to choose a modeling style for a real warehouse.
What you'll learn
- Design star schemas, fact and dimension tables, and a conformed bus matrix
- Implement slowly changing dimensions (SCD Type 1 and 2) correctly
- Know when Data Vault 2.0, One Big Table, or classic Kimball fits the problem
- Make different teams' metrics reconcile through conformed dimensions
Tracks & courses
Full navigation is in the sidebar. Here's what each track gives you and the courses inside it.
Dimensional Modeling Foundations
The shared vocabulary every metrics layer assumes — grain, facts, dimensions, SCDs, conformed dimensions, and the bus matrix.
Dimensional Modeling Fundamentals
Grain, facts, dimensions, measures, keys — the engine-agnostic vocabulary every metrics layer assumes. Build the words you'll use for the rest of Strata 7.
10 ch · 2h 27m
1 freeSlowly Changing Dimensions
When a customer changes their region, every historical fact silently lies — unless you've modeled the change. SCD types 1/2/3/6, effective-dating, bitemporal, and the production anti-patterns that bite teams in their first year.
8 ch · 1h 44m
1 freeConformed Dimensions & the Bus Matrix
Kimball's organizational technology: how to keep 'customer' meaning one thing across marketing, finance, and product — and how the bus matrix turns conformance from a Slack-thread into a treaty.
6 ch · 1h 16m
1 freeCumulative Table Design
The pattern behind dim_all_users: full-outer-join yesterday to today, coalesce, and carry all of history in one row. Complex types (struct, array, map), the compactness-vs-usability tradeoff, temporal cardinality explosions, why run-length encoding is the reason Parquet won, and how to collapse cumulative history into an SCD Type 2 with window functions and a hand-rolled incremental merge.
4 ch · 1h 42m
1 freeFact Data Modeling
Facts are the biggest data you'll ever touch: immutable events at 10-100x the volume of your dimensions. What makes a fact atomic, why raw logs aren't fact data, when denormalization is the fix (not the bug), deduplication at trillion-row scale, and the blurry line where aggregated facts become dimensions.
6 ch · 2h
1 freeDatelists and Reduced Facts
The compression patterns behind Facebook-scale activity analytics: cumulate user activity into date arrays, pack 30 days of history into one integer with bit math, and reduce daily fact volume 30x with value arrays, turning decade-long analyses from weeks of pipeline time into hours.
4 ch · 1h 48m
1 freeModeling Alternatives: Data Vault & OBT
When stars aren't the answer. Data Vault for audit-heavy, source-volatile environments; OBT for columnar engines with low-cardinality joins; and how to choose between styles.
Data Vault 2.0
The audit-first, source-volatile-first alternative to Kimball. Hubs, links, satellites, hash keys, and PIT tables — and the honest discussion of when Vault is the right tool versus when it's cosplay.
10 ch · 2h 8m
1 freeOne Big Table & Wide-Table Design
The columnar-engine-native modeling style. When 200-column denormalized tables go from heresy to the right answer — and where OBT's quiet costs (update amplification, fan-out, governance) live.
8 ch · 1h 28m
1 freeChoosing a Modeling Style
The decision framework that spans Kimball, Vault, and OBT. By the end, the choice is defensible — you can name the inputs, predict the costs, and explain why one approach wins for your specific team, sources, engine, and change rate.
6 ch · 1h 4m
1 freeMetrics Layer Foundations
The engine- and tool-agnostic concepts behind every metrics store. Why metric drift happens, what a semantic layer actually is, the four primitives (measure, metric, entity, dimension), and how a metrics engine compiles a request into SQL over a semantic graph.
Why a Metrics Layer Exists
Metric drift is an org failure with a technical fix. This course makes the case for a metrics layer: the three-definitions-of-MAU war story, the headless BI thesis, what a metrics layer is NOT, and why adoption fails at month 6 rather than at install.
7 ch · 1h 28m
1 freeMeasures, Metrics, Entities, Dimensions
The four primitives every metrics store is built on, made precise. Why a measure is not a metric, how entities are the join graph you already have, the ratio-metric order-of-aggregation trap, cumulative metrics and the time spine, and the metric DAG you never drew.
9 ch · 1h 53m
1 freeQuery Planning Over a Semantic Graph
How a metrics engine turns 'total_revenue by region' into a multi-join SQL plan. Join path resolution, the fan-out join trap, symmetric aggregates, aggregation pushdown, time grain resolution, slice-and-dice semantics, and why you cache a tuple, not a metric.
8 ch · 1h 41m
1 freeAnalytical Patterns
The repeatable analyses behind almost every pipeline: aggregation, cumulation, and window patterns, growth accounting, retention J-curves, funnel construction, and the GROUPING SETS family that powers pre-aggregated dashboards.
Related topics
Start Semantic & Metrics Layer free
The first chapters of every course are free to read — no account needed.