I spent 6 weeks turning a temporal join debugger into a historical data modeling workbench
I’m a Data Engineer and over the last few years I’ve spent a lot of time working with historized data, SCD2 dimensions, monthly snapshots and temporal reporting.
The original idea was simple:
I kept running into temporal join issues and built a small tool to debug them.
The problem?
Nobody wakes up thinking:
“I need a temporal join debugger today.”
After sharing it and watching how people interacted with it, I realized the actual problem was much broader.
Many of the hardest data issues I’ve seen were not SQL bugs or Spark bugs.
They were historical modeling problems:
Snapshot reproducibility
SCD2 alignment
Late arriving corrections
Historical relationship changes
Event-to-state alignment
Dimension completion
Over the last few weeks I expanded the project into a Historical Data Modeling Workbench.
Current features:
Historical Modeling Advisor
Model Review for SQL, dbt and PySpark
Target Table Validation
Historical Modeling Pattern Catalog
Interactive examples and learning pages
Current challenge:
I’m trying to figure out which problem is painful enough that engineers would actually pay to solve it.
So I’d love to hear from others working in analytics engineering, data warehousing or lakehouse projects:
What historical data problem causes the most pain in your environment?
Project: