I hit a small trust issue while building a country-stat explorer: the word “latest” can hide a multi-year gap.
The World Bank WDI global forest-area series (AG.LND.FRST.ZS) records 33.01% of land area in 1992 and 31.10% in 2023. In the snapshot I used, 2023 is still the latest global observation. I kept missing years blank instead of interpolating them.
The reproducible CSV, chart, and fetch script are here: https://github.com/ChessShark1000/world-bank-forest-area-wdi
I’m building Global Data Tracker around the same idea: keep observation years and sources close to country metrics, and mark modeled counters as estimates. https://globaldatatracker.com/
For founders working with public datasets, how do you surface “last updated” without implying the number is live?
Appreciate the honesty here, most people only share the wins.
Solid lesson. Which channel has worked best for you so far?
Two dates is the right default and almost nobody does it. In SEO data the same confusion shows up constantly: a rank tracker shows "current position 4" but the crawl ran three days ago and the page moved to position 8 yesterday. Google Search Console itself reports with a 2-3 day lag, and people treat the latest available date as live.
We built UtilitySEO to scan sites and one design decision we got right early was showing "scanned at [timestamp]" next to every finding, not "current issues." A finding from a scan that ran before the fix deployed is not current — it is historical. Same principle as your period-date vs filing-date split.
The interpolation point bites hardest. A smooth line between two known data points feels like signal and reads like confidence, but it is a guess carrying a chart's authority.
I deal with a milder version of this with SEC 13F filings: fund holdings as of the end of a quarter, filed up to 45 days later, and people read them as "what fund X owns now". What helped:
Your "leave missing years blank rather than interpolate" call is the right one too. A gap reads as honest; a smooth line reads as live.