Circling back to my last post - I’ve spent a fair amount of time tightening that system and I wanted to close the loop on what actually stuck.
The final fix wasn’t more infrastructure, but more structure. I split the snapshot process into two explicit phases: a daily dispatch step that deterministically collects and deduplicates pinned locations, and a controlled worker phase that processes those locations with bounded concurrency, retries, and progress tracking. Instead of “run everything now and hope,” the job now behaves more like a queue that drains safely under pressure.
The other major fix was being strict about time. “Yesterday” is now computed in each location’s local timezone rather than implicitly in UTC, which eliminated a whole class of subtle partial failures where nothing crashed but data silently disagreed across views.
The system is now intentionally boring. If upstream APIs throttle, the job slows down instead of failing halfway through. If a single batch fails, the rest still completes. I can tune batch size and concurrency without redeploying. Most importantly, I’ve stopped waking up to mismatched dates and silent data drift.
This still isn’t infinite scale, but it’s honest scale. I’m comfortable with a few hundred daily snapshots on the current setup, and I won’t pay for more infrastructure until usage actually demands it. The main lesson for me was that “simple” cron jobs tend to hide complexity until they don’t.