I thought I had plenty of runway… then my daily cron job hit rate limits at ~100 locations. Much sooner than anticipated.
The Weather Recap takes a daily snapshot for each pinned location so users can compare yesterday’s weather against what the forecast said would happen. The original implementation was intentionally simple: a single Vercel cron job that ran once per day, looped through all stored locations, called Open-Meteo, and wrote the results to Redis. I assumed this would be fine for a long time — maybe up to a thousand daily snapshots. In reality, it started failing at around 100 locations.
I realized something was wrong when I opened the app one morning and saw that about half my locations had a snapshot dated “today,” while the other half were still stuck on “yesterday.” Nothing had crashed. There were no errors visible to users. But accuracy calculations, overview tables, and the landing screen no longer agreed with each other. It was a partial failure .
The root cause turned out to be too many concurrent upstream requests causing 429 rate limits, no backpressure or retry logic, and no separation between what needed to be snapshotted and how fast that work should happen. On top of that, timezone edge cases meant “yesterday” wasn’t consistently defined across locations.
Instead of immediately upgrading infrastructure, I refactored the system into two explicit phases. First, a daily dispatch step collects only pinned locations from active users, deduplicates them by rounded lat/lon, and enqueues them into a Redis queue for the day. Then, a worker process drains that queue in small batches with limited concurrency, retries transient failures, and persists progress so one failure doesn’t poison the entire run. “Yesterday” is now computed in the location’s local timezone rather than UTC.
The result is much more boring. If Open-Meteo throttles, the job slows down instead of breaking. If a batch fails, the rest still completes. I can tune batch size and concurrency without redeploying. And I don’t pay for more infrastructure until I actually need it, or so I hope .
I’m not completely out of the woods yet. This morning I woke up to a full snapshot failure, which forced one more round of fixes. I’m cautiously optimistic it’s solved now, but I’ll keep monitoring as daily snapshot volume grows — this may still top out around a few hundred locations. The upside is that accuracy gaps self-heal over time, so missing a couple days isn’t catastrophic.
If anyone else has run into cron + rate-limit + timezone issues like this, I’d love to hear how you handled it.