Governments publish product recalls for free. That sounds like there's no problem to solve â until you actually try to use the data.
I spent June and July building RecallStream: one normalized feed for 32 official recall sources across 20+ countries (openFDA, CPSC, NHTSA, EU Safety Gate, UK OPSS/MHRA/FSA, RappelConso, Health Canada, and more). Dashboard, watchlists, email/Telegram alerts, webhooks, REST API. Solo, bootstrapped, live in production.
Why anyone cares: if you import, distribute, or sell physical products, "did anything in my catalogue just get recalled?" is a daily manual chore across a dozen agency websites. Miss a Class I recall and it costs more than any subscription. The tools that solve it properly are sold with a demo call attached; the free government feeds each stop at one country.
What actually turned out to be hard â and it wasn't the API:
1. Every agency has its own idea of what a recall is. Different schemas, different severity scales, some without dates, some without a product identifier at all. Normalization ended up being ~16 materialized views of SQL rather than TypeScript mappers â I deleted about 4,800 lines of mapper code when I moved the mapping into the database.
2. The same recall gets published several times. One product, three agencies, three notice numbers, weeks apart. Naive dedup gives customers three alerts for one event. Clustering them into a single incident is the part I'd call the actual product.
3. "Public data" is not the same as "reachable data." Several government portals sit behind bot protection that blocks datacenter IPs. Some feeds only work from a residential/ISP address. A few "APIs" are HTML pages with a table. One national feed changed its endpoint mid-project and just started returning nothing.
4. Non-English sources. A recall in French or Polish is useless to a US importer's keyword watchlist unless you classify and translate it before matching, which means an enrichment pass on every record before an alert can fire.
Business side: subscriptions, self-serve, no sales call. Free plan (1 watchlist, US core sources, no card), then $49 / $199 / $499. Paddle as merchant of record so I don't touch VAT. Billing went live this week â revenue is $0 as I write this, so treat everything above as engineering notes, not a success story.
Where I'd like feedback:
- Pricing: is $49 the right first paid step for a solo marketplace seller, or should the entry tier be cheaper and thinner?
- Positioning: I've been selling "recall monitoring" to compliance teams, but the API gets more interest from devs. Two audiences, one landing page â split it or pick one?
Happy to go deeper on the ingestion architecture if that's useful to anyone doing multi-source scraping.
I like that you describe the difficult part as turning many inconsistent public sources into a single reliable incident rather than treating data collection as the product.
I'll be interested to see which use case customers consistently pay for first. That pattern will probably reveal whether the business is fundamentally about compliance monitoring, operational risk, or developer infrastructure.