I’ve been focusing mostly on scrapers and data normalization, and one thing that turned out to be more complex than expected is standardizing job data across companies.
Things like location (city vs city + state vs remote vs hybrid), salary formats (ranges, currencies, missing data), tech stacks embedded in descriptions are really different from one to the other.
Even companies using the same ATS (Greenhouse, Lever, etc.) structure things quite differently.
Feels like a big part of the value here will actually come from cleaning and structuring the data, not just aggregating it.
Still early, but starting to see where the real complexity (and opportunity) is.