
SheetSanitizer
Safely clean messy CSV and Excel files — entirely in your br
When I started working on SheetSanitizer, I thought the goal was simple.
Find messy spreadsheet data.
Fix everything possible.
Done.
The more real spreadsheets I tested, the more I realized that approach was actually dangerous.
Some problems are deterministic.
Trailing spaces.
Invisible Unicode characters.
Email casing.
Those can safely be fixed.
But many problems aren't.
For example:
Is 03/04/2024 March 4 or April 3?
Are two identical rows duplicates, or two legitimate transactions?
Should mixed date formats actually be normalized?
Many tools simply guess.
I decided mine shouldn't.
So instead of trying to automate everything, I built the workflow around one rule:
Never guess when the software cannot know.
Every issue is placed into one of two categories.
Safe Fix
Deterministic issues that can be corrected without changing the meaning of the data.
Review
Anything that requires human judgment.
The software explains the issue, but the decision stays with the user.
That one design decision ended up changing almost every part of the product.
Over the past few months I've been designing the workflow, testing lots of edge cases, refining the detection engine, and using AI coding tools to help implement many of those ideas faster. The product decisions, testing, and iteration have been mine, but AI has definitely been part of the development process.
Ironically, after spending so much time improving detection accuracy, the first meaningful feedback I received wasn't about the engine.
It was about the report screen.
One tester told me the Safe Fix / Review distinction clicked immediately, but after the scan they froze for a few seconds because there was simply too much information on screen.
That reminded me that a reliable engine and an understandable product are two very different things.
So before adding more detectors, I'm focusing on making the workflow easier for someone seeing it for the first time.
If you've built a product where correctness and usability were constantly pulling in different directions, I'd genuinely love to hear how you approached that trade-off.
P.S. If you'd like to try it with a sample spreadsheet or tell me where the UX falls apart, it's here:
About
I kept seeing the same problem: spreadsheets looked clean until a tiny formatting issue, hidden character, or ambiguous date quietly broke an import, report, or analysis. Most spreadsheet cleaners try to fix everything

1 Comment
That tester reaction seems especially useful because it shifts the question from whether the engine can identify the right issues to whether users can actually act on what it finds.
Was that one reaction enough to change your roadmap, or are you seeing the same comprehension problem appear in other user behavior as well?