Title:
I built too many features into my PDF app. So I cut most of them.
Body:
I’ve been building a PDF tool called PDflow for a while.
It started with a very simple idea: PDF → Excel.
Then I kept adding things.
Merge, split, compress, watermark, templates, OCR, different conversion modes… At some point, I had built quite a lot.
And weirdly, the more complete the product became, the less sure I was about it.
Not because it didn’t work.
Because I started asking myself:
Why would anyone use this instead of Adobe, WPS, Smallpdf, iLovePDF, or one of the hundred PDF tools already out there?
I didn’t have a convincing answer.
“Free” isn’t a moat.
“Local processing” is useful, but probably not enough.
And trying to win by adding more PDF features felt like a race I was never going to win.
The thing that changed my mind came from testing harder PDF-to-Excel files.
A lot of them technically converted fine. The Excel opened. Nothing crashed.
But then you looked closer.
A row was shifted.
Wrapped text became two rows.
A table broke across pages.
A SKU matched the wrong quantity.
And the worst part was: the file still looked successful.
A user recently left a comment saying:
“Having to hunt for silent layout errors is what ruins the whole experience.”
That sentence really stuck with me.
Because maybe the real problem isn’t:
“Can this PDF be converted?”
Maybe it’s:
“Can I trust the result without checking everything again?”
So I decided to cut most of PDflow.
For now, I’m hiding merge, split, compression, watermarking, PDF-to-Word, PDF-to-PPT and the rest of the toolbox.
I want to focus on one workflow:
PDF → extract → detect suspicious data → review against the source → fix → export
If something looks wrong, PDflow should tell you where to look, and ideally show exactly where that value came from in the original PDF.
I’m also rebuilding part of the architecture, because the first version mixed UI, worker state, recovery logic and export logic too tightly. I spent way too much time fixing one thing and breaking another.
This time I’m trying to be much stricter.
No more features just because they sound useful.
First I want to answer one question:
Can PDflow make difficult PDF-to-Excel work meaningfully easier to verify?
If not, I’d rather find out now.
If you work with PDFs that “convert successfully” but still force you to manually check everything afterward, I’d genuinely like to hear what usually goes wrong.
One source of truth worth checking before you infer anything from layout: if the PDF is tagged (the structure PDF/UA requires for accessibility), the table is already in the file as rows and cells, and a table that breaks across two pages should still be one table there. Chrome can write these tags (Puppeteer's page.pdf has a tagged option for it), scans and a lot of ERP exports don't have them. So it won't cover your hard files on its own, but where the tags exist they settle the shifted-row and wrapped-cell cases instead of guessing them. Do you know what share of your problem files are tagged?
That’s a really useful point — and no, I don’t know the share yet.
I’ve mostly been looking at visual/native extraction and OCR paths, so I haven’t measured how many of the difficult samples already contain usable table structure tags.
I’m going to add that to the corpus audit before doing more layout inference: tagged vs untagged, and whether the existing structure is actually usable.
If the tags are reliable, I agree — it makes much more sense to treat them as a stronger signal than trying to reconstruct the same structure from geometry.
Cutting the toolbox is the decision that makes the product real. What made a similar cut stick for me was writing the one job and the who-not-for in the same pass: one workflow only, plus an explicit line that if someone wants the full suite (merge, watermark, "also convert Word"), they are the wrong user. Otherwise the first "can it also…" request puts the hidden features back.
Your trust question is sharper than another converter: "can I trust the result without checking everything again?" I'd put that sentence above the feature list and refuse any new verb until one hard file type passes it. The silent-error comment is better research than a feature request. A practical lock: pick 5 ugly PDFs, name the exact failure you will surface (shifted row, broken cross-page table), and ship only the review step that points at the source. If that doesn't save a check, you found out before rebuilding the rest.
This is exactly the kind of constraint I need right now.
The “who-not-for” part especially resonates. I’ve been defining the one job, but not explicitly saying that someone looking for a full PDF suite is probably not the user I should optimize for.
I’m actually moving toward almost exactly the test you described: a small set of ugly PDFs, each with a known failure mode, and proving that the review step can surface the problem and point back to the source before I rebuild the rest.
If that doesn’t meaningfully reduce the manual check, I’d rather learn that there.
Update: pushed a fix based on similar feedback — column explosion on messy quotations went from 42 cols down to 7, and it's actually usable now (not just "technically an Excel").
Next I want to stress-test it on structures I haven't seen — multi-level headers, merged cells, tables that span pages. If either of you has an ugly one lying around (sanitized is totally fine), I'd genuinely like to try it and show you the before/after.