Hi Indie Hackers đź‘‹
I just launched PDF to Markdown, a tool that converts PDFs into clean, structured Markdown.
This started as a personal problem.
I often work with PDFs that look structured — headings, lists, tables, code blocks — but once you try to reuse the content, everything falls apart.
Most tools do one of these things:
Dump raw text with broken line breaks
Flatten everything into paragraphs
Lose tables, lists, or code formatting
Or treat the PDF as an image and OCR everything
For developer docs, research papers, and technical PDFs, this makes Markdown basically unusable without a lot of manual cleanup.
I wanted something that produced Markdown I could actually commit to a repo.
PDF to Markdown is a web tool that focuses on structure first, not just text extraction.
It tries to preserve:
Heading hierarchy
Paragraph boundaries
Ordered / unordered lists
Code blocks
Tables (with sensible fallbacks when Markdown isn’t expressive enough)
You upload a PDF, and you get editable Markdown with a side-by-side preview.
No installs, no local setup.
PDFs are surprisingly hard.
They’re not documents — they’re drawing instructions.
Instead of assuming PDFs are “text with layout”, I treated them as layout + geometry + heuristics, then used AI only where rules break down.
A few conscious trade-offs I made:
I don’t support scanned PDFs yet — OCR is a different problem with different failure modes
I prefer predictable output over trying to perfectly replicate visual layout
When Markdown can’t express something well (e.g. complex tables), I fall back to clean HTML instead of mangling it
This keeps the output usable rather than “technically correct but painful”.
So far, it’s been most useful for:
Developers converting API docs or specs
Researchers turning papers into Markdown notes
Indie hackers migrating legacy docs
Anyone who wants Markdown that doesn’t require an hour of cleanup
Scanned/image-only PDFs aren’t supported yet
Very design-heavy PDFs (magazines, brochures) won’t convert well
It’s optimized for content, not pixel-perfect layout
This is still early.
If you work with PDFs regularly, I’d really appreciate:
PDFs that break the converter
Feedback on output quality
Suggestions on what matters most in Markdown output
👉 https://pdftomarkdown.pro
Thanks for reading — happy to answer any technical questions.