PDF to Markdown

Convert PDFs into clean, structured Markdown

Visit Website
January 2, 2026 I built a PDF to Markdown converter because nothing else worked reliably

Hi Indie Hackers 👋
I just launched PDF to Markdown, a tool that converts PDFs into clean, structured Markdown.

This started as a personal problem.


The problem

I often work with PDFs that look structured — headings, lists, tables, code blocks — but once you try to reuse the content, everything falls apart.

Most tools do one of these things:

  • Dump raw text with broken line breaks

  • Flatten everything into paragraphs

  • Lose tables, lists, or code formatting

  • Or treat the PDF as an image and OCR everything

For developer docs, research papers, and technical PDFs, this makes Markdown basically unusable without a lot of manual cleanup.

I wanted something that produced Markdown I could actually commit to a repo.


What I built

PDF to Markdown is a web tool that focuses on structure first, not just text extraction.

It tries to preserve:

  • Heading hierarchy

  • Paragraph boundaries

  • Ordered / unordered lists

  • Code blocks

  • Tables (with sensible fallbacks when Markdown isn’t expressive enough)

You upload a PDF, and you get editable Markdown with a side-by-side preview.

No installs, no local setup.


Some technical notes (for builders)

PDFs are surprisingly hard.

They’re not documents — they’re drawing instructions.

Instead of assuming PDFs are “text with layout”, I treated them as layout + geometry + heuristics, then used AI only where rules break down.

A few conscious trade-offs I made:

  • I don’t support scanned PDFs yet — OCR is a different problem with different failure modes

  • I prefer predictable output over trying to perfectly replicate visual layout

  • When Markdown can’t express something well (e.g. complex tables), I fall back to clean HTML instead of mangling it

This keeps the output usable rather than “technically correct but painful”.


Who this is for

So far, it’s been most useful for:

  • Developers converting API docs or specs

  • Researchers turning papers into Markdown notes

  • Indie hackers migrating legacy docs

  • Anyone who wants Markdown that doesn’t require an hour of cleanup


Current limitations (being honest)

  • Scanned/image-only PDFs aren’t supported yet

  • Very design-heavy PDFs (magazines, brochures) won’t convert well

  • It’s optimized for content, not pixel-perfect layout


I’d love feedback

This is still early.

If you work with PDFs regularly, I’d really appreciate:

  • PDFs that break the converter

  • Feedback on output quality

  • Suggestions on what matters most in Markdown output

👉 https://pdftomarkdown.pro

Thanks for reading — happy to answer any technical questions.

Comment

About

I built this because I frequently need to turn PDFs into Markdown for writing, documentation, and LLM workflows. Most existing tools either output messy text, lose structure, or treat everything as plain OCR. I wanted s