Hey Indie Hackers π
I just launched Extraction.app β a tool that extracts structured data from any type of document using AI and natural language prompts.
About 4 months ago, someone in the construction industry told me about a frustrating workflow: they were juggling multiple document formats (contractor invoices, receipts, forms, reports) and spending hours manually copying data into spreadsheets to create financial reports.
When I asked why they didn't use existing tools, they said: "We know AI exists. We know it can help. But nothing works the way we actually need it to."
That conversation stuck with me.
The real problem:
People know AI can process documents, but current options either:
Only work with specific document types (invoice-only, receipt-only tools)
Require technical setup or custom development
Use generic LLMs that hallucinate data (not great for financial numbers)
What I built:
A flexible extraction platform that works with any document type β not just invoices.
Want to extract data from:
Invoices and receipts? β
Stacks of CVs to compare? β
HR documents to compile? β
Contracts, forms, reports? β
You describe what you want in natural language, and the AI extracts it into structured data (JSON, CSV, Excel, whatever you need).
Why I'm posting here:
Feedback needed: Does document data extraction solve a real problem you have? (Beyond just invoices)
Use case discovery: What document types would you want to extract data from? I'm curious what workflows this could improve.
Marketing reality check: I'm a builder. Getting people to discover this has been harder than building it. Where do ops/finance/HR people hang out online?
Try it if you're curious:
Free tier, no credit card. Genuinely want to know if this is useful.
π extraction.app
Honest question:
Am I solving a real problem, or did I just build something because it was technically interesting? I'd love brutal honesty.
Happy to answer questions about the tech, the journey, or just chat π
This looks interesting! Document extraction is definitely a real problem. We're building Docuct in a similar space. Curious β how do you handle different document formats and layouts?