
Protosense
A Precise OCR & Translation Suite
Good day, builders!
I went live with a productivity app – Protosense (Google Play). It's a document management suite with a focus on OCR & data extraction.
How the idea was born
I was searching for a niche that fit my personal skills and had enough ground to justify building a solution to compete with the existing tools out there.
A routine market analysis ("compulsively spamming keywords into Google Play" in the non-LinkedIn language) revealed that the top results for the OCR apps are fragmented between two types of apps,
Free, but ad-ridden wrappers over outdated OCR libraries with 2.5 stars
Paid, but genuinely well-built, practical, and intuitive OCR solutions
I decided to try my best to fit into the second category.
How the idea evolved
After finalizing the initial proof-of-concept, I noticed two usability flaws,
I often had to crop an image, as I didn't want to extract the full body of text
Not every photo in my gallery is a document scan, or has any text at all
At that moment, Protosense already had a chatbot as a general-purpose solution for image understanding, but it was outside the image-to-text user flow. As someone who uses chatbots on a daily basis, it was clear to me that simply handing a chatbot to a user and hoping they won't switch to a different chatbot app they're already paying $19.99 for felt like a road to nowhere.
To address both, I embedded a new "Precise Extraction" into the existing flow. Instead of extracting every character from a picture, now I can type a couple of words to get,
Only the data from the first paragraph, converted into a table
Detect the names of traditional dishes I never saw before
Recognize the botanical names of the plants that grow in my yard
In other words, "Read the world", as the slogan goes.
To make it truly universal, translation also became available as a part of the same data extraction flow, closing the language barrier gap and opening the door for travelers and digital nomads to get the most out of the tool.
The «Early Access» stage
At this point, the app has to offer,
Document management. Simple and logical. Workspaces host documents, documents host pages.
Encryption. A non-negotiable feature when it comes to an app handling documents, which also opened a path for an experimental syncing feature between multiple devices.
OCR. A VLM-in-the-loop helps to understand the context and process blurry & smeared details from a scan or a poorly lit photo.
Precise Extraction. A feature to extract not only text, but any data the user's camera can catch.
Translation. Multiple target languages are available, alongside an experimental "Any" feature for translating the extraction into non-supported languages (or humorous & extinct!)
These features are available in the «Early Access» release channel for people to try and comment on.
The development stack
I'm new to the IndieHackers community, and the first thing that caught my attention is that you really like mentioning the dev tools you're using, which sounds only natural to me!
The client is built with Expo and WatermelonDB for the underlying database. It works well with Appwrite, the BaaS provider.
The backend is built with an Express API server & webhooks, Appwrite, and RevenueCat. The LLM calls go through a Bifrost gateway (a solid & Python-free alternative to LiteLLM), while Cloudflare helps to keep things secure.
A personal note
Thank you for making it through the post! How was it? Too salesy? Too casual? Do tell, I bite and listen!
P.S. If you're developing an Android app and are considering giving the "open testing" branch a go, definitely do. Google Play algorithms will not give it a boost, but you'll have a much easier time sending the app to people you know.
About
I wasn't satisfied with what's available in the top results of Google Play when I looked up "OCR" utilities and built a document management suite focused on solving the points others had missed.

Comment