Over the past few months, I’ve been building an AI-native platform called DOVENI.
Like many founders, I initially thought the biggest challenge would be choosing the right LLM.
It wasn’t.
The real challenge was making AI outputs trustworthy.
Some of the hardest problems had nothing to do with prompts:
The model is only one piece of the system.
The product around the model is what users actually experience.
That completely changed how I think about building AI software.
For those of you building AI products:
What ended up being much harder than you expected?
P.S. If anyone is curious, DOVENI is the project that taught me these lessons:
https://doveni.app
Same lesson, different product. I'm building a startup idea validation service — the AI part was easy to set up. The hard part was making the output trustworthy enough that a founder would actually make a decision based on it. That's why I switched to manual research with AI just for structure. Users don't trust AI verdicts. They trust cited sources and human judgment.
The false-positive problem seems especially brutal with AI products. A model can produce an impressive result, but if users have to second-guess whether each finding is actually meaningful, the impressive part doesn't matter much. Curious which of those problems took the longest to get under control.