My first startup documentary took 5 days and cost about $70 — and that was the "cheap" way. I bounced between Gemini for research and the story bible, ElevenLabs for voice, Veo for clips, and CapCut to stitch it together. Every tool handoff introduced drift: the voice didn't match, the visual style shifted scene to scene, and I kept re-exporting to patch it.
I wasn't doing that again for video #2.
So I built Brand Narratives AI: paste a URL, GitHub repo, or pitch deck, and it produces a documentary-style or explainer video in under 5 minutes. Same pipeline end to end — Rails backend, Gemini for the story bible, Veo for the cinematic footage, ElevenLabs for voice, FFmpeg for assembly.
Two honest problems I haven't fully solved: Veo still can't render text reliably (working around it with Remotion), and generation costs $1-2 per 30 seconds, so margins are tight at low price points.
The proof it's not just a toy for me: someone found it searching on Brave and paid $12 for a video with zero involvement from me. That's the first time a stranger validated it was worth money.
Next up — cheaper/faster rendering (moving off Heroku to Cloud Run), fixing text-in-video, and a use case I'm excited about: turning employee handbooks into explainer videos for companies that can't afford a video team.
If you've ever needed a demo/explainer video and given up because of cost or time, I'd love to know what stopped you — trying to figure out where this is most useful right now.
Would love to see this side by side — could you share a before/after example, the manual 5-day version next to what the tool produced in 5 minutes? Curious how close the output style and pacing actually get once you remove the manual re-exporting and drift between tools
Are you insane?
haha why would you say that? Am totally fine, I think.