
Our team kept hitting the same gap: an agent could help us generate an image or a clip, but finishing a 30-second video still meant moving files between several tools and rebuilding the edit by hand. The awkward part was not another model prompt. It was keeping the assets, the canvas, and the timeline in one place.
So we built BeatDesign, an Apache 2.0 open-source workspace that runs locally. It has a node canvas for exploring prompts and media, a timeline editor for assembling the result, and a shared asset library between them. You can trim and split clips, adjust audio, layer images, style SRT captions, and export an MP4 in the browser.
The design choice we cared about most was making the same project available to both a human editor and an agent. BeatDesign exposes 29 local MCP tools for project, asset, canvas, generation, and editing operations. Claude Code, Codex, and other MCP clients can move work forward; a person can inspect the timeline and change it. That is the workflow we wanted when we started.
There is a tradeoff: local-first does not mean every operation is offline or free. Basic asset handling, editing, preview, and export do not need an API key. AI generation and analysis need a model connection and incur usage charges. We use BeatAPI as the default connection, and the provider layer is open to extension.
To try it, clone https://github.com/BeatAPI/BeatDesign and run pnpm install, pnpm db:push, and pnpm dev with Node.js 22+ and pnpm 10+. The product site is https://design.beatapi.io/.
I would be interested in how other builders handle the handoff from generated assets to a finished video. Do you keep everything in one project, or still move the edit to a separate app?
Disclosure: I am part of the team building BeatDesign and BeatAPI. The product screenshots attached are from the actual app, not model output.
The human-plus-agent handoff is the interesting boundary here: 29 MCP tools sound like plenty, but the real test is whether a tired human can fix a trim without losing project state. Tracking the edits people make after an agent pass seems like a great roadmap signal—and possibly the first timeline where “undo” deserves its own product manager 😄
I like the tired-human test 😄 My thinking is that the agent should do most of the work, then the human makes the final tweaks and approves it. That’s a direction I’d like to improve next: making it easier to see what the agent changed and adjust the details.
Have external users actually finished videos faster by keeping the agent, assets, and timeline together, or is that workflow advantage still mainly validated by your own team?
Mostly our own team so far. We’ve run the Canvas → timeline → MP4 flow end to end, but we don’t have external-user time-to-finish data yet. That’s the next thing to measure: finished videos and time lost moving assets between tools.
That’s a clean next test once external users come in. If you’re open to it, what’s the best email to reach you on?
Thanks, Aryan — karmen@beatapi.io is best. Happy to talk through the external-user test.
I'm on the other end of this. Nothing generated, my app demo is just screen recordings off the phone.
But the "moving files between several tools" part is exactly the same pain, trim here, captions there, export somewhere else.
Does the timeline work fine with plain screen recordings, or is it built around generated clips?
Yes—plain screen recordings work. You can import local MP4/MOV/WebM footage, trim or split it, then preview and export MP4 without an API key. Browser codec support varies, so if yours fails to preview, tell me the format.
local video workspace for agent-generated clips is a smart wedge - generation is solved, editing is the bottleneck nobody owns. we handle the conversion side of video on swapfile.live (all in-browser, nothing uploaded) and the edit step is exactly where our users get stuck next. are you planning to open the workspace to non-agent uploads too?
Yes—non-agent uploads already work. Bring a local video into a Project, mix it with generated clips if you want, and edit/export without an agent or API key. What edit do your users reach for first after conversion?
Putting the timeline beside the agent tools feels like the right boundary. The human can inspect the edit without losing the agent's project state, and the local first split between editing and paid generation makes the cost tradeoff clear. I would track how often people touch the timeline after an agent pass since that should reveal which operations deserve better tools.
That's a useful metric. We don't track it yet, but I'd start with human fixes to trims, clip order, and captions after an agent pass. Repeated corrections would tell us which MCP tools still need work.
The shared project state is the key differentiator. An editable manifest of asset IDs, clip ranges, captions, and unresolved decisions could make agent-to-human handoff recoverable after failures. Is that state exposed through the MCP tools?
Yes. The project, assets, Canvas, timeline, and generation history stay local, so another MCP-capable agent can reconnect and continue editing. We don't yet store unresolved human decisions as a separate record; that's a real gap in the handoff.