18
32 Comments

We built a local video workspace because our agent could generate clips but could not edit them

Our team kept hitting the same gap: an agent could help us generate an image or a clip, but finishing a 30-second video still meant moving files between several tools and rebuilding the edit by hand. The awkward part was not another model prompt. It was keeping the assets, the canvas, and the timeline in one place.

So we built BeatDesign, an Apache 2.0 open-source workspace that runs locally. It has a node canvas for exploring prompts and media, a timeline editor for assembling the result, and a shared asset library between them. You can trim and split clips, adjust audio, layer images, style SRT captions, and export an MP4 in the browser.

The design choice we cared about most was making the same project available to both a human editor and an agent. BeatDesign exposes 29 local MCP tools for project, asset, canvas, generation, and editing operations. Claude Code, Codex, and other MCP clients can move work forward; a person can inspect the timeline and change it. That is the workflow we wanted when we started.

There is a tradeoff: local-first does not mean every operation is offline or free. Basic asset handling, editing, preview, and export do not need an API key. AI generation and analysis need a model connection and incur usage charges. We use BeatAPI as the default connection, and the provider layer is open to extension.

To try it, clone https://github.com/BeatAPI/BeatDesign and run pnpm install, pnpm db:push, and pnpm dev with Node.js 22+ and pnpm 10+. The product site is https://design.beatapi.io/.

I would be interested in how other builders handle the handoff from generated assets to a finished video. Do you keep everything in one project, or still move the edit to a separate app?

Disclosure: I am part of the team building BeatDesign and BeatAPI. The product screenshots attached are from the actual app, not model output.

on September 29, 2026
  1. 1

    That gap is familiar. Generation is the easy part; the painful part is the last mile of timing, cuts, and captions when files bounce between tools. Keeping the edit workspace local and next to the agent sounds right if you care about iteration speed. Curious whether you keep a simple timeline model or something more prompt-driven for the cuts themselves.

  2. 1

    Nice separation of the agent and human editor. The failure mode I'd watch is state drift between the agent's asset graph and the timeline after a human makes a change. An append-only project event log (asset id, operation, input hash, timestamp) plus idempotent MCP mutations makes retries and replays much safer than asking the model to infer current state from a screenshot. Have you considered exposing a small “validate project” tool that returns broken references, missing media, and renderability before export? That kind of preflight could make the handoff reliable without forcing the agent to own every edit.

  3. 1

    The handoff between generated clips and a human editor is exactly the part that gets glossed over. We see a similar gap with sports practice footage: athletes record constantly, then rarely revisit it because turning raw clips into a useful next step takes work. We’re building Vidi AI to watch a practice clip and point out a concrete fix for the next session. Curious how you’re thinking about feedback from the human editor becoming context the agent can use on the next pass.

  4. 1

    The human-and-agent handoff is the most interesting design constraint here. A useful way to keep that reliable is to make each project save explicit checkpoints: source assets, timeline state, and export settings, each with a small validation pass before the next tool runs. That gives the human a clear recovery point when an agent operation fails instead of forcing a full rebuild.

  5. 1

    That's so cool!

  6. 1

    Behind the machine we must need one real human to check accuracy. Sometimes machines can do mistacks which wipe out total reputations.

  7. 1

    oh, wow, interesting. I wonder how it compares to e.g. remotion

  8. 1

    The shared project model between a human editor and an agent seems like the key differentiator—not just putting generation and editing side by side. I’d watch where handoffs break down: asset naming, timeline state, or the final review step. That friction could tell you which MCP actions deserve the strongest guardrails and undo support.

  9. 1

    This is an interesting approach. The “generated asset → finished output” gap is definitely where a lot of AI workflows still feel fragmented.

    We’re seeing a similar problem on the AI visibility side: brands can generate tons of content with AI, but they often have no idea what AI models are actually saying about their brand when users ask relevant questions.

    I’ve built a tool that lets brands track how they’re being mentioned in AI responses and compare their visibility against competitors — essentially, seeing who gets recommended and how often.

    Curious to see where you take the agent + human editing workflow.

    Also I am documenting my experiment so you can check out my tool if you have time:) Currently it's free and stateless ✅

  10. 1

    One gap nobody's raised yet: brand rules. When an agent does most of the edit, the fixes a human makes afterwards are usually rules the agent never knew, not creative calls.

    We have a few strict ones for UtilitySEO's ads: no logo on the creative itself, and single UI elements in 3D perspective rather than screen recordings. They live in notes our agent reads each session, not in the tools, so nothing stops a wrong edit. It just gets caught later.

    If a BeatDesign project could hold those as constraints (fonts, safe zones, no-logo areas, caption style) and the MCP tools checked edits against them, the tired-human pass gets much shorter.

    Captions are the other one. Our site runs in 7 languages, and one timeline exporting a caption track per language would save rebuilding the edit each time.

    Can a project carry rules the MCP tools enforce, or only assets?

  11. 1

    This is a really interesting boundary for agents. Generating the asset is getting easier, but maintaining state across the full workflow is where things start getting hard.

    The shared project between the agent and human feels especially important here. Curious how you’re thinking about failures mid-workflow — if the agent makes the wrong edit or loses context, can you replay exactly what it changed and recover cleanly?

  12. 1

    The human-plus-agent handoff is the interesting boundary here: 29 MCP tools sound like plenty, but the real test is whether a tired human can fix a trim without losing project state. Tracking the edits people make after an agent pass seems like a great roadmap signal—and possibly the first timeline where “undo” deserves its own product manager 😄

    1. 1

      Haha, “undo deserves its own product manager” 😄 — I can totally see that.

      Your point about tracking what humans change after an agent pass is interesting. I’m actually working on a somewhat similar problem, but from the AI visibility side.

      I built a tool that lets brands see how AI models mention and recommend their brand compared with competitors — basically, what AI is saying about you and how often you show up

      I’m documenting the experiment as I build it, so if AI visibility/AI search is something you’re interested in, and have time you can check out my post on indie hackers : )

    2. 1

      I like the tired-human test 😄 My thinking is that the agent should do most of the work, then the human makes the final tweaks and approves it. That’s a direction I’d like to improve next: making it easier to see what the agent changed and adjust the details.

  13. 1

    Have external users actually finished videos faster by keeping the agent, assets, and timeline together, or is that workflow advantage still mainly validated by your own team?

    1. 1

      Mostly our own team so far. We’ve run the Canvas → timeline → MP4 flow end to end, but we don’t have external-user time-to-finish data yet. That’s the next thing to measure: finished videos and time lost moving assets between tools.

      1. 1

        That’s a clean next test once external users come in. If you’re open to it, what’s the best email to reach you on?

        1. 1

          Thanks, Aryan — karmen@beatapi.io is best. Happy to talk through the external-user test.

  14. 1

    I'm on the other end of this. Nothing generated, my app demo is just screen recordings off the phone.
    But the "moving files between several tools" part is exactly the same pain, trim here, captions there, export somewhere else.
    Does the timeline work fine with plain screen recordings, or is it built around generated clips?

    1. 1

      The bigger point you made about moving files around really resonates with me. I’m working on a similar problem, but on the AI search side.

      I built a tool that shows brands how AI models mention and recommend them compared to their competitors

      I’m documenting the whole experiment here on Indie Hackers as I build it. So feel free to check it out if you have time :)

    2. 1

      Yes—plain screen recordings work. You can import local MP4/MOV/WebM footage, trim or split it, then preview and export MP4 without an API key. Browser codec support varies, so if yours fails to preview, tell me the format.

      1. 1

        I see. Thanks for the answer. I'll try it.

  15. 1

    local video workspace for agent-generated clips is a smart wedge - generation is solved, editing is the bottleneck nobody owns. we handle the conversion side of video on swapfile.live (all in-browser, nothing uploaded) and the edit step is exactly where our users get stuck next. are you planning to open the workspace to non-agent uploads too?

    1. 1

      Yes—non-agent uploads already work. Bring a local video into a Project, mix it with generated clips if you want, and edit/export without an agent or API key. What edit do your users reach for first after conversion?

  16. 1

    Putting the timeline beside the agent tools feels like the right boundary. The human can inspect the edit without losing the agent's project state, and the local first split between editing and paid generation makes the cost tradeoff clear. I would track how often people touch the timeline after an agent pass since that should reveal which operations deserve better tools.

    1. 1

      Yeah, I think that human + agent boundary is really interesting. The edits people make after the agent finishes probably reveal a lot about where the agent still falls short.

      I’m actually exploring a similar idea from a different angle with AI visibility. I built a tool that lets brands see how AI models mention and recommend them compared with their competitors.

      Basically, instead of guessing how your brand is showing up in AI responses, you can actually have a clear picture.

      I’m documenting the experiment here on Indie Hackers as I build it, so feel free to check it out if you have time :)

    2. 1

      That's a useful metric. We don't track it yet, but I'd start with human fixes to trims, clip order, and captions after an agent pass. Repeated corrections would tell us which MCP tools still need work.

  17. 1

    The shared project state is the key differentiator. An editable manifest of asset IDs, clip ranges, captions, and unresolved decisions could make agent-to-human handoff recoverable after failures. Is that state exposed through the MCP tools?

    1. 1

      The idea of making the agent’s state recoverable is especially interesting. I’m working on a different side of the AI ecosystem — I built a tool that lets brands track how AI models mention and recommend them compared with their competitors.

      I’m documenting the experiment as I build it here on Indie Hackers. Different problem, but the common thread is making what’s happening inside AI systems more visible and measurable. you can check it out if you have time:)

    2. 1

      Yes. The project, assets, Canvas, timeline, and generation history stay local, so another MCP-capable agent can reconnect and continue editing. We don't yet store unresolved human decisions as a separate record; that's a real gap in the handoff.