Klyde Labs

Turn long videos into topic-based clips, automatically

Visit Website
June 12, 2026 How I built a video clipping tool that doesn't slice at random (solo, as a non-team-of-one)

Most AI clipping tools work like this: scan a long video, find the "high-energy moments," cut there. That's fine for a podcast where any 30 seconds is interchangeable. It's useless for structured content — a course, a lecture, a coaching session — where a clip that starts mid-thought makes zero sense to anyone who didn't watch the original.

I kept hitting this with my own long-form video, so I built KlydeLabs to clip by topic instead of by timer. Here's how the pipeline actually works.

1. Word-level transcription, not just a transcript

The foundation is WhisperX with forced alignment. Plain Whisper gives you a transcript with rough timestamps; forced alignment pins down every word to the millisecond. That precision is what makes accurate karaoke-style caption highlighting possible — the word lights up exactly when it's spoken, not a half-second late. I store this in a word-timing audit table so I can validate alignment quality instead of trusting it blindly.

2. Finding topic boundaries

Once I have a precise transcript, the real work is segmenting by idea rather than by silence or volume. The goal is clips that each cover one coherent topic and could stand on their own as a standalone short.

3. The "cold open" problem (the hard part)

This is the bit nobody talks about. A clip can be on-topic and still be confusing if it opens with "...and that's why this matters" — a referent the viewer never saw. So before a clip ships, it runs through a gate that checks two things: is the opener self-contained, and does it reference something that isn't in the clip? If it fails, the boundary moves. This single piece of logic is the difference between "AI clip" and "watchable short."

4. Burning it together

Final step: render the clip with the word-level captions burned in, ready to post to Reels / Shorts / TikTok with no further editing.

What surprised me

I assumed the captions would be the hard part. They weren't — WhisperX handles that well. The genuinely hard problem was the boundaries: knowing where a topic actually starts and ends in a way that produces a clip a stranger can follow. That's most of where my time went, and it's the part I think is actually defensible.

I'm building this solo, mostly because it's the tool I wanted for my own content.

Question for the room: for those who've shipped AI-content tooling — how do you validate "good output" at scale when the thing you're judging (does this clip make sense?) is fundamentally subjective? Right now I lean on the cold-open gate + manual spot-checks, but I'd love to hear how others have approached the eval problem.

1 Comment

  1. 1

    One thing I'd be careful with:

    The interesting question may not be whether the clips make sense.

    It may be what outcome the clip is actually being trusted to create for the user.

    Those sound similar, but they can lead to very different product decisions.

About

Every AI clipping tool just hunts for "viral moments" and slices at random. KlydeLabs exists to repurpose structured video by topic — so each short actually stands on its own.