TimedSubs

Turn scripts and voiceovers into subtitle files

Visit Website
June 15, 2026 I built TimedSubs because subtitles kept breaking in my YouTube workflow

I make faceless YouTube videos — the kind where the script is written first, the voiceover gets recorded, and the visuals are cut on top. By the time I reach subtitles, the hard part is supposedly done: the script is approved, the audio is final, and all I need is timed subtitle files.

That last step is where everything kept falling apart.

I'd run the audio through whatever subtitle tool I had and get back something close but never right. Timing drifted a few hundred milliseconds at a time until, ten minutes in, the captions were landing a full line behind the voice. Words got dropped, or quietly "improved" into something I never wrote. The SRT/VTT exports came out messy enough that I'd spend longer cleaning them up than I spent recording the voiceover. For a 60-second clip I could eyeball it. For a 20-minute video, checking every cue by hand was its own job.

The problem isn't transcription. It's that I already had the script.

Here's what bugged me about every tool I tried: they're all transcription-first. They listen to the audio and try to guess the words. But I wasn't guessing — I had the exact script sitting right there, already approved. I didn't need a machine to re-derive my own words and get them slightly wrong. I needed timing, a QA pass, and clean exports.

Once I framed it that way, the gap was obvious. I couldn't find a tool built for the case where the text is already correct and only the timing is missing. So I built it.

What TimedSubs actually does

TimedSubs takes a finished script plus the matching voiceover and turns them into subtitle files you can drop straight into your editor. That's the whole job.

It is not a YouTube downloader, not a generic transcription app, and not a video editor. The one rule it never breaks: your script is the source of truth. The audio is used only as evidence for where each line lands in time. The system is designed so the model doesn't get to overwrite your words — if what it hears disagrees with the script, the script wins, and the disagreement gets flagged instead of silently "fixed."

A few things I cared about under the hood:

- Script-locked alignment. Your submitted text stays exactly as written — no paraphrasing, no dropped punctuation, no quiet rewrites.

- Visible QA before export. It checks for overlapping cues, negative or zero-length durations, timing that runs past the end of the audio, lines too short to read, and reading speed that's too fast. You see the problems before you ship, not after a viewer leaves a comment.

- Real export formats. SRT, VTT, SBV, ASS, TXT, JSON, or a ZIP of everything.

- Heavy work runs off the web request. Alignment happens in a background worker, not crammed into a page load, so long files don't time out halfway through.

Why long audio is where it earns its keep

This is the part I most want feedback on. Short clips were never the real pain — you can hand-fix a 90-second video. The pain is the 15- and 20-minute pieces, where one early drift cascades and you're left scrubbing through hundreds of cues hunting for the one that knocked everything out of sync.

Because TimedSubs keeps your script intact and aligns against it in segments, longer audio doesn't fall apart the way it does when you throw the whole file at a transcription model and hope. It keeps your text, exposes the problem cues, and hands you a deliverable instead of a guess. It also works better for non-English scripts, where simple line-length and reading-speed rules often break first.

Where it's at

It's live now, and I'm still hardening the edge cases that only show up in real creator workflows. I'd rather have ten people tell me where it breaks than launch quietly.

Sample (real aligned output, no signup): https://timedsubs.com/en/sample

If you make long-form or faceless videos: does the script-first difference actually land for you from that sample, or does it read like just another subtitle tool? That's the thing I most need to know before I push harder on it.

6 Comments

  1. 1

    I'd be careful with one thing.

    The interesting question may not be whether the script-first difference is real.

    It may be what decision creators are actually making when they choose a subtitle tool in the first place.

    Those sound similar, but they can lead to very different conclusions about positioning, proof, and what deserves emphasis.

    I wouldn't make that call casually from the current signals.

    1. 1

      That’s a really useful distinction, and I agree.

      For me, “script-first” is not meant to be a standalone benefit. It’s more the mechanism behind a bundle of practical outcomes: when creators already have a script and voiceover, starting from the script should mean less cleanup, more consistent wording, more reliable exports, and a smoother production workflow.

      But you’re right that the positioning question is different: do creators actually think about the problem this way when choosing a subtitle tool, or do they describe the buying reason in more direct terms like saving time, fewer subtitle fixes, reliable exports, or workflow control?

      So I wouldn’t treat “script-first” as a proven category yet. I’m treating it as a positioning hypothesis that needs to be tested against how creators actually talk about the pain.

      Which makes me curious: when you or creators you know last picked or switched a subtitle tool, what was the actual trigger? Was it a specific video getting stuck at the subtitle step, or something more direct like speed, price, exports, or staying inside the editor they already use?

      That’s the part I’m trying to get closer to.

      1. 1

        That's exactly why I'd be careful.

        The interesting part isn't the trigger itself.

        It's what conclusion deserves confidence once you think you've found it.

        That's where founders can end up optimizing around a very convincing interpretation that turns out to be the wrong one.

        I wouldn't try to unpack that properly in a thread.

        If you're curious, drop your email and I'll send over the tighter version.

        1. 1

          That makes sense. I see the risk: not just finding a trigger, but getting too confident in the wrong interpretation of what that trigger means.

          IH won’t let my account post an email address yet. The contact is support at timedsubs dot com.

          I’d definitely be curious to read the tighter version.

          1. 1

            Sent you a note.

            I think the interpretation decision matters more than the trigger discussion itself right now.

            1. 1

              Got it, thanks for sending it over. I’ll read the note.

About

I make YouTube faceless videos, and subtitles kept breaking: timing drift, missing words, messy SRT/VTT exports, and too much cleanup after the script and voiceover were already done.