1
0 Comments

Building SubtitleGenerator: From User Research to Product Hunt #3

On August 22, SubtitleGenerator—an AI subtitle generator built to take a video all the way to publish-ready subtitles—was #3 Product of the Day on Product Hunt.

What mattered to me was the loop behind that result. Research shaped the product, the product shipped, and the launch gave me both a verified outcome and a new question from a real person.

A subtitle generator should finish the whole job

Before building SubtitleGenerator, I spent weeks reading public complaints about leading AI subtitle tools and testing the workflow myself. The same pattern kept appearing. A model returned a mostly-right transcript, but the creator still had to turn it into subtitles they could trust and publish.

A timed transcript is useful raw material. The words still need review and correction. The track may need translation. The subtitles need to fit the video, and a finished video or subtitle file still has to come out at the other end.

That became the product definition: generate, fix, translate, style and export subtitles in one browser workflow.

Fix is the sharpest difference inside that path. It marks uncertain words, keeps the remaining review count visible, and gives the review an “All clear” finish line. But Fix is not the whole product. It is the trust layer inside a larger delivery system. Translation, styling and export belong to the same promise: an AI subtitle generator should finish the subtitle job instead of stopping when text appears.

From research to #3 Product of the Day on Product Hunt

I kept the Product Hunt launch tied to those decisions. The tagline was “From video to publish-ready AI subtitles—all in one browser.” The gallery used real product screens to show the no-signup entry, Fix, “All clear,” styling, translation and export. My maker comment ended by asking what still slows people down after an AI transcript is generated.

Product Hunt put that complete workflow in front of a broader external audience. The result was not validation of every product assumption. It was a verified outcome from the first loop. Public complaint research became a product design, the design became a working end-to-end product, and that product became #3 Product of the Day on Product Hunt.

Then one reply opened the next loop.

What one real comment revealed about speaker labels

Gal Dayan answered my question with a specific pain point: “speaker labeling when two people talk over each other or overlap slightly.”

He described an interview or podcast workflow where the text could read fine. Assigning each line to the right speaker still became manual work after the “AI part” was technically done. This was not a customer testimonial. It was a public commenter describing work he repeatedly has to finish himself.

Product Hunt discussion about speaker labeling for interview and podcast subtitles after the SubtitleGenerator launch

I do not see that as a disconnected feature request. It is another place where subtitle generation can stop before the job is delivered. If the words are right but the creator still has to decide who said every line, the AI has produced readable text—not finished multi-speaker subtitles.

Speaker attribution extends the same product definition. Fix answers which words still need attention. Speaker labels answer who each line belongs to. Translation carries the result into another language. Styling and export turn it into something the creator can publish.

I did not turn one comment into a shipping promise. I went back to competitor pages, keyword data and current search results. That research showed that the problem language and search interest extend beyond one comment. Across those sources, the recurring language included speaker labels, speaker identification and multi-speaker transcription. That did not prove that speaker labeling should jump to the top of the product roadmap or become a standalone page.

The research also made the opportunity more precise. Speaker labeling is not an empty market; several products already address parts of the task. The gap worth testing is whether automatic labels, fast correction, renaming and reassignment can work as one creator workflow—especially with overlapping speech—and remain coherent through translation and export.

What is designed—and what still needs proof

I am not starting from zero technically. Before launch, a provider test on a clean 12-second, two-speaker clip identified all four speaker turns correctly, with confidence scores from 0.94 to 1.00. It also returned the segments and word-level assignments needed to build an editor.

Since launch, I have turned that evidence into an approved interaction design. It keeps speaker labels inside the existing editor, reveals them only for multi-speaker projects, and gives creators a correction path for renaming, adding, merging and reassigning speakers.

That is meaningful product progress, but it is not a shipped feature. The design has not been validated in the real editor or connected to live speaker data, saved drafts and exports. The clean test also says nothing about real interviews, overlapping speech or noise.

The next step is not “turn on speaker labels.” It is to prove that:

  • the approved labels, rename, add, merge and reassignment interactions are clear on desktop and mobile;
  • the system holds up across real interviews, podcasts, overlap, crosstalk and noise;
  • corrected speaker assignments stay coherent through translated tracks and exports;
  • creators value the workflow enough—and whether automatic accuracy or fast correction matters more.

That is the positive feedback loop I wanted from a launch. Research defined the full job. I built and launched the end-to-end workflow. Product Hunt produced a verified result. One real reply exposed the next place the same definition has to hold. Follow-up research and the clean test narrowed the problem, and the approved interaction design turned it into a concrete hypothesis. What comes next is validation, not a shipping claim.

You can see the original exchange on the SubtitleGenerator Product Hunt page.

For interviews or podcasts, when do subtitles feel finished: when the words are right, when the speakers are right, or only when the export is publish-ready?

on August 24, 2026