
FluentCap
Real-time captions and translation for any video or audio
Hey IH! π
Quick update on FluentCap β the free desktop app for real-time transcription & translation.
What we shipped: Two pronunciation features β IPA (International Phonetic Alphabet) lookup and a Speak button.
Why we built it: Our users are mostly language learners who watch content in foreign languages with live subtitles. We kept hearing the same frustration: "I study words but can't recognize them when natives speak."
Turns out there's solid research behind this. When you learn a word with wrong pronunciation, your brain stores a wrong sound model. When the real sound comes in, it doesn't match β you don't understand.
The fix was surprisingly simple:
Select a word from the transcript β show its IPA transcription (offline, 25+ languages)
Add a "Speak" button β hear it pronounced correctly via the system's native TTS
No new API costs, no server calls for IPA (bundled dictionaries), and the Speak button uses the free Web Speech API. Literally zero marginal cost per user.
The learning workflow: Watch content β encounter unknown word β select it β see IPA β click Speak β shadow the pronunciation. All in-flow, zero context switching.
Business model reminder: FluentCap is free forever with BYOK (Bring Your Own Key). Users connect their own STT API keys. Providers give $50-200 in free credits (~140-750 hours). After that, users pay providers directly at wholesale rates ($0.15-0.40/hour).
π Full article: How IPA and Speak Help You Master Pronunciation
π Website: https://fluentcap.live
β¬οΈ Download: https://fluentcap.live/download
Would love feedback from anyone learning a second language! π
Just shipped Highlights for FluentCap, addressing consistent user feedback about a workflow gap.
When transcribing long content like 2-hour lectures or movies, users noticed important moments but had no systematic way to mark and return to them. Valuable information got buried in thousands of words.
What We Built
During live transcription, Instant-Apply Mode: select text and highlight applies immediately without menu interaction. Speed matters during live sessions, and context menus introduce friction.
When reviewing past sessions in History Mode, a context menu appears with Highlight and Copy options. To remove highlights, hover to reveal the Remove button.
All highlights collected in a Highlights tab in the sidebar, grouped by session. Copy All exports everything as structured text. Clear All resets with confirmation.
The key feature is cross-session navigation. Click any highlight in the Collection, and FluentCap switches to that session, scrolls to the exact location, and focuses the highlight visually. Users can jump to any saved moment across their entire history in seconds.
Technical Challenge
STT providers sometimes revise text after sending it, adding punctuation or fixing capitalization. Highlights could break if the text changes. We implemented a Hybrid Matching Strategy: check exact position, then case-sensitive search, then case-insensitive fallback.
Observed Use Cases
Language learners mark vocabulary, export to Anki via Copy All. Researchers mark quotes and use navigation to compile citations. Business users mark action items during meetings without breaking focus.
What We Learned
The Instant-Apply versus context menu split based on mode was the right call. Live sessions need speed. History review wants deliberate control. Same interaction for both would compromise one.
Cross-session navigation turned out more valuable than expected. Users said this is what makes highlighting genuinely useful.
All data local. No server involvement. Free like all FluentCap features.
Feedback welcome.
Homepage: https://fluentcap.live
Feature details: https://fluentcap.live/blog/fluentcap-highlight-feature
2 Likes
3 Comments
3 Comments
-
1
Congratulations on your launch. It looks impressive! What channels are you exploring to attract early users?
-
1
I just posted videos on TikTok and Reddit and did SEO for my website. Actually, I donβt know how to bring my application to end users
-
1
It's quite common for early-stage apps to struggle with visibility. While TikTok can generate views, its results can be unpredictable, and search engine optimization (SEO) requires time to bear fruit. For many apps, the key challenge is connecting with users who are already discussing the problem the app addresses. This is where Reddit can be effective, especially if the approach focuses on fostering discussions. Iβd be happy to share a couple of strategies that typically help reach genuine users without relying on ads.
-
-
We just published a comprehensive guide on why delayed captions work for language learning, backed by Second Language Acquisition research.
The core problem: When subtitles appear simultaneously with audio, your brain takes a shortcut. It reads instead of listens. Eye-tracking research shows viewers spend 68-84% of their viewing time looking at subtitles, not the video content. Your ears become secondary input.
The solution: Delay captions by 1-1.5 seconds. This timing is optimal because it's long enough to force genuine listening and hypothesis formation, but short enough to maintain comprehension flow.
When you watch with delayed captions, your brain must:
Process the audio first (the text simply isn't there yet)
Form a hypothesis about what was said
See the caption appear and confirm or correct
Vanderplank (2016) identified this as the "text-dependence" problem in his research on captioned media learning. Learners develop excellent reading comprehension but underdeveloped listening skills. Delayed captions break this pattern by forcing active auditory processing.
The results are significant. After 6-8 weeks of consistent practice, most learners report 30-50% improvement in pure listening comprehension without any text support.
We included a detailed 4-week training progression in the article, moving from 0.5s delay in Week 1 to 1.5s delay by Week 4.
Full article with training plan: https://fluentcap.live/blog/delayed-captions-listening-training/
Try FluentCap free: https://fluentcap.live/
Research sources:
Vanderplank (2016): https://doi.org/10.1017/S0261444815000142
Eye-tracking in subtitled video: https://www.tandfonline.com/doi/full/10.1080/14790718.2018.1465571
Bjork Lab on Desirable Difficulty: https://bjorklab.psych.ucla.edu/research/
1 Like
Comment
I built FluentCap because existing caption tools are either expensive, limited to specific platforms, or lack essential features.
FluentCap is completely FREE with: dual audio capture (mic + system), real-time AI transcription (50+ languages), instant translation, automatic meeting notes, and fully customizable styling.
It works everywhere - YouTube, Netflix, Zoom, podcasts, any audio source.
Whether you're learning languages, taking meeting notes, or need accessibility captions - FluentCap makes it simple and free.
Try it now: https://fluentcap.live/
Read more:
https://fluentcap.live/blog/introducing-fluentcap/
1 Like
Comment
About
I built FluentCap because caption tools are expensive and limited. FluentCap is FREE with: dual audio capture, real-time transcription, instant translation, meeting notes, and custom styling. Works with any audio source.


1 Comment
The interesting part here isn't the IPA or TTS feature itself β it's that you're keeping the learning loop inside the transcript instead of sending users to another tool. That makes the feature feel like part of the product's core workflow rather than an add-on.