Hey IH community! I’m the solo founder behind Fidget and I wanted to share why I started this project, how it works, and get your thoughts on our unique approach.
Why Fidget Exists
I’ve spent countless hours scrubbing through long tutorials, webinars and conference talks - searching for "that one bit" I really needed. In a world where AI can write essays or generate images in seconds, wasting time manually navigating videos felt like a waste. As a programmer that's written and worked with AI, I knew there had to be a better way, so I built Fidget to reclaim those lost hours.
What Makes Fidget Different
Most video summarizers today simply transcribe audio or grab on-screen text. Fidget goes beyond that with a few unique systems, which all combine into a "multimodal approach."
From refining early prototypes, I've nailed down the following unique selling points:
Audio Understanding: Going beyond raw text using sentiment analysis
Visual and Scene Cues: Performing visual analysis of the video's key frames to understand what's actually happening.
Context Windows: Group together clusters of audio and visual cues in a sliding-window approach to persist meaning based on most-recently encountered context.
Metadata Analysis: Instead of ignoring important information in the metadata, we use it to infer wider context that feeds into the "context window" system.
Flexible Input: Plan on supporting many different video sources e.g. YouTube URL, MP4 files, Zoom recordings etc...
API SaaS: Open up the core functionality and share via API.
Why Multimodal Matters
By combining audio, visual and even metadata, Fidget’s summaries are more "aware" of the content. It’s not just a transcript - it’s a smart distillation of what matters most.
I'd be happy going into the details with the results of some of my early prototype tests.
Where We’re Headed
Right now, Fidget is in active development. Early prototypes have been built (to test the concept) and I’m refining the multimodal engine, improving language coverage, and getting feedback on edge-cases (think: noisy audio, multiple speakers, dense technical slides etc...). If you’re interested in following Fidgets development, I’d love your feedback - and you’ll be front-of-the-line for the official launch later this year! (Summer 2025.)
🚀 Join our waitlist: https://getfidget.pro/
Feedback Welcome!
Thanks for reading - I'm looking forward to showing you how Fidget develops!
Wow, it looks interesting.