Minutehand

Automatic meeting transcripts for your AI agent to read

Visit Website
August 10, 2026 My AI agent created my launch video's background score in Python.

The soundtrack on my launch video was written in Python. Not recorded, not licensed, not pulled from a library. And not "AI-generated". It was generated though.

There was no soundfont or MIDI synth on the machine, which ruled out playing a .mid file. So the waveforms got synthesised directly, a sample at a time: square waves with a variable duty cycle for the bass and the lead, a triangle underneath for the low end, and sample-and-hold noise shaped into hats and a snare. Those are the same four voices an NES sound chip had, roughly what a C64 or an arcade cabinet was working with.

Which is why it sounds like 1987. The constraint chose the genre rather than the other way round. It runs at 112bpm in D minor, sparse for the first few seconds and opening up as the app appears, and it is a bit over two hundred lines with no audio library behind it.

I did not write that code. I described what I wanted and an agent wrote it, which is the actual point of this post. What I did do was reject it twice, and that turns out to be the whole job.

Attempt one for the second video was the first track sped up. Same riff, same key, faster. It is exactly what you get when your brief is "now do one for the other video" — the machine has no reason to think you meant a new piece rather than a variation. So the second track is now genuinely separate: C major instead of D minor, a bright arpeggio instead of the pentatonic riff, octave-jumping bass, offbeat open hats. Different piece, same synth.

Attempt two was worse in a more interesting way. Right tempo, right feel, and it looped a single bar of material for a full minute. It sounded like it was building to something that never arrived. A bed still needs an arrangement, and "make it sound like X" does not imply one.

So attempt three got a bar-by-bar plan written down before a note was generated: intro, groove, melody enters, strip back to a snare roll on the five chord, then the drop with the counter-line and the melody an octave up, a closing tag under the end card. The chords never come round in the same order twice. The section lengths are chosen so the last flourish lands exactly on the video's fade to black.

That is the bit worth passing on. Working this way, the taste is the bottleneck, not the skill. I cannot write a synthesiser and I did not need to.

What I needed was to know that a loop is not an arrangement, and to be willing to say "no, again" twice to something that already sounded fine.

The captions came out of Remotion, React components rendered as transparent stills and composited over the footage with ffmpeg, which is far quicker than pushing three thousand frames through a headless browser.

A video edit, a title sequence and an original score are now things you can describe, listen to, reject, and iterate on in an afternoon, without opening a DAW. The gap between having taste and being able to act on it got a lot smaller this year.

The video, if you want to hear it: https://youtu.be/50XHOzKj9Uo

Minutehand is the app it is for. macOS menu bar, records meetings, files plain text Markdown transcripts into a folder your AI agent reads.

https://minutehand-app.com

Comment

About

All our context now lives inside our LLM of choice. And yet meeting tools insist on offering AI summaries based on single meetings. All Claude needs is the meeting transcript — so it can review it within that context.