Earlier this year I had a working prototype that turned a text prompt into a marketing graphic. I made about 20, looked at them, and realized I wouldn't put a single one on my own account.
You know the look. A half-empty layout, a gray gradient in the background, a button drawn on that leads nowhere, and usually a number the model just invented. It screamed "a machine made this." So I deleted it and spent the next few weeks rebuilding from scratch.
My bar the whole time was uncomfortable and simple: would I actually post this on my own account? For most results the answer was no for a long time.
What changed the output:
Every design starts from a hand-made template instead of a blank canvas, so it looks designed, not random.
It reads your brand kit (colors, font, logo) and applies it automatically.
It checks its own text for contrast and legibility and fixes what fails, instead of shipping unreadable copy.
You refine by chatting ("make the headline bigger", "shorter") instead of opening design tools.
It's live at postmint.de. One sentence in, a finished post out, in Instagram / Story / Carousel / LinkedIn sizes. Free tier is 10 posts a month, no card.
I'd genuinely value feedback from this crowd on three things:
Run one generation. Does the result clear the "would you post this" bar, or where does it still look off?
Did anything in the flow make you hesitate before you even generated? (My worst number is people who sign up and never try it.)
The landing page and positioning: does it land, or what's confusing?
Tear it apart. I'd rather hear it here than not.
I like the standard you used for evaluating the output.
I'm curious whether "would I post this?" eventually became a measurable product constraint for you, or if you still think it's something that only human judgment can reliably answer.
The issue for me was, as I build mostly SaaS tools, posts always contain some graphs, or boxes with stats. AI generation wants to build a website as an image with buttons and standard CTAs and in my experience does not get the difference easily.
I build a testing environment where I logged my issues, compared the default AI results and the results from my tool and adjusted.
In short, it's always a human judgment, since we post for humans (for now), but you can automate a lot of the issues away.
Appreciate the context.
That distinction between what can be automated and what still requires judgment is the interesting challenge here.
Would be good to continue the conversation outside the thread.
What's the best email to reach you on?
Sorry for the late reply, you can shoot me a mail to me@tobias-schaefer .com
Thanks! I’ve just sent it over.
Looking forward to hearing your thoughts whenever you have a chance.