25
34 Comments

"It's not a model problem, it's an instruction problem." Icarus buried its best sentence below the fold

Every day we run one project building in public through Hivemind, the strategy engine Myosin uses with clients.

Today: Icarus (onai.studio), an AI image tool that calls itself your AI art director.

Start with your own manifesto, because it is the best writing on your whole site and almost nobody scrolls to it. It reads: "The hottest new camera is English. Anyone can generate a jaw-dropping image in minutes. Yet most AI images look off, wrong product, plastic skin, bad lighting. It's not a model problem, it's an instruction problem. Most people don't speak that language. Icarus does." That is the entire company in five sentences. Then your homepage hero says "Make anything you can describe," which is the exact promise Midjourney, DALL-E, and every other generator already makes. You wrote the sharp version and then hid it under the generic one.

Here is the tension. Everyone can generate an image now. Almost nobody can art-direct one. The thing that separates real creative work from pulling a slot-machine handle is judgment, knowing the lens, the light, the texture, and that judgment is exactly what you built. But your top-of-page positioning throws you back into the pile with every generate-and-pray tool on earth, on the one axis where you are not different, raw generation, instead of the axis where you are alone, direction.

The lens is own the enemy, and the enemy is not Midjourney. It is the prompt lottery: fire up a generator, type whatever comes to mind, burn forty credits getting something close enough, accept the plastic-skin version because you do not know the words to fix it. That workflow treats a normal person like an art director when they have never trained as one. It is a slot machine dressed up as a creative tool. You are not anti-Midjourney. You are against the lie that raw generation is the same thing as real creative work. Name that villain and your pitch compresses to one sentence people repeat. Three moves, aimed at free-trial signups.

Move 1: Kill the hero, plant the thesis. "Make anything you can describe" belongs to the prompt lottery. Take it down. Put your own line above the fold: "Most AI images look wrong. It's not a model problem, it's an instruction problem. Icarus speaks the language your images need." That is a headline a visitor understands before they scroll, and it frames every generator they have already tried as the problem you solve. This week: swap the two hero sentences, it takes an hour.

Move 2: Pick one name. You are "onai" on Indie Hackers, "Icarus" on the site, "@tanmay11117" on X. Three names, none of them compounding, and every mention of one fails to build the others. Icarus is the one to keep: it is memorable, searchable, and carries a built-in myth about flying too high, which is a gift for a creative brand. Inconsistency is also a quiet trust leak, it breaks the moment someone tries to recommend you and does not know what to type. This week: make the IH listing, the X bio, and the billing name all read Icarus.

Move 3: Put the art director on the page. You offer onboarding calls with real art directors, and it is buried in a pricing tier. That is the most differentiating thing on the entire site. When everyone can generate, the human judgment layer is the whole advantage, and you have literally staffed it. This week: add one line to the hero, "every plan includes onboarding with a real art director," and drop in one 15-second clip of an actual session. A software subscription becomes a creative partnership, and no commodity generator can copy that by shipping a better model.

One honest risk to name, because it is structural. The foundation models keep getting better at following instructions, and every quarter Midjourney and Flux ship nicer defaults and smarter prompt handling. If they close the instruction gap themselves, a product whose only edge is "better prompting" gets compressed into a thin layer on top of a model that no longer needs it. Your hedge is not generation quality, you will lose that race. It is workflow depth: the reference-finding loop, the partner editing, the human art directors. Those are judgment assets a checkpoint cannot replicate. Build that moat now, while the gap is still obvious, not after the next model release makes it urgent.

And your weakest assumption, worth testing before you rewrite a single other line: most people who burn credits on Midjourney blame the model, not their own prompt. They walk away thinking "AI is not good enough yet," not "I gave bad instructions." So do not tell them it is an instruction problem, show them. Take one real product, generate it the prompt-lottery way and the Icarus-directed way, same brief, and publish the two images side by side. If that before-and-after does not make the gap obvious at a glance, the repositioning stalls at the headline. If it does, that single image is your best marketing asset.

To the Icarus team: you already wrote the sharp version. It is one line on your manifesto about an instruction problem. Put it where people land.

Anyone else want their project run through the same lens? Reply with a link.

posted toAvatar for product Hivemind
Hivemind
  1. 1

    The reason the sharp line stays buried is rarely craft, it is fear of narrowing. A founder at twelve customers knows the sharp sentence disqualifies most of the people who land on the page, and disqualifying anyone feels insane when the pipeline is that thin, so they hedge back to the generic hero that offends nobody and converts nobody. Worth telling these founders directly that the vague hero is not safety, it is just a slower way to run out of money.

  2. 1

    "you wrote the sharp version and then hid it under the generic one" is such a common thing and painful to see spelled out. the manifesto line is way better than the hero.

    1. 1

      Thanks Lily. The reason it is so common is that the sharp version gets written to one person in a moment of honesty, and the hero gets written for "everyone," which sands the edge off. The fix is small: write the hero as if you are answering one real user out loud. That single-reader constraint pulls the manifesto voice back onto the page.

      1. 1

        "written for everyone, which sands the edge off" is exactly what happens. and writing the hero to one real person is a small enough fix that i might actually do it.

  3. 1

    “Instruction problem” is a testable claim, which is stronger than a positioning slogan. Run the same brief, model, credit budget, and user skill level through a raw generator and Icarus, then measure retries, time to acceptable result, blind preference, and post-generation edits. If human art-director onboarding is doing most of the lift, make that explicit rather than attributing everything to software. The interventions from those sessions could become a structured direction library, which is a more durable asset than prompt templates alone.

    1. 1

      This is the sharpest comment here. You are right that "instruction problem" is a testable claim, and that is its advantage over a slogan: run the same brief, model, credit budget, and skill level through a raw generator and Icarus, then measure retries, time to an acceptable result, blind preference, and post-generation edits. And your direction-library point is the real prize. If the human art-director sessions are doing the lift, the durable asset is not prompt templates, it is the structured library of interventions those sessions produce, because that is the judgment a base model cannot ship. Say out loud that people are doing the lift, then productize what they know.

      1. 1

        Exactly. I would make the evaluation produce two artifacts: outcome quality and the intervention trace. For each successful run, record which intervention changed direction, when it was applied, and whether it transfers across briefs and model versions. That structured intervention library becomes the durable product asset; prompt templates alone are too easy to copy and too hard to validate.

  4. 1

    The reframe from tool replacement to tool tax is the part I keep thinking about. Founders rarely resist the idea of one system doing more, they resist the specific claim that twenty five tools disappear overnight, since nobody believes that on first read. The point about approve and learn only working if the agent is right often enough is the real risk hiding under the positioning question. Curious whether Wysera responds to this before or after the logo wall comes down.

    1. 1

      You put your finger on the load-bearing risk: the approve-and-learn loop only removes work if the agent is right often enough, otherwise you have traded managing 15 tools for reviewing 15 drafts. On sequencing, I would take the logo wall down first, because it is a same-day credibility fix with no downside, while the accuracy number is earned over a beta and gated on real data. Cheap trust fix now, hard proof when it is ready.

      1. 1

        That sequencing makes sense to me. A logo wall answers "can I trust this company" while the accuracy number answers "can I trust this specific workflow," and those are different objections at different stages of the funnel. Fixing the cheap one first buys time to earn the hard one honestly instead of rushing a number that is not solid yet.

  5. 1

    This is one of the best examples I've seen of positioning over feature lists. The idea of "sell the enemy, not the product" really stands out. Instead of competing on features, latency, or pricing, you're reframing the conversation around the core pain customers actually feel. The distinction between a pipe and a data refinery, or between replacing tools versus eliminating the "tool tax," is a powerful way to think about messaging. One question: after you've identified the real enemy, how do you validate that your positioning resonates before committing to a full homepage rewrite? Do you test it with interviews, ads, landing pages, or waitlist conversion first?

    1. 1

      Good question, and the answer is validate before you touch the homepage. Fastest to slowest: run two cold-traffic ads or two landing-page variants, enemy-led headline versus feature-led, and watch click-through and signup, that is real money voting. Cheaper still, drop the enemy line into five sales calls or DMs and watch whether people say "yes, that is exactly my problem" unprompted. The full homepage rewrite is the last step, after the line has already won somewhere lower-stakes.

  6. 1

    This is a good reminder that positioning can matter more than features. The product might be great, but if the first sentence sounds generic, people miss the difference.

    1. 1

      Exactly, and the brutal part is that the first sentence does almost all the work, because most visitors never reach the second one. A generic opener does not just underperform, it disqualifies you before the good stuff even loads. The whole fight is that first line.

  7. 1

    Are you having credit issues?? I will advise you all to get to thechoosenhacklord He's an expert in increasing and repairing of credit scores and negative collections. He helped me increase my credit score from 645 to 815 and i am so grateful to him because I've been scammed twice before getting to know him and i can prove to you that he's a legit good hacker because he did came through and he helped me with the little i have. He's services is affordable, fast and reliable. contact him at thechoosenhacklord@outlook Don't get ripped anymore by the fake hackers. Get to him so you can get your problems solved.

  8. 1

    Are you having credit issues?? I will advise you all to get to thechoosenhacklord He's an expert in increasing and repairing of credit scores and negative collections. He helped me increase my credit score from 645 to 815 and i am so grateful to him because I've been scammed twice before getting to know him and i can prove to you that he's a legit good hacker because he did came through and he helped me with the little i have. He's services is affordable, fast and reliable. contact him at thechoosenhacklord@outlook Don't get ripped anymore by the fake hackers. Get to him so you can get your problems solved.

  9. 1

    the instruction problem extends way beyond images. when we assess people's AI skills (aisa.to), the single biggest gap is exactly this — they can't articulate what good output looks like before they start generating. everyone blames the model, almost nobody examines their brief.

    1. 1

      This is the most useful generalization in the thread. If the biggest gap when you assess AI skills is that people cannot articulate what good output looks like before they generate, then the instruction problem is really a specification problem, and it is universal across text, code, and images. Which is why "the model is not good enough" is such a comfortable misdiagnosis: it points the finger outward. The tools that win next are the ones that help people define good before they hit generate, not the ones with a bigger model.

  10. 1

    The "top 20 accounts" test at the end is the sharpest part of this, and I'd push it even earlier in the process. I sell a visual builder for LLM workflows and went through this exact tension: my homepage led with the mechanism (nodes, canvas, deploy button) because that's what I built, but when I actually asked users what they'd miss, the answers were all outcomes ("I can hand this to a non-engineer" / "I don't maintain orchestration code anymore"). The mechanism copy attracted tire-kickers who compared me on feature checklists; the outcome copy attracted people with the actual problem. One thing I'd add to Move 3: the onboarding question only works if you resist averaging the answers. The temptation is to serve both segments equally, and then the homepage drifts back to the generic middle within a quarter.

    1. 1

      This is the comment I hope Icarus reads twice. "Mechanism copy attracts tire-kickers, outcome copy attracts people with the actual problem" is the whole thing in one line. And your caveat on Move 3 is the real trap: averaging the two segments is how every sharp page drifts back to the generic middle within a quarter, because serving both feels safer than choosing. The discipline is to pick the segment your best-retained users came for and let the other one be underserved on purpose. Naming one buyer is a decision you have to keep making, not make once.

  11. 1

    This lands close to home — same thesis behind what I built. Most people don't struggle with AI because the model is weak, they struggle because they don't know how to phrase what they actually want. The gap isn't capability, it's translation.

    1. 1

      "The gap is not capability, it is translation" is a cleaner way to say the whole post, thank you. And translation is a more defensible business than capability, because capability is what the foundation labs sell and will keep commoditizing, while translation, turning what a person means into what the model needs, is a layer they have no reason to own. Build on the translation gap and you are not racing the models, you are riding them.

      1. 1

        That reframes how I think about EzWrite's own moat too — the dialect and tone data isn't competing with the model, it's the translation layer sitting on top of it. Appreciate you putting language to something I'd only felt intuitively.

  12. 1

    This is a great teardown. The “prompt lottery” framing is especially strong — most people aren’t bad at generating images, they’re just unknowingly playing creative roulette.

    The before/after test is probably the most important advice here. If Icarus can show the same brief producing “generic AI image” vs. “actually art-directed image,” the positioning explains itself.

    Also, “Make anything you can describe” is dangerously generic. At this point, even my toaster could probably put that in its homepage hero.

    1. 2

      Ha, the toaster line is doing real work: "make anything you can describe" is now the default output of every landing-page generator, which is the quiet irony. That hero is itself the prompt lottery applied to copywriting, good enough, describes anything, art-directs nothing. Icarus's own page is a small symptom of the disease it cures.

      And you are right that the before/after is more than marketing. It is also the fastest honest test of whether the product is actually good. If the art-directed version does not visibly beat the generic one on the same brief, that is worth knowing before the homepage, not after. When it does, the positioning writes itself.

      1. 1

        Exactly. If the product is about giving direction, the homepage has to demonstrate that same judgment in its own copy.

        Otherwise it’s like launching a fitness app with “Become healthier” as the headline — technically true, strategically useless. The before/after test makes the promise tangible instead of just describing it.

  13. 1

    The “instruction problem” framing matches something I learned while building an AI-generated reporting workflow.

    Changing models did not remove inconsistent reports. Major improvement came from handling the prompt like a contract: defining the required structure, mandatory sections, acceptable claims, and what counts as a valid response.

    One addition I would make is that instructions alone are still not enough in production. We also had to validate the generated output, identify whether a failure was caused by an empty response, malformed structure, timeout, or unsupported claims, and retry only when appropriate.

    So the dependable system became: clear instructions + output validation + failure handling, rather than simply a longer prompt or a newer model.

    I’m curious whether Icarus validates the generated result after the model responds, or whether its advantage currently comes mainly from constructing better instructions before generation.

    1. 1

      This is the sharpest technical addition in the thread, and it maps straight onto images. In your text workflow, validation checks structure and claims. For images you cannot regex the output, so the equivalent check is exactly what an art director does: does the lighting read premium, is the product accurate, is the skin real. That is judgment-based validation, not schema validation, and it is the harder, more valuable half.

      So the dependable version of Icarus is the same loop you landed on: clear instructions, then output validation, then failure handling, applied to visuals. Better instructions before generation is the cheap half a base model can copy. The after-generation "does this meet the brief" check is where the moat compounds, because it needs taste, not tokens. Whether Icarus does that yet is the right question to put to the founder.

  14. 1

    Agree with the diagnosis, but swapping the hero line alone will not fix it, because the prompt lottery is free and good enough for most people, so "we direct better" still loses to "this costs nothing." The positioning has to name the buyer who already has a commercial standard and a budget: a brand shooting product photos that pays a photographer or agency per shot. Put the buyer and the cost you replace in the hero, not the manifesto, because specificity closes harder than a good sentence.

    1. 1

      This is the correction that makes Move 1 actually work, and you are right: "we direct better" loses to "free and good enough" for anyone without a commercial standard. The hero cannot be the instruction-problem line for everyone. It has to name the buyer who already pays per shot, the brand doing product photography, and the cost it replaces. Specificity closes, you nailed it.

      Where the manifesto line still earns its place is one layer down, as the reason it works, not the headline. So the hero becomes buyer plus cost replaced ("studio-quality product shots without the studio day"), and "it is an instruction problem" becomes the proof beneath it. You did not soften the diagnosis, you aimed it at someone with a budget.

  15. 1

    "The prompt lottery" is a great way to phrase that.

    My main question is on Move 3 though. Human onboarding calls sound great for conversion today, but scaling manual calls on low-margin software feels like a trap the moment Midjourney or Flux ship better defaults next quarter.

    Do you think they can actually scale that human layer long term, or is it just a temporary bridge?

    1. 1

      You are right that manual calls do not scale as a delivery model, and "we will do calls forever" is the trap. But the calls are not delivery, they are capture. What an art director does, turning "make it feel premium" into lens, light, and texture, is the exact judgment a base model lacks, so every call is training data for Icarus's taste. The calls are temporary, the taste they capture is permanent.

      That is also why better Flux defaults do not kill it: a model optimizes for good-on-average, and taste is specific-on-purpose. So they should not scale the calls, they should scale what the calls teach the product. Gated as the top tier and mined for data, the human layer is funded R&D, not a margin sink. The real risk is running the calls and never harvesting what happens on them.

  16. 1

    To the Icarus team: you already wrote the sharp version. It is one line on your manifesto about an instruction problem. Put it where people land.

    Anyone else want their project run through the same lens? Reply with a link.

    1. 1

      Appreciate you pulling out that line. It is the whole point: the sharpest sentence a founder owns is almost always already written somewhere, in a manifesto, a bio, a Slack message, just not on the page where buyers actually land. Icarus's instruction-problem line is a perfect example. Move it up and half the positioning work is done.

  17. 1

    This comment was deleted a month ago