We are two people from Slovakia who build AI room design software, and we have now put a photograph of somebody's real living room through our own pipeline more than a couple of million times. MeltFlex AI has passed 213,000 users, our biggest market is the United States rather than anywhere near us, and almost none of what made that work was the part I would have guessed on day one.
This is the post I wanted to read in 2024, when the category looked like an obvious wrapper play and the obvious conclusion was that it would be commoditised by Christmas. It half was. Here is what happened to the other half.
The demo converts. The product is a different company.
The demo is easy, and in AI room design it is the only thing that is. You upload a photo of a room, you get back a photorealistic room, people say wow, they share it. Every model release since has made that demo easier and cheaper to build, which is exactly why the category filled up with tools that are one prompt template deep.
What we learned the expensive way is that the wow is not the product. The wow is a coupon for about ninety seconds of attention, and what you do with those ninety seconds decides whether anyone comes back. Two things reliably converted that attention into retention for us, and neither of them is a model capability.
The first is that the room has to be theirs. The second is that the furniture has to be buyable. Everything hard we have done in two years sits under one of those two sentences.

One room, in and out of our pipeline. The balcony opening, the two ceiling fittings, the service line overhead, the kitchen run and the floor direction all survive. The light does not: nothing in the brief asked for a sunset, and that is a bug wearing a compliment.
Constraint one: it has to be their room
Hand a general image model a photograph of a kitchen and ask for something Scandinavian and it will hand you a beautiful Scandinavian kitchen. It will not be yours. The window will have moved, the run will have grown, the doorway will be somewhere more convenient. Users do not articulate this. They just quietly stop using you, because the image was pretty and useless, and their actual question was never "what does Scandinavian look like."
So the first real engineering constraint is that the photograph is a cage, not a mood. The geometry gets held while surfaces and contents get rebuilt inside it.
The useful part for anyone building here is how you test that, because you cannot eyeball it and you certainly cannot trust a cherry-picked before and after. What we run is an invariance check. Take one room, generate it under briefs that are as far apart as we can make them, then diff only the things a builder would have to be paid to move: window openings and their position on the wall, door side, ceiling line, any bulkhead or beam, the length of a fitted run, appliance positions, floor direction. Furniture is supposed to change. If anything in that list drifts while the style swings, the output is an invention wearing the room's approximate character, and it will fail the user at exactly the moment they try to make a decision with it.
Two categories still beat us, and I would rather say so here than have somebody discover it. Fitted joinery gets redesigned, because it behaves like architecture and looks like furniture, so it gets restyled along with the sofa. And rooms get relit without being asked, which sounds cosmetic and is not, because a north-facing room returned in golden hour is a sales pitch rather than a plan.
Constraint two: the boring moat is a furniture catalogue
This is the part I would push hardest to anyone evaluating AI room design as a business rather than as a weekend build.
A gorgeous render is free. It arrives with every model release, for everyone, including your competitors, including the incumbent design suite that adds it as a feature next quarter. There is no defensibility in the picture.
The defensibility is in what the picture connects to. Our renders resolve to actual products from retailers, with prices and links, so the output is a costed plan rather than a mood board. Building that is deeply unfun work, and that is precisely why it holds.
What it actually took, honestly:
A per-brand ingestion effort, because there is no such thing as a furniture data standard. Every retailer exposes its catalogue differently. Some have a clean listing API sitting behind their own storefront. Some only give you complete data on the product page. Several actively do not want to be read by machines, and for a couple of brands the only reliable path was a slow, well behaved crawl that respects the fact you are a guest. We now carry roughly a hundred products per brand across a widening set of them, and every new brand has been its own small project rather than a config file.
Photo normalisation, which I did not budget for at all. A catalogue assembled from ten retailers looks like ten catalogues: lifestyle shots, room sets, watermarked tiles, inconsistent crops, mixed backgrounds. Products only read as a coherent shopping list when they sit on the same white background at the same scale, so we ended up re-rendering something like seven hundred product images into consistent packshots using our own image pipeline. That is a weekend of scripting and then two weeks of failures, retries and manual review.
Matching, which is where the honesty has to live. We return the exact item where the catalogue holds it and a close relative where it does not, and we say so in the product rather than hoping nobody notices. Overclaiming here is the fastest way to lose a user permanently, because they find out at the checkout, not in the app.
Prices, currencies and availability, which rot continuously. A catalogue is not a dataset you acquire. It is a maintenance obligation you take on.

Our own workspace: every item in the render resolved back to a real product with a price. This is the part competitors can copy in principle and mostly do not, because it never finishes.
What we priced wrong
We launched cheap, because we were nervous, which is the standard founder illness.
The first correction was upward. Standard went from twenty to twenty nine, pro from forty to fifty nine, and we grandfathered every existing subscriber rather than migrating them. Nothing bad happened. Conversion did not collapse. In hindsight the low price was doing damage on both sides: it starved the margin we needed for GPU-heavy generation, and it signalled toy rather than tool to exactly the professional users who turned out to be our best retained cohort.
The second correction was sideways, and it mattered more. Our largest market is the United States and a meaningful share of our signups are not, and a flat euro price is a very different product in Manila than in Munich. We moved to regional pricing banded by country income group. If you sell a visual tool to homeowners and small studios globally, a single worldwide price is not neutral, it is a filter, and it filters out people who would have paid you something.
Third, we gated the wrong things for a while. Our free tier is deliberately thin, and for a period the shopping catalogue available to free users was narrower than the one paying users get. That is defensible as a business decision and it was bad product design, because the shoppable output is the thing that makes the product make sense. Cripple the demonstration of your differentiator and you are just another render toy in the free tier, which is where every user forms their opinion.
The distribution lesson: our own pages competed with each other
We are an SEO-heavy company. Two hundred plus articles, a landing page for each tool, an entire cluster around rooms, styles and materials. It works, and it produced two lessons I did not expect.
The first: for a large number of the AI room design queries we care about, Google prefers our blog post over our own tool page. We built dedicated landing pages targeting the head terms, wrote supporting articles around them, and then watched the article outrank the page it was supposed to support. Internal links did not fix it. What the query wanted was an article, and the algorithm was right about that. We eventually stopped fighting it and started treating the winning article as the landing page, with the tool sitting inside it, which converts perfectly well.
The second: a big slice of our impressions now come from queries no human types. Long, conversational, oddly specific, zero clicks, respectable position. That is retrieval traffic from assistants fanning out a user's question, not demand. If you read those numbers as an audience you will write content for an audience that does not exist. We now separate them before deciding anything.
Neither of these is an AI-era novelty exactly, but both changed how we spend a week.
Three surfaces we did not plan for
We built a consumer web app. Then usage kept arriving through doors we had not built.
Professionals wanted the engine, not the interface, so there is now an API, a command line tool and an MCP server, plus an extension that renders straight out of a 3D modelling viewport. None of that was a roadmap item. Each one existed because a specific paying user described a workflow our interface got in the way of.
Property developers turned out to be the cleanest business case in the whole company. Furnishing off-plan apartments for a sales gallery is traditionally a per-unit cost with a lead time measured in weeks. Several developers and agencies here now do it through us, and the appeal has nothing to do with the technology. It is that the finance director can put a number on it. Consumer AI tools are sold on delight. B2B is sold on a line item somebody was already paying, which is a much shorter conversation.
What we thought was hard versus what actually was
We expected
It turned out to be
Image quality
Solved by the model layer, and improving without us
Prompt engineering
Real but shallow, a few weeks of work, no defensibility
Keeping the user's actual room
Genuinely hard, and the thing that determines retention
Furniture catalogue and matching
The most expensive, least glamorous, most durable part
Pricing
Two full corrections and still not finished
Distribution
Our own pages competing, and traffic that is not people
The limits we publish on purpose
We say all of this in our own marketing, which sales people find perverse and which has cost us nothing measurable.
There are no reliable measurements. Nothing underneath is a measured model. It is an image conditioned on a photograph and it will quietly adjust a proportion to make a composition work, so nobody should order a sofa or brief a joiner off a render.
Built-in joinery gets redesigned, as above.
Rooms get relit unprompted, as above.
The quality selector's top label promises more than the exported file delivers, and the sharpest tier sits behind a paid plan. Invisible on a phone, obvious on a large monitor.
The reason for publishing that list is not virtue. In a category where every competitor's landing page claims the same six things, the specific admission is the only credible sentence on the page. It also front-loads the disappointment, which is much cheaper than a refund and a review.
If you are starting in this category now
Five things, compressed.
Assume the model layer is a commodity, because in AI room design it already is, and plan for the version of your product where the render is free and instant, because that version is arriving.
Find the constraint that makes the output true rather than pretty. Ours is that it has to be their actual room. Yours will be different and it will be the thing you cannot buy from a provider.
Pick the unfun asset early. Ours is a furniture catalogue. It compounds while you sleep and it is the reason a better funded competitor still has to spend a year to reach parity.
Charge more than feels comfortable, then charge differently in different places.
Say what your product cannot do, in your own words, before a reviewer says it in theirs.
We are two people in Bratislava selling mostly to Americans, in a category that everyone said would be commoditised. The commoditisation part was correct. It just happened to the part we were never going to win anyway.
The model changes under you, and nobody tells the user
One operational thing nobody warned us about. In a normal product, your software changes when you deploy. In this one, the layer doing the actual generation can change underneath you, and the first person to notice is a user whose kitchen came back wrong.
So we keep a fixed set of reference rooms, deliberately awkward ones, and re-run the whole set whenever anything in the generation path moves. Not to score beauty, which is unmeasurable, but to check the boring invariants: did the window stay, did the run stay, did the room keep its proportion, did the style brief still land. It takes a couple of hours to look through and it has caught regressions that no error rate would have shown, because nothing failed. It just quietly got worse.
If you are building on somebody else's model, budget for this. Your regression suite is not a set of assertions, it is a set of pictures somebody has to look at, and there is no way around the looking.