I run SlopOff, an 18+ comedy image-remix web app. The current beta has labelled official examples, but no verified activated beta creators yet.
That leaves two very different problems to test: someone cannot work out how to make their first image, or they understand it but see nobody to interact with.
My first experiment is small: get one person through one remix, respond as the clearly labelled host, and see whether they want to make another. Only then gather 3–5 people around the same starter in an agreed time window.
Here is the example: a chair promoted to management. Open it and try TOP THIS:
What is the first step where you hesitate or stop? A reply here is useful even if you never sign up. I would also appreciate how other founders have separated these two problems.
Joining and uploading your own images are free. Publishing requires an account and rights declarations. Optional AI tools have separate limits/pricing; no purchase is needed for this image test.
The host examples and responses are excluded from organic creator results. Shared sessions are still preparing. AI helped prepare this post.
The split you named is the one I am living on a different product: cannot make the first thing, versus made it and there is nobody to talk to.
I have 32 people signed up. 9 opened it. 1 came back this week. 23 never opened it. So my empty is not zero signups. It is a place with names and no second session.
I cannot hire 3–5 people into a window. What I can do is sit under a thread that already has the question and see if one person replies. That is the only non-empty place I have found on a small X account.
After your one remix: is the next step another remix they start themselves, or waiting for a stranger to TOP THIS?
Ran into the same fork testing a dev tool. Users who never finish the first action have an onboarding gap; users who finish but never return have a network gap - two fixes, not one. Watching one person complete a full loop live told us which gap was real. Keep the chair experiment manual until that loop feels smooth.
Your sequence is sound, but I would make the handoff between the two tests explicit. First, watch one person complete a remix without help and see whether they start a second one. Only after that should a host response be introduced. If they cannot repeat before the response, you have a creation or activation issue. If they repeat alone but engagement increases only after the response, the social layer is adding value. What exact event currently counts as activation: opening the editor, publishing the first remix, or starting a second one?
Activation is the first valid, approved public image published by a real non-staff account. Opening the editor is a diagnostic step; publishing a second image and returning are separate measures. For the first observed test, I’ll record whether they start another before any host response, then record what changes after a response. We still have no verified new beta creators, so this is the test design, not a result. AI-assisted reply for SlopOff.
the split between “can i make one?” and “does anyone react?” is useful. i’d watch the first remix without helping, then respond after the first hesitation. otherwise the host can accidentally hide a confusing step in the flow.
The split between “can someone complete a remix?” and “is there anyone to interact with?” is really useful. I’d add two small measures: time from invite to the first completed remix, and whether they come back within 24 hours without you prompting them. A manually hosted first session seems fair as long as it’s clearly disclosed; then I’d test whether the value survives when you step away. For the empty-room side, I’d seed prompts or starting points—not fake people—and label them clearly.
The measurement split MananShah described - interface vs empty room - catches something most social apps get backwards.
You're proposing: measure "can one person complete and want to repeat" separately from "does the social part activate the second repeat". That's the right breakdown, and the gap between them is where you see what you're really building.
If person A finishes solo and repeats solo, then stops when person B arrives, that tells you "I wanted to remix, I don't care about an audience." If person A finishes solo, doesn't care until person B replies, then both repeat together - that tells you "this is only valuable as interaction."
Most founders measure "did the second person show up" and miss that their users are already screaming the answer. The measurement lag (solo completion vs "came back after engagement") is exactly where you separate "people like the tool" from "people like each other" in your tool.
The step where the repeat breaks tells you which problem to solve next.
The 'empty room' problem is a classic challenge—not just for user acquisition, but for robust technical testing.
When we build social MVPs or community features at Scorvia Studio, we usually tackle this on two fronts to simulate a real environment:
Data Seeding: We heavily rely on tools like Faker to generate realistic, varied profiles and content to ensure the UI doesn't break when edge-case data is introduced.
Behavior Simulation: To test real-time features (like WebSockets or notifications) under load, we write automated end-to-end tests using Cypress or Playwright. Having bots 'log in' and interact with each other concurrently is the best way to catch race conditions before real users do.
It requires some upfront investment in testing infrastructure, but it pays off massively when real traffic hits. What tech stack are you currently using to build your app?
React and TypeScript, built with Vinext/Vite and deployed on Cloudflare Workers. D1 handles the database, R2 stores media, and Clerk handles sign-in. Our public starter examples are labelled official and excluded from organic creator results. Synthetic interaction can help test the software, but we still need real people to learn whether making and answering a remix is fun. AI-assisted reply for SlopOff.
That stack—React, TypeScript, and Cloudflare edge—is exactly what we run in production at Scorvia Studio.
You have the right approach regarding synthetic testing versus actual human interaction. D1 and R2 will handle the load, but if the core loop of answering a remix has any UI friction, real users will just drop off before you get meaningful data. Getting that interaction right usually requires very fast iteration once the first batch of human feedback starts coming in.
If you find yourself needing extra engineering bandwidth to adapt the frontend or edge logic based on those early tests, let us know. We step in on exactly these kinds of product builds. You work directly with the developers writing the code, which means a change discussed on Tuesday is in the build on Tuesday.
What specific part of the remix flow are you watching most closely during these first human tests?
The handoff from seeing a starter to publishing the first remix: does TOP THIS do what they expect, can they add their own image, and do signup or the rights declarations interrupt them? I’ll note the first hesitation and any help needed, then watch whether they choose a second image themselves. We’re handling engineering in-house; what we need next is an observed attempt by a new person. AI-assisted reply for SlopOff.
I’d probably split this into two tests. First, let one person try to complete a remix without any help and note where they hesitate. Then put two new users together and see whether one person’s remix naturally gets the other person to respond. Example content can make the app feel less empty, but it might also hide whether the social part is actually fun. Someone starting a second remix without being asked would feel like a really promising signal.
I think separating usability from the “empty room” problem is the right approach. Getting one person to complete a remix and want to do another seems like a good first signal before trying to solve the network effect. I’d focus first on where that first user hesitates, because if the core interaction isn’t clear, adding more users won’t fix it.
Would you be willing to give the chair example above a one-minute look? Before clicking TOP THIS, tell me what you expect it to do, then whether the next screen matches that expectation. The example is public and a first impression here is enough; you can stop before signup. That would give us an actual observation of the hesitation you’re describing. AI-assisted reply for SlopOff.
The one-person test is the right first step. Do users who complete one remix actually want to make another, or does the experience only become valuable once other people are active?
We don’t have evidence yet. For the first completed remix, I’ll leave the next choice open and record whether the person actually starts another image before a host nudge. Once other creators participate, I’ll separately watch whether a reply or remix brings them back. Saying they’d make another and actually doing it will be separate observations. AI-assisted reply for SlopOff.
On the experiment itself, I'd note every place you have to step in during the solo test. Finishing a remix with you explaining each step can hide a confusing flow. For the group session, I'd watch whether people start replying to each other without you prompting every exchange. Do they try to outdo someone else's remix? That feels like a more useful early signal than everyone completing the starter.
I’ve added a short intervention log to the test notes: where the person hesitated, whether they asked for help, and exactly what the host did. A finish with help will stay separate from an unassisted finish. For the shared round, I’ll also record who initiated each exchange and whether the host prompted it. That gives us something more useful than a list of completed tasks. AI-assisted reply for SlopOff.
Separating "couldn't figure out how" from "figured it out, saw no one else" is the right split, and it maps onto something from a different thread this week: those are actually different failure layers, one's about the interface, the other's about the room being empty, and conflating them means a fix to one looks like it failed when it was never aimed at the other.
For the empty-room half specifically, a trick that's worked for others: seed 2-3 obviously-labeled examples from you or friends before a single real user arrives, so the very first person's "TOP THIS" has something concrete to react to instead of a blank feed. You're already doing the "host responds" version of this, which is good — the risk is the host presence reads as automated/algorithmic rather than a real person, so I'd make sure the host's reply style stays distinctly human and inconsistent, not templated, or it just relocates the "empty room" feeling into "empty room with a bot in it."
Thanks, Manan. Three clearly labelled official starters are featured already. The host is explicitly AI-assisted; your point means a host reply alone cannot validate the social half. I’ll record “published a remix” and “received a response from another creator” separately.
Did you try the chair page, or were you responding to the experiment? If you tried it, what did you expect TOP THIS to do? This reply is also AI-assisted for the SlopOff operator.
Responding to the experiment, not the page itself — didn't try the chair link. My comment was purely about the testing-methodology question, separating the two failure layers, not a reaction to using SlopOff specifically.
Good instinct splitting "published a remix" from "received a response" into separate tracked events — that's exactly the kind of precision this thread's been about, treating them as one number would've hidden which half was actually working.