19
30 Comments

I Paid for 2,150 Survey Responses. Then I Threw 441 of Them Away.

I wanted to do some original research into how people actually think about AI companions.

There are already plenty of reviews and affiliate sites in this space. I was interested in something different: the people using these products.

Why do people use AI companions? What do they actually think about them? Where do they draw the line between an AI interaction and a real relationship?

My thinking was pretty simple: if I could understand even a little more about how people view these products, I could write better reviews and potentially provide more useful feedback to the developers building them.

So I decided to spend my own money on original research.

I paid for 2,150 responses

I commissioned a survey of 2,150 U.S. adults through SurveyMonkey Audience.

The questions covered things like AI companion usage, loneliness, romance, emotional attachment, privacy, and whether having a romantic relationship with an AI could count as cheating.

Then I got the data back.

At first, some of the results looked incredible.

And I don't necessarily mean that in a good way.

While analyzing the responses, I started noticing a pattern. When I looked at individual respondents, a large group seemed to be repeatedly selecting the first available option.

Again and again.

It looked less like 441 people who happened to have remarkably similar opinions and more like people trying to finish the survey as quickly as possible and get paid.

My first thought was basically:

I just wasted my money.

I had paid for honest answers. Apparently, some of the respondents didn't have quite the same idea.

I considered keeping them

I don't want to pretend I immediately took the high road and deleted them.

I actually considered keeping them.

After all, I paid for 2,150 responses.

And the unfiltered results gave me some crazy findings.

Before removing the suspicious responses, 30.3% of respondents appeared to be current AI companion users. More than a quarter, 27.1%, said they could genuinely fall in love with an AI.

Those are fantastic numbers if your goal is to write headlines.

And technically, I could just report what the survey returned. I wouldn't have to invent anything. I could say these were the results and leave it at that.

But I knew there was a problem.

The suspicious group wasn't subtle. Of those 441 responses, 98% came from Android devices. Their median completion time was only 39 seconds, and they arrived within a concentrated 2.5-hour window.

So eventually I accepted that keeping them would defeat the entire reason I commissioned the survey in the first place.

I wasn't trying to manufacture the most exciting statistics possible.

I was trying to understand this niche.

So I removed all 441.

And some of my best numbers disappeared

That left me with 1,709 responses.

Suddenly, the percentage of current AI companion users dropped from 30.3% to 12.3%.

The percentage saying they could genuinely fall in love with an AI dropped from 27.1% to 8.3%.

Those aren't small corrections.

The dataset told a substantially different story once the suspicious responses were removed.

It hurt a little. 1,709 is still a large sample, but I paid for 2,150. That's my hard-earned money sitting in those 441 discarded rows.

Eventually, though, I started looking at the loss differently.

Documenting what happened and explaining why I removed those responses probably adds more credibility to the research than pretending I never noticed the problem.

The research became more useful than I expected

I published the study through AI Girlfriend Coach, the site I'm building around AI companion testing and research.

Since then, I've been able to take the remaining data in quite a few directions.

The research has been picked up by other publications, and I've seen people use the findings to develop their own insights about AI relationships.

I also learned that original research doesn't have to replace the rest of what I'm doing.

I'm still reviewing AI companion platforms. But now I'm also giving specific feedback to developers, preparing research PDFs for companies that want them, and creating guides based on what I'm learning about the people actually using these products.

For me, those things complement each other.

Understanding the product is useful.

Understanding the user is useful.

Putting the two together is much more interesting.

What I'd do differently next time

I'd spend considerably more time planning before commissioning another survey.

Not just writing the questions, but thinking carefully about the target population, quality controls, what I actually want to learn, and how I'll evaluate the responses once they arrive.

My biggest lesson is probably that paying for original data doesn't automatically give you good original data.

And sometimes the most valuable thing you can do with data you've paid for is throw some of it away.

Has anyone else here commissioned surveys or other original research for a product? I'd be interested to hear how you handled response quality, especially when using third-party panels.

on September 11, 2026
  1. 4

    Survey quality control is a nightmare nowadays with automated bot/click farms. Adding attention-check questions upfront is definitely a lifesaver.

    1. 2

      Yeah, that’s what caught me off guard. At first the responses just looked like unusually strong results. It was only when I went into the individual rows that the pattern became obvious.

      1. 2

        That's exactly how it works. The aggregate numbers look clean until you go one level deeper. Surface-level stats can hide a lot. The individual row check is what separates honest research from a good-looking dashboard.

  2. 3

    I did something similar with my own data and realized that the worst kind of junk is the stuff that looks normal: I had a blank page returning a success result instead of an error... and a few other spots like that. As for the survey question: falling in love with AI isn't going to happen, it already drives me crazy as it is :)

    1. 1

      Exactly. The obvious bad data is probably the easy part. It's the stuff that looks completely normal in the aggregate that worries me more now.

  3. 2

    Thanks for showing both sets of numbers. As a reader, I can see how much the result changed after filtering, and why you made that choice. That explanation is useful alongside the final percentages.

  4. 2

    Ran into the same thing with event data - the weekly rollup looked clean until I opened the raw rows and saw the same session logged twice. What fixed it: a validation rule that flags valid-but-impossible rows before they hit the report. Good call deleting the 441 instead of keeping a tidier story.

  5. 2

    The 39-second median completion time is what makes the filter defensible. That detail alone justifies the cut. The bigger lesson here is that paying for data and getting good data are two different things. Documenting the filter publicly actually adds credibility rather than removing it. Most people would have kept the 441 and written the headline. The fact that you didn't is what makes the remaining 1,709 worth trusting.

  6. 2

    The 39-second median and the 2.5-hour window are the kind of detail that makes the cut defensible, so publishing the filter rule alongside the numbers is worth as much as the numbers.

    I had a smaller version of this with my own marketing data. My posts were getting reach that looked fine on the dashboard, and it would have been easy to report that as progress. When I sat down and counted the thing that actually mattered - posts that started a two-way conversation with someone in my target market - the count was zero out of 32. Writing that down honestly is what made me change what I was doing instead of posting more.

    One thing I'd add for anyone repeating this: decide the disqualification rules (straight-lining, minimum completion time) before you look at the results, so you're not choosing them after you've seen which way they move your headline.

  7. 2

    Discarding 441 responses is a useful reminder that volume can hide a broken sampling process. I’d keep a small audit table for every exclusion—speeding, duplicate patterns, failed attention checks, or contradictory answers—so the cleanup itself becomes evidence about acquisition quality. The next useful metric seems less like total completes and more like cost per decision-grade response by source.

  8. 2

    The money you spent on those 441 rows was actually the cost of discovering your measurement system wasn't measuring what you thought.

    You built a second measurement layer on top: device patterns, completion latency, temporal clustering, response uniformity. That's what actually caught the signal/noise boundary. The survey platform gave you data; your observation system gave you measurement clarity.

    This is why so many surveys fail quietly. Organizations measure "number of responses" while thinking they're measuring "quality signal." The uncomfortable part is that the most valuable metric you discovered wasn't in the dataset - it was "which responses should we ignore." That's a different measurement problem entirely, and it's the one that separates research from theater.

  9. 2

    Respect for tossing ~20% instead of forcing a clean story. First-option streaks look like consensus until you open the raw rows. One thing that helped me later: bake a couple of reverse-coded or "pick this exact answer" checks into the survey so the junk is easier to spot before you've already paid for the full batch. Still stings — you just catch it earlier.

    1. 1

      That's a really good idea. I definitely learned this one the expensive way 😅. I'll be building checks like these into the next survey instead of relying on finding the patterns afterward.

  10. 1

    The part worth stealing here is that you published the 441 and what removing them did to the numbers, rather than only the clean result.

    I had the opposite lesson on this site this week. I used a figure in a comment that I had never actually measured, someone else reasoned from it, and I had to go back and retract it publicly. Being wrong was survivable. The real cost was that anyone who had already built on the number had no way to know.

    One thing that would make your dataset harder to misquote: put the exclusion rule and the before and after figures together as plain text near the headline number, 30.3% against 12.3% with the reason sitting beside it. Otherwise the unfiltered version is the one that travels, because it is the more quotable number and there is nothing on the page arguing with it.

    1. 1

      That's a really good point, especially about the more quotable number being the one most likely to travel. I documented the exclusion in the methodology, but hadn't really thought about making the before/after context harder to separate when someone quotes it. That's something I'll look at improving.

  11. 1

    The 39-second median completion time on Android within a 2.5-hour window is the kind of quality signal most people would miss. The fact you caught it and removed those responses rather than publishing the 30.3% headline is the exact opposite of how most content in this space works.

    The cleaned 12.3% is also a more useful number for anyone building in this space. The 30.3% would have set wrong expectations for developers. The 12.3% tells them something they can actually build around.

    One thing that helped me when commissioning customer research: adding a timing threshold as a filter upfront (anyone under X seconds auto-excluded) plus a logic check question buried mid-survey that nonsensical answers can't pass. Still doesn't catch everything but it cuts the noise before it reaches your dataset.

    Did you find any pattern in which types of questions attracted the most speed-clicking? I'd expect the more sensitive questions (love, cheating) to have more genuine variance than the baseline demographics.

    1. 1

      I didn't see it concentrated around specific questions. What stood out was the response pattern across the survey itself, with the same option repeatedly selected. That's actually something I'd like to look at more closely next time though, especially whether certain question types produce more low-quality responses.

  12. 1

    The move almost nobody makes: go back to SurveyMonkey with exactly the evidence you laid out here, 98% Android, 39-second median, a 2.5-hour window, and demand replacement completes, because panels carry fraud terms and founders almost never invoke them. On the next one, bury two attention checks mid-survey with one reverse-scored and negotiate a minimum completion-time floor before they bill you. The methodology note is worth more than the 441 rows anyway, since that is the part other publications end up citing.

    1. 1

      I actually hadn't considered going back to SurveyMonkey with the evidence and asking about replacement completes. That's worth looking into, especially since the pattern was so concentrated. Thanks for suggesting it.

  13. 1

    Looking to help indie founders and gain hands-on experience!
    Hi everyone,I recently joined this community to learn how indie hackers think and build successful projects. I am highly motivated and looking for opportunities to help founders with their daily tasks, operations, or any manual work you need assistance with.My goal is to learn from your experience and add value to your business. If you have any work or small tasks available, please let me know. I am ready to jump in and help!

  14. 1

    The straight-lining catch is the part that sticks with me. Panel data rarely fails loudly — it just quietly inflates whatever sits in the first option slot. If "yes" was first on your AI-companion question, that 30.3% is doing a lot of heavy lifting.

    Did the 441 cluster on the long Likert grids, or spread evenly across every question? That's usually the tell between respondent laziness and a survey-design problem.

    Respect that you ate the cost instead of running with the inflated number. What convinced you the remaining sample was still trustworthy after cutting a fifth of it?

    1. 1

      That's a fair question. The main reason was that the 441 formed a very distinct cluster across several signals, not just one odd answer. Once they were removed, I didn't find the same combination of straight-lining, completion speed, device concentration, and timing in the remaining responses. That said, this experience definitely made me more cautious about treating any panel dataset as automatically clean.

  15. 1

    The expensive lesson is not only to add attention checks. It is to decide the exclusion rules before looking at how they change the result.

    When removing one cohort changes the headline, the filter becomes part of the finding. I would publish a small sensitivity table: the raw estimate, the estimate after each exclusion rule, and the number of rows each rule removes. That lets readers see whether the conclusion survives reasonable cleaning choices.

    It also protects the researcher from an uncomfortable bias: accepting a rule because it produces the more plausible or more interesting number. Data cleaning is strongest when another person can rerun the rules and reach the same dataset.

    1. 1

      This is a really good point. In my case, the exclusion rule came after I noticed the pattern, which isn't ideal even if the cluster itself was pretty striking. Predefining the rules before collecting the data is definitely something I'd change next time. I also like the sensitivity-table idea.

  16. 1

    Was this survey done through Prolific by any chance? I might be wrong, but I think I saw your research there and I may have even taken the test.

    The thing that came to my mind is that some platforms are starting to have a problem with people farming these tasks. I know some cases where people use AI tools to complete surveys and similar jobs.

    I don’t know if that’s what happened here, but the pattern you described (very fast completion + same answers + time clustering) made me think about it.

    1. 1

      It was actually through SurveyMonkey Audience, not Prolific. But what you're describing is definitely interesting. I can't say whether AI or task farming was involved in my case, so I don't want to speculate on the cause. What caught my attention was really the combination of the answer pattern, completion speed, and timing. That's what made me dig into the individual responses.

  17. 1

    The research clearly creates credibility, but the commercial test seems downstream. Have AI companion companies actually paid for the research or developer feedback yet, or is the strongest evidence still media pickup and audience interest?

    1. 1

      I actually see the commercial value a little differently. I didn't commission the survey as something to sell directly. The goal was to better understand the people using and thinking about AI companions, so that research can inform the reviews and guides I'm already producing.

      For me, the test is whether original research makes the site more useful and credible over time. The media pickup is an early positive signal, but it's probably too soon to know what the downstream commercial impact will be.

      1. 1

        That credibility-to-commercial-impact gap is interesting. If you’re open to it, what’s the best email to reach you on?