I wanted to do some original research into how people actually think about AI companions.
There are already plenty of reviews and affiliate sites in this space. I was interested in something different: the people using these products.
Why do people use AI companions? What do they actually think about them? Where do they draw the line between an AI interaction and a real relationship?
My thinking was pretty simple: if I could understand even a little more about how people view these products, I could write better reviews and potentially provide more useful feedback to the developers building them.
So I decided to spend my own money on original research.
I paid for 2,150 responses
I commissioned a survey of 2,150 U.S. adults through SurveyMonkey Audience.
The questions covered things like AI companion usage, loneliness, romance, emotional attachment, privacy, and whether having a romantic relationship with an AI could count as cheating.
Then I got the data back.
At first, some of the results looked incredible.
And I don't necessarily mean that in a good way.
While analyzing the responses, I started noticing a pattern. When I looked at individual respondents, a large group seemed to be repeatedly selecting the first available option.
Again and again.
It looked less like 441 people who happened to have remarkably similar opinions and more like people trying to finish the survey as quickly as possible and get paid.
My first thought was basically:
I just wasted my money.
I had paid for honest answers. Apparently, some of the respondents didn't have quite the same idea.
I considered keeping them
I don't want to pretend I immediately took the high road and deleted them.
I actually considered keeping them.
After all, I paid for 2,150 responses.
And the unfiltered results gave me some crazy findings.
Before removing the suspicious responses, 30.3% of respondents appeared to be current AI companion users. More than a quarter, 27.1%, said they could genuinely fall in love with an AI.
Those are fantastic numbers if your goal is to write headlines.
And technically, I could just report what the survey returned. I wouldn't have to invent anything. I could say these were the results and leave it at that.
But I knew there was a problem.
The suspicious group wasn't subtle. Of those 441 responses, 98% came from Android devices. Their median completion time was only 39 seconds, and they arrived within a concentrated 2.5-hour window.
So eventually I accepted that keeping them would defeat the entire reason I commissioned the survey in the first place.
I wasn't trying to manufacture the most exciting statistics possible.
I was trying to understand this niche.
So I removed all 441.
And some of my best numbers disappeared
That left me with 1,709 responses.
Suddenly, the percentage of current AI companion users dropped from 30.3% to 12.3%.
The percentage saying they could genuinely fall in love with an AI dropped from 27.1% to 8.3%.
Those aren't small corrections.
The dataset told a substantially different story once the suspicious responses were removed.
It hurt a little. 1,709 is still a large sample, but I paid for 2,150. That's my hard-earned money sitting in those 441 discarded rows.
Eventually, though, I started looking at the loss differently.
Documenting what happened and explaining why I removed those responses probably adds more credibility to the research than pretending I never noticed the problem.
The research became more useful than I expected
I published the study through AI Girlfriend Coach, the site I'm building around AI companion testing and research.
Since then, I've been able to take the remaining data in quite a few directions.
The research has been picked up by other publications, and I've seen people use the findings to develop their own insights about AI relationships.
I also learned that original research doesn't have to replace the rest of what I'm doing.
I'm still reviewing AI companion platforms. But now I'm also giving specific feedback to developers, preparing research PDFs for companies that want them, and creating guides based on what I'm learning about the people actually using these products.
For me, those things complement each other.
Understanding the product is useful.
Understanding the user is useful.
Putting the two together is much more interesting.
What I'd do differently next time
I'd spend considerably more time planning before commissioning another survey.
Not just writing the questions, but thinking carefully about the target population, quality controls, what I actually want to learn, and how I'll evaluate the responses once they arrive.
My biggest lesson is probably that paying for original data doesn't automatically give you good original data.
And sometimes the most valuable thing you can do with data you've paid for is throw some of it away.
Has anyone else here commissioned surveys or other original research for a product? I'd be interested to hear how you handled response quality, especially when using third-party panels.
Survey quality control is a nightmare nowadays with automated bot/click farms. Adding attention-check questions upfront is definitely a lifesaver.
Yeah, that’s what caught me off guard. At first the responses just looked like unusually strong results. It was only when I went into the individual rows that the pattern became obvious.
Respect for tossing ~20% instead of forcing a clean story. First-option streaks look like consensus until you open the raw rows. One thing that helped me later: bake a couple of reverse-coded or "pick this exact answer" checks into the survey so the junk is easier to spot before you've already paid for the full batch. Still stings — you just catch it earlier.
That's a really good idea. I definitely learned this one the expensive way 😅. I'll be building checks like these into the next survey instead of relying on finding the patterns afterward.
The expensive lesson is not only to add attention checks. It is to decide the exclusion rules before looking at how they change the result.
When removing one cohort changes the headline, the filter becomes part of the finding. I would publish a small sensitivity table: the raw estimate, the estimate after each exclusion rule, and the number of rows each rule removes. That lets readers see whether the conclusion survives reasonable cleaning choices.
It also protects the researcher from an uncomfortable bias: accepting a rule because it produces the more plausible or more interesting number. Data cleaning is strongest when another person can rerun the rules and reach the same dataset.
The money you spent on those 441 rows was actually the cost of discovering your measurement system wasn't measuring what you thought.
You built a second measurement layer on top: device patterns, completion latency, temporal clustering, response uniformity. That's what actually caught the signal/noise boundary. The survey platform gave you data; your observation system gave you measurement clarity.
This is why so many surveys fail quietly. Organizations measure "number of responses" while thinking they're measuring "quality signal." The uncomfortable part is that the most valuable metric you discovered wasn't in the dataset - it was "which responses should we ignore." That's a different measurement problem entirely, and it's the one that separates research from theater.
Was this survey done through Prolific by any chance? I might be wrong, but I think I saw your research there and I may have even taken the test.
The thing that came to my mind is that some platforms are starting to have a problem with people farming these tasks. I know some cases where people use AI tools to complete surveys and similar jobs.
I don’t know if that’s what happened here, but the pattern you described (very fast completion + same answers + time clustering) made me think about it.
It was actually through SurveyMonkey Audience, not Prolific. But what you're describing is definitely interesting. I can't say whether AI or task farming was involved in my case, so I don't want to speculate on the cause. What caught my attention was really the combination of the answer pattern, completion speed, and timing. That's what made me dig into the individual responses.
The research clearly creates credibility, but the commercial test seems downstream. Have AI companion companies actually paid for the research or developer feedback yet, or is the strongest evidence still media pickup and audience interest?
I actually see the commercial value a little differently. I didn't commission the survey as something to sell directly. The goal was to better understand the people using and thinking about AI companions, so that research can inform the reviews and guides I'm already producing.
For me, the test is whether original research makes the site more useful and credible over time. The media pickup is an early positive signal, but it's probably too soon to know what the downstream commercial impact will be.
That credibility-to-commercial-impact gap is interesting. If you’re open to it, what’s the best email to reach you on?