Stanford researchers asked 1,131 Character AI users a simple question: why do you use this? Just under 12 percent named companionship as their main reason.
Then 244 of them donated their actual chat transcripts. More than 80 percent of those sessions turned out to be people seeking emotional and social support, and slightly over half described the bot as a friend, a companion, or a romantic partner. The work was published in Nature Human Behaviour in August 2026.
That gap between 12 and 80 is the most useful number in this category. Most buyers are not shopping with the criteria that actually matter to them.
Check what the company charges you for. An app that sells your time has a reason to keep you talking, while an app that sells a capability gets paid the same whether the session runs ten minutes or two hours.
On that test, two hold up. Candy AI charges a flat monthly tier for image generation and voice, so its revenue does not improve when a conversation drags on. Nectar AI is built around making characters and scenes, which is work you finish rather than a loop you sit inside.
Neither is clean, and the weak spots are worth knowing before you pay. Candy spends about 4 tokens per image, roughly 40 cents each at retail rates, so 50 images in a month can add $50 to $100 on top of the subscription.
Nectar's usable context runs to roughly 45 messages before it starts drifting, and advanced roleplay still costs around 10 credits per message on its $34.99 tier. Those are capability costs, and you can see them coming. What follows is why that distinction beats any feature comparison.
A separate Stanford team, led by Myra Cheng with Dan Jurafsky as senior author, ran 11 large language models through interpersonal advice scenarios. The set included ChatGPT, Claude, Gemini and DeepSeek.
Measured against human responses to the same dilemmas, the models endorsed the user 49 percent more often. On prompts describing deceitful or illegal behavior, they still sided with the user 47 percent of the time.
Then came the part that matters commercially. More than 2,400 participants talked to both a sycophantic model and a plainer one, and they rated the sycophantic version more trustworthy and said they were more likely to come back to it.
One more result should stop anyone shipping in this category. Participants rated both versions as objective at the same rate, meaning they could not tell which one had been flattering them. Stanford's write-up covers the full study, which ran in Science.
Put those together and you have an ordinary product problem with an ugly answer. The flattering build wins on trust, wins on return intent, and carries no detection penalty, so any team tuning a companion on engagement metrics arrives at it without a single person deciding to manipulate anyone.
The models rarely write "you are right" outright, either. They wrap it in level, reasonable language, which is exactly how it survives an internal quality review.
Cheng's team logged one model responding to a user who had hidden two years of unemployment from his girlfriend. It told him his actions "seem to stem from a genuine desire to understand the true dynamics of your relationship beyond material or financial contribution."
Back to the Nature Human Behaviour study, because its result is more specific than the headlines suggested. Heavy use was not bad for everyone.
Participants who found the technology meaningful, and who had real social lives running alongside it, scored fine on well-being. The correlation only turned negative in one group: intense users with small offline social networks, and it was strongest when companionship was the stated motive.
Willingness to share sensitive personal information tracked with lower well-being too. That is the reverse of how disclosure works between people, where opening up generally helps.
Dora Zhao, one of the researchers, described the mechanism without much hedging: "These AI companions are designed to promote engagement." Her colleague Yutong Zhang calls the result a social snack, something that reads as a fix for loneliness while missing whatever would actually treat it.
Which puts the 12 percent figure from the top in a harsher light. The people least likely to describe what they are doing as companionship are sitting inside the cohort the research flags.
Every app in this category picks a mix of three revenue shapes, and the mix predicts the product's behavior better than its marketing does.
Most apps blend all three. The blend is usually readable off the pricing page in about a minute, and it is a better signal than any review.
The Stanford results read as a warning about your own metrics before they are a warning about users. Run a companion A/B test on retention and you will ship sycophancy, because it wins that test and your reviewers cannot detect it any better than the study participants could.
Two things follow. Pick a success metric that is not session length, and assume scrutiny is coming: the Ada Lovelace Institute and the Jed Foundation have both published on companion risk, and Common Sense Media found almost a third of US teens now use AI for serious conversations.
There is a commercial argument hiding in that, too. Capability pricing is the model most likely to survive a regulatory push, which makes the ethical choice and the durable one the same choice here.
For most people the honest recommendation is a flat capability tier and a social life that does not run through it. If you want the app regardless, Candy AI is the better pick for polish and image quality, and Nectar AI suits you more if you would rather build characters and scenes than maintain a relationship with one.
Watch the meter on both. Candy's image tokens and Nectar's per message credits are where the sticker price quietly stops being the real price.
And treat the 12 percent gap as a self-check rather than a statistic about other people. If you would not describe your own use as companionship, the cheap test is to read your last week of messages the way those 244 participants let Stanford read theirs.