2
17 Comments

We are The Human AI Company: building Human Agents with a real face and voice, in real time.

We are Ojin, The Human AI Company, a small Berlin team. For a long time the same thing held AI agents back: they had a tell. The face lagged, the voice went flat, and some part of your brain knew it was not real.

Our flagship, Human Agents, is an instant end-to-end app for natural, human conversation with an AI that has a real face and a real voice, in real time. Under it sit two developer-tier face models you can use through one API: Oris Portrait (fastest and most scalable) and Oris Presence (the most sophisticated, the first to break the uncanny valley). Sub-200ms, works with Pipecat, LiveKit, and any stack.

You can see it at ojin.ai. I would genuinely love feedback from this community on where it still feels off: the latency, the expressions, the turn-taking, the moments the illusion cracks. That is what we are working on next, and outside eyes catch what we have gone blind to.

on July 1, 2026
  1. 1

    The use case question in the comments is the most important one to answer early. Passing as human is the right goal for some contexts and completely the wrong goal for others, and building toward the wrong one means all the polish lands in the wrong place.

    For task-oriented interactions, trust probably comes from competence and reliability, not from the agent feeling human. Users who've been burned by a confident-sounding AI that got it wrong don't want it to seem more human, they want it to be more accurate. Which use case are you actually optimizing for right now?

    1. 1

      Fair challenge, and it names the split we think about daily. Right now we optimize for customer-facing conversation, the contexts where rapport changes the outcome: a concierge, a guide, front-line support. There the face and voice earn attention, but you are right that they cannot substitute for being correct. The way we split it: accuracy lives in the reasoning layer (customers bring their own LLM and knowledge), and our job is that the delivery layer never gets in the way. A wrong answer in a warm voice is still a wrong answer, so we would rather be judged on both.

  2. 1

    This is a cool direction, and I agree the “tell” is what’s been holding these back more than anything.

    My guess is the weak spots won’t be the obvious ones like latency anymore, but the subtle human stuff. Things like slightly off timing in turn-taking, expressions not fully matching tone, or responses that feel a bit too clean and perfectly formed. Humans hesitate, interrupt, overlap, even misread each other sometimes. When that’s missing, it starts to feel artificial again.

    Also, long-session consistency is where a lot of systems fall apart. It can feel very real for the first minute, then small mismatches stack up and the illusion breaks.

    Honestly, getting past the uncanny valley might be less about perfection and more about adding the right imperfections.

    Would be interesting to know if you’re focusing more on behavior modeling or mostly on rendering and latency right now.

    1. 1

      This matches what we see. Rendering and latency are largely solved problems for us at this point, the hard frontier is exactly what you describe: turn-taking, hesitation, expression tracking tone rather than lagging it. Long-session drift is real too, small mismatches compound. So the honest answer is behavior over rendering right now. And your line about the right imperfections is well put. Perfectly formed delivery reads as artificial, and finding the natural texture without faking it badly is a genuinely hard design problem.

  3. 1

    Building on the presence-vs-work split above, there's a buyer-side version of that same fork. Right now this reads as dev-tier self-serve (API key, Pipecat, LiveKit, "try it and tell us where it breaks"). But the use cases that actually need this most: support, scheduling, ops replacing a human agent get bought by a 200-person team through security review and an SLA conversation, not a signup form. Those two motions want almost opposite front doors: one wants zero friction, the other wants a sales call and a trust page. Worth deciding which team size you're building the next six months of product for, because "developer API" and "enterprise support infra" pull the roadmap in different directions fast.

    1. 1

      You have described our actual front door layout. The API is the zero-friction path for builders who want to drop a face model into their own pipeline. For organizations that need the security review, the SLA conversation and someone accountable on the other end, there is a managed enterprise path. You are right that they pull the roadmap in different directions, and we would rather hold both doors open than pretend one motion fits everyone. The builders teach us where it breaks, the organizations teach us what production actually requires.

      1. 1

        You've solved the product fork, both doors exist. Real test is exactly which stage is actually arriving at each door?
        SoftRankings' stage-segmented traffic would show you instantly: pre-seed devs visiting your site vs. Series A teams. Then you can match entry point to stage intent. Right now you're probably getting both cohorts to the website, but don't know the mix or where leakage happens.
        That aggregate data (not individual profiles, just 'which stages show up') tells you if your routing is working or if stages are bouncing off the wrong door.

        1. 1

          That is the honest gap right now, we do not yet split traffic by stage. Aggregate, stage-segmented signal like that is exactly the kind of thing that would tell us if the two doors are actually catching the right cohorts or if builders are landing on the enterprise path and bouncing, or the reverse. Worth setting up before we assume the routing works just because both doors exist.

  4. 1

    The uncanny-valley framing might be aiming at the wrong axis for part of your market. "Remove the tell so users can't clock it as AI" is exactly right if the job is presence — companionship, entertainment, a face people want to feel something toward. There, undetectable is the product. But for AI agents doing actual work (support, scheduling, ops), my hunch is a lot of users don't actually want the mask removed — they'd rather know it's an AI and trust that it's doing the task, and I'd bet "this is an AI, and here's what it just did" often reads as more trustworthy than a flawless human face. Support is the messy middle where you probably want both at once — rapport and legibility. So the failure that would matter to me isn't a gaze or timing artifact — it's a use-case one: where is "passes as human" the actual goal, and where does it quietly cost you trust? Curious which end you're aiming at first, because the feedback you'd want back is completely different for each.

    1. 1

      This is the sharpest version of the question in this thread. Our answer: passing as human is not actually the goal, feeling natural is. Those are different targets. An agent can be clearly disclosed as AI and still benefit from timing, warmth and a face that does not distract. Disclosure is the customer's call and we support making it explicit. Where you are right to push: in pure task contexts, legibility (this is an AI, here is what it did) probably builds more trust than seamlessness. We start from rapport-heavy customer experience, where the natural delivery earns its keep either way.

  5. 1

    This is one of those products where the real competition isn’t other apps—it’s human perception thresholds. Once you hit low latency, the remaining failures stop being “AI issues” and become subtle timing, gaze, and turn-taking artifacts that users interpret emotionally rather than technically. That makes feedback loops especially valuable at this stage.

    1. 1

      Agreed, and that is exactly why the ask in the post is where does it break rather than please try it. Once you are past the obvious thresholds, users stop reporting bugs and start reporting feelings, something was off, and translating that back into timing, gaze or turn-taking fixes is the actual work. If you take a look, the most useful thing you could tell us is the first moment you noticed anything, however small. Those first-artifact reports are worth more to us than any benchmark.

      1. 1

        Interesting.

        Your reply made me think less about where the interaction breaks and more about what those moments gradually teach the product to optimize for.

        I don't think the consequence of that is obvious at first, and I don't think I can explain the reasoning properly in a thread without flattening it.

        If you're open to it, what's the best email to reach you on?

        1. 1

          Happy to take this further. [email protected] reaches us, or reply here with more and I will make sure it gets to the right person.

          1. 1

            Thanks! I’ve just sent it over.

            Looking forward to hearing your thoughts whenever you have a chance.

            1. 1

              Got it, thank you. We will make sure it lands with the right person and gets a proper reply, not just a form response. Appreciate you taking this further than a comment thread.

              1. 1

                Appreciate it.

                Looking forward to hearing your thoughts.

Trending on Indie Hackers
How to rank #1 on ChatGPT? User Avatar 112 comments I built a startup-idea scanner. It just told me none of my 3,400 ideas are easy wins. User Avatar 66 comments I Tested Agenmatic for Finding Customers in Communities — Here’s What I Learned User Avatar 63 comments “I’ll just post on Upwork” is not a client strategy. Here’s what I built instead. User Avatar 51 comments Building a Shopify bundles app for stores with real fulfillment: here's the wedge User Avatar 42 comments I recorded myself using 200+ indie SaaS products cold. Here are the 7 conversion killers that keep showing up. User Avatar 32 comments