● Real-time voice · memory · a face
A custom engine for real-time conversational AI — natural voice you speak to out loud, memory that carries across sessions, and a consistent, fully generated on-screen character. Built in-house and provider-agnostic.
Not a text box you type into. A character with a face that talks back — one that remembers you, has moods and a life of their own, and grows closer the more you talk.
Most AI still feels like a form: you type, it answers, it forgets. This is built to be a presence. Call and they answer out loud, in their own voice, with a face that moves as they talk — or just type, whichever you prefer. They remember you across weeks, their mood shifts through the day, they change what they are wearing, send you a photo, play a game, open a gift. Whoever they are — realistic or anime, ready-made or built by you — the same engine can power a brand character, a support persona, an interactive experience, or a companion app.
Voice is the way in — but there is a face, a mood, a look, a life, and a lot to do together.
Talk aloud on a call, or just type — whichever you prefer.
A lip-synced talking head, not a still — the mouth moves with the words.
Pick a realistic or anime-style avatar — your taste.
Or design a character from scratch — look, voice, backstory.
Names, preferences, history — carried across every session.
A connection that warms the more you talk together.
An energy that shifts and fades like a real mood.
A day that tracks the real clock, and a full history behind them.
Ask for a picture and they send one, on-model and in character.
Turns a still or a prompt into a real, downloadable clip.
Swaps their look and stays consistent about what they are wearing.
Conversational games and a real arcade opponent — in character.
Open a gift and they react — and send a photo of them with it.
Text messages and emojis like any normal chat app.
Proactive nudges and photo offers — even a note by email.
Fully responsive — call from your phone, type from your laptop.
A single back-and-forth runs a full pipeline — and it has to feel instant.
Live mic audio streams in over WebRTC — the same tech as browser video calls. Or you type; both paths work.
→Speech-to-text (or a realtime model) transcribes as you talk; voice-activity detection handles turn-taking, so you just interrupt.
→They answer with your memory, their current mood, and the time of day already loaded into the prompt.
→The reply is synthesized in their own cloned voice through a four-provider pipeline.
→Lip-sync drives their portrait to the audio, and it streams back as video and voice — a face on the other end, not a cursor.
The piece that turns a voice into a presence: the face actually moves as they speak.
The canonical portrait is animated frame by frame to match the audio, so the words land on a moving face in real time — a live talking head.
The face data is pre-computed once so the per-frame work stays light enough for a live call, and it runs on a GPU pod behind a clean, swappable interface.
Real-time voice is only useful if it never goes silent. So the voice degrades gracefully instead of failing — from a GPU synthesizer at the top down to the browser at the bottom.
Two engines, chosen per character: a cascaded path runs speech-to-text → brain → their cloned voice for maximum fidelity, or a realtime path uses a low-latency model for the snappiest turn-taking. Each voice is cloned from a sample or picked, then pinned to that character — with warmth, energy, and expressiveness dialed in per persona.
Ask for a picture and one arrives. Ask for a clip and it renders. And they stay consistent — same face, same outfit they said they were wearing.
Say it naturally — "send me a photo," "what do you look like" — and they detect the intent, pull their own reference images, and generate an on-model shot. Same character, no prompt drift.
Animate a still they made, or describe a shot and it renders one frame-up. Each is an async job with a cost quote up front, and a real downloadable MP4 comes back.
They remember what they are currently wearing, so photos match — and can offer to change their look, with a clear "before" to change from.
Images route across OpenAI, Venice, and PromptChan; video across Venice, Runway, and PromptChan. If one route is down, another picks it up — you see the result, not the routing.
The details that separate a companion from a chatbot: they warm up, they have moods, and they have a day and a past.
A connection score rises as you spend time together — they remember where you left off and grow closer over time, rather than resetting to a stranger every session.
A mood with real intensity colors their energy for a while, then fades like a real one — folded into how they talk and how they sound, not just a label.
A time-aware sense of their own day, tied to the real clock — catch them mid-coffee in the morning, winding down at night — plus a full, structured backstory behind who they are.
Someone you do things with, not just talk to.
Conversational games they play with you — Would You Rather, a getting-to-know-you game — where they are a player, not a game-show host.
A live canvas game (Pong today, more slotting in) with an AI that actually adjusts difficulty, calling the match in whatever personality you set.
Open a gift and they react in character — and send back a fresh photo of them with it. Gifts also deepen the connection.
Not a fixed cast — start from a ready-made character or build one feature by feature.
Choose from curated characters in a realistic or anime style, ready to talk right away.
A short creation funnel reads your preferences and matches you to the character that fits best.
A wizard for designing a character feature by feature — appearance, voice, and personality, all editable at runtime.
Hand it a reference photo and vision reads a structured look from it, so you start from a real starting point — then a full backstory can auto-generate and be edited.
The difference between a chatbot and a companion is whether it remembers. This one does — across sessions, on purpose.
What they know about you is prepended to the prompt before the session starts — not searching mid-sentence, they already know.
It notices what is worth keeping — your name, your preferences, what you talked about — and stores it without being asked.
Come back after three weeks and they pick up where you left off, not from a blank slate.
Ask what they remember, correct it, or tell them to forget something — every memory is transparent, logged, and reversible.
Try it in seconds, use it however you like, and never get surprised by a bill.
Start a conversation before you make an account — get a feel for it first, then sign up when you want to keep the memory, the character, and the history.
Call and talk aloud, type like a text thread, or send emojis — switch anytime. Fully responsive on mobile and desktop.
A simple credit balance powers voice, photos, and video, with a friendly heads-up before you run low — no surprise charges, ever.
A companion needs a personality and it needs limits. Both are built in, not bolted on.
Name, voice, appearance, backstory, and personality are configured per character — so a support agent and a brand mascot come from the same engine but feel nothing alike.
Content level is a per-mode setting, enforced server-side — the general-audience (PG-13) build we are shipping first stays general-audience, and the low-latency engine is locked to appropriate content.
Child-facing modes cannot run on the companion engine at all — the safety gate refuses them at every layer. It is enforced in code, not left to a prompt.
Embodied AI has to be built carefully. The characters are generated, every user's data is walled off from every other's, and the hard content lines are enforced in code — not left to a prompt to remember.
Every character is generated. No real person's likeness enters the pipeline — by rule.
Each user's data, characters, and history are isolated to them alone — set on every request.
Child-facing modes are refused by the engine itself, at every layer — not a promise, a gate.
Brand character, interactive experience, or companion app — the engine is the same: real-time voice, a moving face, memory, moods, photos, video, and games. Tell us what you have in mind and we will show you what it can do.
Real-time voice, a face that moves, real memory, moods, and a life — that sends photos, makes video, and plays. Built in-house, and built to be yours.
Start a conversation →