Work/Real-Time AI Companion
Real-Time AI Companion

An AI you can actually talk to.

Real-time voice · memory · a face

A custom engine for real-time conversational AI — natural voice you speak to out loud, memory that carries across sessions, and a consistent, fully generated on-screen character. Built in-house and provider-agnostic.

What it is
Real-time conversational AI
Type
Voice + memory + generated character
Built on
WebRTCmulti-provider
Status
In beta
LIVE VOICE SESSIONLISTENING
Ada · a generated character
Real-time voiceRemembers youSynthetic character
— CONVERSATION —
YouMorning — did I mention my trip next week?
AdaYou did — Thursday, right? Want a reminder?
Turn-taking~300ms · natural, interruptible
VoiceWebRTC realtime
Memorypersistent, cross-session
Characterconsistent · synthetic-only
// the idea

Not a text box you type into. A character with a face that talks back — one that remembers you, has moods and a life of their own, and grows closer the more you talk.

Most AI still feels like a form: you type, it answers, it forgets. This is built to be a presence. Call and they answer out loud, in their own voice, with a face that moves as they talk — or just type, whichever you prefer. They remember you across weeks, their mood shifts through the day, they change what they are wearing, send you a photo, play a game, open a gift. Whoever they are — realistic or anime, ready-made or built by you — the same engine can power a brand character, a support persona, an interactive experience, or a companion app.

// what it can do

It doesn't just talk.

Voice is the way in — but there is a face, a mood, a look, a life, and a lot to do together.

📞

Call or text

Talk aloud on a call, or just type — whichever you prefer.

🙂

A face that moves

A lip-synced talking head, not a still — the mouth moves with the words.

🎨

Real or anime

Pick a realistic or anime-style avatar — your taste.

Build your own

Or design a character from scratch — look, voice, backstory.

🧠

Remembers you

Names, preferences, history — carried across every session.

💗

Grows closer

A connection that warms the more you talk together.

🎭

Has moods

An energy that shifts and fades like a real mood.

A life + backstory

A day that tracks the real clock, and a full history behind them.

📸

Sends photos

Ask for a picture and they send one, on-model and in character.

🎬

Makes videos

Turns a still or a prompt into a real, downloadable clip.

👗

Changes outfits

Swaps their look and stays consistent about what they are wearing.

🎮

Plays games

Conversational games and a real arcade opponent — in character.

🎁

Gifts + a photo

Open a gift and they react — and send a photo of them with it.

😊

Send emojis

Text messages and emojis like any normal chat app.

📬

Reaches out

Proactive nudges and photo offers — even a note by email.

📱

Any device

Fully responsive — call from your phone, type from your laptop.

// one turn, end to end

What happens when you speak.

A single back-and-forth runs a full pipeline — and it has to feel instant.

01 · you

You speak

Live mic audio streams in over WebRTC — the same tech as browser video calls. Or you type; both paths work.

02 · hear

It hears

Speech-to-text (or a realtime model) transcribes as you talk; voice-activity detection handles turn-taking, so you just interrupt.

03 · think

It thinks

They answer with your memory, their current mood, and the time of day already loaded into the prompt.

04 · voice

It speaks

The reply is synthesized in their own cloned voice through a four-provider pipeline.

05 · face

The face moves

Lip-sync drives their portrait to the audio, and it streams back as video and voice — a face on the other end, not a cursor.

// the face

A talking head, not a still.

The piece that turns a voice into a presence: the face actually moves as they speak.

Real-time lip-sync

The canonical portrait is animated frame by frame to match the audio, so the words land on a moving face in real time — a live talking head.

Built to stay fast

The face data is pre-computed once so the per-frame work stays light enough for a live call, and it runs on a GPU pod behind a clean, swappable interface.

// the voice pipeline

Four voices, one fallback chain.

Real-time voice is only useful if it never goes silent. So the voice degrades gracefully instead of failing — from a GPU synthesizer at the top down to the browser at the bottom.

1
GPU synthesis — Chatterbox TurboRuns on a RunPod GPU instance: fast, high-quality, and the home of the cloned voice. The default when it is up.
first choice
2
Self-hosted — CoquiNo per-call cost and no external dependency, with built-in and custom voices. Takes over if the GPU tier is unavailable.
no API cost
3
API — ElevenLabsPremium synthesis over an API for when the self-hosted options are down — per-call cost, exceptional quality.
premium fallback
4
Browser — Web Speech APIZero latency, zero cost, no server. Lower fidelity, but there is always a voice — even when every external service is unreachable.
always available

Two engines, chosen per character: a cascaded path runs speech-to-text → brain → their cloned voice for maximum fidelity, or a realtime path uses a low-latency model for the snappiest turn-taking. Each voice is cloned from a sample or picked, then pinned to that character — with warmth, energy, and expressiveness dialed in per persona.

// photos, video & wardrobe

They can show you, not just tell you.

Ask for a picture and one arrives. Ask for a clip and it renders. And they stay consistent — same face, same outfit they said they were wearing.

Sends a photo on request

Say it naturally — "send me a photo," "what do you look like" — and they detect the intent, pull their own reference images, and generate an on-model shot. Same character, no prompt drift.

Image-to-video and text-to-video

Animate a still they made, or describe a shot and it renders one frame-up. Each is an async job with a cost quote up front, and a real downloadable MP4 comes back.

Outfits that stay consistent

They remember what they are currently wearing, so photos match — and can offer to change their look, with a clear "before" to change from.

Multi-provider, invisibly

Images route across OpenAI, Venice, and PromptChan; video across Venice, Runway, and PromptChan. If one route is down, another picks it up — you see the result, not the routing.

Photos on requestImage-to-videoText-to-videoMP4 outWardrobe continuityConsistent avatarMulti-provider
// it feels like a person

Presence, not just replies.

The details that separate a companion from a chatbot: they warm up, they have moods, and they have a day and a past.

affinity

They warm up to you

A connection score rises as you spend time together — they remember where you left off and grow closer over time, rather than resetting to a stranger every session.

mood

They have moods

A mood with real intensity colors their energy for a while, then fades like a real one — folded into how they talk and how they sound, not just a label.

life + backstory

They have a life

A time-aware sense of their own day, tied to the real clock — catch them mid-coffee in the morning, winding down at night — plus a full, structured backstory behind who they are.

// games & gifts

Things to do together.

Someone you do things with, not just talk to.

games with them

They play along

Conversational games they play with you — Would You Rather, a getting-to-know-you game — where they are a player, not a game-show host.

the arcade

A real opponent

A live canvas game (Pong today, more slotting in) with an AI that actually adjusts difficulty, calling the match in whatever personality you set.

gifts

Give them something

Open a gift and they react in character — and send back a fresh photo of them with it. Gifts also deepen the connection.

// pick one, or make your own

Choose a character, or design one.

Not a fixed cast — start from a ready-made character or build one feature by feature.

Ready-made, realistic or anime

Choose from curated characters in a realistic or anime style, ready to talk right away.

Matched to your taste

A short creation funnel reads your preferences and matches you to the character that fits best.

Or build your own

A wizard for designing a character feature by feature — appearance, voice, and personality, all editable at runtime.

Autofill from a photo

Hand it a reference photo and vision reads a structured look from it, so you start from a real starting point — then a full backstory can auto-generate and be edited.

// memory

They know you before you speak.

The difference between a chatbot and a companion is whether it remembers. This one does — across sessions, on purpose.

Loaded before they wake up

What they know about you is prepended to the prompt before the session starts — not searching mid-sentence, they already know.

Remembers on its own

It notices what is worth keeping — your name, your preferences, what you talked about — and stores it without being asked.

Survives weeks away

Come back after three weeks and they pick up where you left off, not from a blank slate.

You stay in control

Ask what they remember, correct it, or tell them to forget something — every memory is transparent, logged, and reversible.

// how you use it

Made to be easy to live with.

Try it in seconds, use it however you like, and never get surprised by a bill.

no barrier

Chat before you sign up

Start a conversation before you make an account — get a feel for it first, then sign up when you want to keep the memory, the character, and the history.

your way

Voice, text, or emoji

Call and talk aloud, type like a text thread, or send emojis — switch anytime. Fully responsive on mobile and desktop.

clear pricing

Runs on credits

A simple credit balance powers voice, photos, and video, with a friendly heads-up before you run low — no surprise charges, ever.

// persona & guardrails

You define who they are — and who they are not.

A companion needs a personality and it needs limits. Both are built in, not bolted on.

persona

Voice, look, and manner

Name, voice, appearance, backstory, and personality are configured per character — so a support agent and a brand mascot come from the same engine but feel nothing alike.

content controls

Gated by mode

Content level is a per-mode setting, enforced server-side — the general-audience (PG-13) build we are shipping first stays general-audience, and the low-latency engine is locked to appropriate content.

child safety

Walled off, by design

Child-facing modes cannot run on the companion engine at all — the safety gate refuses them at every layer. It is enforced in code, not left to a prompt.

// built responsibly

Synthetic, private, and gated.

Embodied AI has to be built carefully. The characters are generated, every user's data is walled off from every other's, and the hard content lines are enforced in code — not left to a prompt to remember.

01

Synthetic subjects only

Every character is generated. No real person's likeness enters the pipeline — by rule.

02

Scoped per user

Each user's data, characters, and history are isolated to them alone — set on every request.

03

Safety in code

Child-facing modes are refused by the engine itself, at every layer — not a promise, a gate.

// where it stands

In active development.

In build
actively developed and bug-tested
Real-time
live voice, a moving face, photos, video, and games
Growing
new capabilities landing regularly
// interested?

Want a talking AI character of your own?

Brand character, interactive experience, or companion app — the engine is the same: real-time voice, a moving face, memory, moods, photos, video, and games. Tell us what you have in mind and we will show you what it can do.

Talk to us →
Custom · provider-agnostic

A companion you can actually talk to.

Real-time voice, a face that moves, real memory, moods, and a life — that sends photos, makes video, and plays. Built in-house, and built to be yours.

Start a conversation →
See all our platforms →
Real-Time AI Companion — a Stark Create build.