Two years ago, this category was a text box. You typed, it replied, and whatever happened next existed because both of you put something into it. That's basically it. Now the same apps ship short clips, templated motion effects, and serialized "episodes" complete with the little "NEW EP" badge streaming apps use to tell you something fresh just dropped. Nobody's actually written that shift down at the category level. It happened slowly enough that each step looked minor at the time, and none of these companies were exactly advertising "we just turned your girlfriend into a TV show."

We've spent close to two years running hands-on reviews across this category. The same handful of features keep showing up in the same order, app after app: text, then images, then voice, then short clips, then templated motion, then, most recently, full serialized episodes. Every app in the space is racing toward video right now on a shared assumption nobody's actually tested out loud, that more production value means more intimacy. Those two things aren't the same. The category is betting a lot of engineering time on the idea that they are, and we think that bet deserves a closer look than it's gotten.
Where it started: a co-authored text box
Replika, the category's oldest surviving name, was released in November 2017 as a plain text-based companion — a product that, as we covered in the story of how Replika grew out of a memorial chatbot, was never designed to be a media platform. You wrote, it learned your patterns, and the relationship got built entirely out of exchanged messages. Nothing else. Character.AI was founded in November 2021 and opened its public beta in September 2022, same basic setup, chat as the whole product, dressed up with user-created personas instead of one fixed companion.
That setup mattered more than it looked like at the time. In a text exchange, the user supplies half of everything: the pacing, the specifics, the tone. A companion app's explicit content, back when this was the whole category, worked the way scenes in tabletop roleplay or fan fiction always have: the app proposed, the user directed, and the result belonged to neither party alone. Nobody watched a scene happen. They wrote it, turn by turn, together, and a bad line could get edited or steered somewhere else before it ever became a thing that happened.
Easy to write that off now. A plain text interface looks almost primitive next to what shipped this year. But the primitiveness was kind of the whole point, honestly. There wasn't much for the app to hide behind. If a companion felt hollow, you noticed right away, because the whole exchange was made of things you'd typed and things it had typed back, no filter, no gloss.
Voice and images: presence, not yet performance
The first real addition wasn't video. It was voice, and it showed up on Replika first, alongside the AR camera mode that still anchors the app's 3D avatar experience today. The timing tracked with demand: Replika added half a million users in April 2020, its fastest growth period on record, as isolated people stuck in lockdown reached for anything that felt like company. Character.AI took a different order: mobile before vocal. The app launched on iOS and Android in May 2023, pulling in over 1.7 million downloads in its first week, and character voices didn't reach general availability until March 2024.
Image generation arrived around the same window, mostly pitched as a keepsake feature. Turn a mood or a memory into a picture you could keep. That's still roughly how Replika's own app copy describes it, actually. None of this changed the underlying deal yet, though. A voice reading back what you typed, or an image generated from a prompt you wrote, was still a response to something the user handed it first. Voice made the companion feel present in a way plain text never quite managed, and a picture gave people something to hold onto after the conversation ended. But neither one moved the authorship anywhere. You were still the reason the scene existed.
The video pivot: from written scene to generated clip
That changed once video generation moved inside the chat window itself. Apps like Yollo AI now put companion chat and a video generator side by side. Describe a scene in text and the app renders a short clip, or animate a still image into motion. Yollo markets this outright as an uncensored video tool, built for exactly the kind of explicit content that used to exist only as typed scenes, now available as generated footage with the character's chat memory attached to it.
More than voice, more than images, this is where things actually turned. A typed scene needs the user's own words to exist at all. A generated clip needs a prompt too, sure, but the output itself, the pacing, the framing, the specific thing that happens on screen, comes from the model and whatever it was trained on. Not from anything the user wrote sentence by sentence. The user slid from co-author down to art director. Smaller job. Even when the finished product looks a lot better than anything a person could type out in the moment.
Directing isn't nothing, to be fair. But it's a different kind of labor than writing, and it leaves people in a different relationship to the result. A director approves a shot. A writer builds the thing being shot. The category quietly swapped one for the other and didn't really say so out loud.
Templated motion: production value becomes a checklist item
Once clip generation was in place, templating followed almost immediately. Templates are what make a generation feature usable at scale instead of a one-off novelty. The current wave of companion video tools ships with aging filters, style transfers, gender-swap effects, and pre-built motion templates that a user can just drop a character into, rather than prompting from scratch. The pitch shifted from "describe what you want" to "pick a look."

And this is where the category's honesty problem really starts, if we're being blunt about it. A template can produce something genuinely striking, or something completely generic wearing a character's face, and both come out of the same button. Our hands-on testing of Yollo's generator suite landed on both ends of that range in a single afternoon: a five-second flower-field clip was, unexpectedly, genuinely convincing. Soft motion, coherent light, an actual sense of a specific place. A templated twerk loop, generated minutes later from the same menu, was not. Same six frames of hip motion, different face bolted on top of it. Production value and intimacy are not the same variable, whatever the roadmap decks say — a gap that shows up across the tools built specifically for companion clips.
None of this is happening by accident, for what it's worth. Grand View Research put the global AI companion market at $28.19 billion in 2024, counting the whole companion category rather than romantic apps alone — a broader slice than the figures in our roundup of the sector's numbers, and money that size pulls product roadmaps toward whatever looks most fundable in a pitch deck. A text box is hard to demo in a meeting. A rendered clip, or a trailer for an in-app "series," is not.
Serialized episodes: the streaming playbook, applied
The clearest sign of where this is heading arrived on July 9, 2026, when Character.AI launched (c.ai) series, its first slate of studio-produced "Microdramas," delivered through a new in-app entertainment tab. Three titles at launch: a romance anime called Last Summer, a paranormal horror story called The Nighttime Game, and a game-world survival drama called Eden Fall. They came out of an in-house studio team with credits spanning Nickelodeon, DreamWorks, Netflix, and Blumhouse, which is not a sentence anyone would have written about this category three years ago. Character.AI paired the launch with (c.ai) fm for serialized audio drama and (c.ai) reads for character-driven fiction, calling the whole push "connected entertainment."

The seam at the end of each episode is what actually makes this different from a normal streaming feed. Finish watching, and over-18 users get dropped straight into a chat with the character they just watched, continuing a story they didn't write a word of.
Character.AI didn't invent that seam, though. It just put the biggest budget behind it. Candy AI has been running the same structure for months: pre-produced Shorts sitting under a "Watch Series" panel on a character's profile, roughly ten episodes of about a minute each, the first one free and the rest unlocked with tokens drawn from the same balance that pays for images and voice. What makes it the sharper example is the crossover. Bring something up from episode three in ordinary chat and the companion picks it up, sometimes with mock offense if you clearly weren't paying attention. The episodes themselves don't branch, and nothing you type changes what happens on screen. But they do get folded back into the relationship as shared memory, which is the whole strategy in miniature: the show supplies the story, and the chat is there to make you carry it around between installments.
Smaller, adult-oriented apps have been reaching for the same shorthand for a while now, too. One companion app on the App Store brands its explicit roleplay scenarios directly as dated "episodes," borrowing the vocabulary of a show rather than a conversation. Character.AI's version is just the first one doing it with an actual production budget and actual Hollywood credits behind it.
Follow that language shift and the strategy gets pretty clear, fast. Chat software doesn't add video features like this. Media companies do. And a media company that happens to also run a chat product mostly wants that chat there to keep you inside the ecosystem between episodes, not to be the main event.
What gets lost when a scene becomes a show
Every one of these apps is quietly betting that better production value equals more intimacy, that a rendered clip lands harder than a paragraph the user helped shape themselves. That's a big enough claim it deserves an actual look instead of just getting taken on faith, and the evidence, honestly, cuts both ways.
Video does unlock something real, and it would be dishonest to pretend otherwise. It takes the burden of description off users who freeze up staring at a blank text box, and a well-executed clip, like that flower-field render, can build a specific, held sense of place that text alone rarely pulls off on the first try. For someone who came to this category for a keepsake or a mood rather than a collaborative scene, that's a real upgrade. Not a gimmick.
What it costs is authorship, though. A watched episode leaves no seam for the user to write themselves into, except the chat that comes after it, and that chat is now competing for attention against an increasingly finished, increasingly scripted piece of media. The templated twerk loop is the honest version of what that costs: technically a video, took no writing at all, produces nothing that resembles a relationship because nothing about it was co-authored. Plenty of users won't care about that difference, and there's nothing wrong with preferring something polished over something you had to write yourself. Fair enough. But it's a different product than the one this category started as, and somebody should say that plainly instead of letting "video" quietly stand in for "better" every single time.
Where this goes next
The category isn't walking video back. The download numbers and the studio hires make that pretty clear. What's more likely is a split: one lane keeps pushing toward passive, high-production episodic content because it's simply easier to make addictive at scale, and another tries to keep some version of the original co-authored contract alive inside the new formats, letting users actually steer a scene instead of just picking a template and watching it play out.
Which lane wins probably comes down to something none of these apps have figured out yet: whether what people were asking for, two years ago, in that plain text box, was better production value, or was just something, someone, paying attention to what they typed. The apps chasing the first answer are easier to fund, easier to demo, easier to pitch. Whether they're the apps people are still using a year from now is a different question. Right now, nobody in this category, including the ones with the biggest production budgets, actually knows the answer to it.
