Two avatar announcements arrived on the same day. Both promise a character that can respond in live video. The tempting shortcut is to rank them by the smallest millisecond number. Their own measurements do not allow that.
Meta introduced Muse Realtime Avatar on September 23 as a system that turns the speech-token stream from Muse Realtime Voice into synchronized video. A reference image can be a portrait, full-body illustration, animal or object. Meta says the system uses recent generated frames as context so a character’s appearance and mannerisms continue across conversational turns.
LemonSlice released Character World Model-1 on the same date. It says users can start from one image and animate the whole character, including gestures and interactions with the surrounding scene. Its “emotion engine” can guide actions and expressions during a conversation. A free web demo is available; API access is listed for Ultra and Enterprise customers.
The numbers do not measure the same interval
Meta reports portrait video at 448×768 pixels and 25 frames per second, with about 870 milliseconds measured from the end of a user’s turn to receipt of the first byte of synchronized voice and video. It also reports 12 concurrent video sessions on one NVIDIA GB200. Those are company-reported serving results, not independent measurements of a public product.
LemonSlice says its model can generate response video frames within 471 milliseconds when interrupted. That starts from a different event and ends at a different point in the pipeline. It cannot be subtracted from Meta’s 870 ms to conclude that one service is faster for an actual user. A fair test would use the same device, network, avatar identity, turn-taking pattern, output resolution and endpoint definition.
Meta’s own user study compared Muse Realtime Avatar with Runway Characters and HeyGen LiveAvatar—not LemonSlice. It reported a preference for Muse on the measures it tested, with one Runway mannerism comparison not statistically distinguishable from parity. That study says nothing about how Muse would perform against CWM-1.
Access and safeguards are the near-term differences
Meta cautions that the examples in its research post show model capability and do not all represent avatars available in the Muse app. Muse is for users aged 18 or older, and Meta says generated video carries an invisible Video Seal watermark. LemonSlice currently offers a public demo and a tier-limited API, making hands-on evaluation more straightforward for some developers.
Teams assessing either system should test not only visual appeal but also interruptions, identity consistency over a long call, consent for uploaded likenesses, safety controls and the total cost of serving simultaneous users. The next credible comparison will come from a shared test protocol, not from two launch-day marketing figures.
The broader generative-video landscape includes tools for making a finished clip, but a live avatar has to react during a conversation. That is a different product test. It also differs from choosing a model through an AI video API with multiple providers, where a developer can compare completed outputs and per-generation costs after the fact.
Sources: Meta AI Research announcement; LemonSlice CWM-1 announcement.