Sidney Primas presents LemonSlice's effort to make real time avatars feel convincingly human during video calls. The system takes a single reference image and adds continuous full body movement, actions, hands, environmental physics, and a visual layer that customers combine with their own language and voice models.
Sidney Primas explains that LemonSlice trains an audio conditioned video model, then makes it causal so generation can look only at prior inputs. The team reduces diffusion from many denoising steps to one for real time output, while addressing error accumulation so an avatar can run continuously for many hours without visible degradation.
Sidney Primas says commercial deployment also depends on efficient hardware use, careful orchestration across GPU and CPU work, and low enough cost for consumer applications. The next step is an emotion engine that selects expressions and actions from audio and text, followed by a longer term end to end model that handles the emotional layer while a separate reasoning model supplies tools and intelligence.
Watch on YouTube



