🎬 JoyAI-Echo — multi-shot video with audio
JoyAI-Echo (JD Joy Future Academy) is a DMD-distilled, LTX-2.3-derived joint audio-video generator: a single 8-step pass produces the picture and its synchronized soundtrack — dialogue, lip-sync, foley and ambience.
Write one prompt per shot. Every shot after the first is conditioned on a paired
cross-modal memory bank (video clips and audio latents carried over from the earlier
shots), which is what keeps a character's face and voice timbre stable across a story.
Reuse the same character description — the authors write ID_A is … — in each shot.
A good shot prompt covers, in order: who (appearance + voice timbre) · action & the spoken line · visual style · camera · background · sound design.
⏱️ A cold run streams ~53 GB of weights onto the GPU before sampling, so the first generation of a session takes noticeably longer than the ones after it. Each extra shot adds a full 8-step pass plus a VAE decode.
Model: jdopensource/JoyAI-Echo · code: jd-opensource/JoyAI-Echo · text encoder: google/gemma-3-12b-it. Released under the LTX-2 Community License — non-commercial use, and generated content must be disclosed as machine-generated.