🎬 JoyAI-Echo — multi-shot video with audio

JoyAI-Echo (JD Joy Future Academy) is a DMD-distilled, LTX-2.3-derived joint audio-video generator: a single 8-step pass produces the picture and its synchronized soundtrack — dialogue, lip-sync, foley and ambience.

Write one prompt per shot. Every shot after the first is conditioned on a paired cross-modal memory bank (video clips and audio latents carried over from the earlier shots), which is what keeps a character's face and voice timbre stable across a story. Reuse the same character description — the authors write ID_A is … — in each shot.

A good shot prompt covers, in order: who (appearance + voice timbre) · action & the spoken line · visual style · camera · background · sound design.

Resolution

1280 x 736 is the model's training resolution; 832 x 480 is ~2x faster.

41 201

⏱️ A cold run streams ~53 GB of weights onto the GPU before sampling, so the first generation of a session takes noticeably longer than the ones after it. Each extra shot adds a full 8-step pass plus a VAE decode.

Example stories — the authors' own prompts (prompts/test_*.json)

Model: jdopensource/JoyAI-Echo · code: jd-opensource/JoyAI-Echo · text encoder: google/gemma-3-12b-it. Released under the LTX-2 Community License — non-commercial use, and generated content must be disclosed as machine-generated.