Tongyi Lab's 14B end-to-end character-animation DiT consumes a driving video directly, avoiding lossy pose or motion extractors. A dual branch, time-aligned RoPE, sparse reference attention, and a viewpoint LoRA provide high-fidelity motion transfer and text-controlled camera changes. The Lite variant distills generation into a causal streaming pipeline.

Model Details

Parameters 14B
License Apache 2.0

Paper

videoanimationdiffusionopen-weightmultimodal