Checkpoint text_both_0721 / epoch001-global_step13000 (EMA) · Wan2.1-T2AV 1.3B, FiLM text conditioning,
trained on both edit directions · 81 frames @ 16 fps, 256×256, 50 sampling steps, CFG 5.0 ·
target's first frame prepended as anchor (stripped before decode). One test-set sample and one train-set sample (overfit check), each run in both directions. All clips include generated audio — unmute to listen.
Addition — “birds”
test setprompt: a video with birds · epMUuqXcgeo_000030
Source (input, birds removed)
Generated (model output)
Target (ground truth)
Removal — “birds”
test setprompt: a video without birds · epMUuqXcgeo_000030
Source (input, birds present)
Generated (model output)
Target (ground truth)
Addition — “trumpet”
train set (overfit check)prompt: a video with trumpet · GfeEN8LONh0_000253
Source (input, trumpet removed)
Generated (model output)
Target (ground truth)
Removal — “trumpet”
train set (overfit check)prompt: a video without trumpet · GfeEN8LONh0_000253