JavisDiT Edit — audio-visual object addition & removal

Checkpoint text_both_0721 / epoch001-global_step13000 (EMA) · Wan2.1-T2AV 1.3B, FiLM text conditioning, trained on both edit directions · 81 frames @ 16 fps, 256×256, 50 sampling steps, CFG 5.0 · target's first frame prepended as anchor (stripped before decode). One test-set sample and one train-set sample (overfit check), each run in both directions. All clips include generated audio — unmute to listen.

Addition — “birds”

test set prompt: a video with birds · epMUuqXcgeo_000030

Source (input, birds removed)

Generated (model output)

Target (ground truth)

Removal — “birds”

test set prompt: a video without birds · epMUuqXcgeo_000030

Source (input, birds present)

Generated (model output)

Target (ground truth)

Addition — “trumpet”

train set (overfit check) prompt: a video with trumpet · GfeEN8LONh0_000253

Source (input, trumpet removed)

Generated (model output)

Target (ground truth)

Removal — “trumpet”

train set (overfit check) prompt: a video without trumpet · GfeEN8LONh0_000253

Source (input, trumpet present)

Generated (model output)

Target (ground truth)