JavisDiT audio-visual object edit — text_both_base_0728

Checkpoint epoch017-global_step71000, on the first 20 held-out val.jsonl clips. Every clip has generated audio as well as video — unmute to judge the full result.

How to read this
Two things to keep in mind.

Summary — did the edit move the right way?

Mean of (L1 → target − L1 → input) over the 20 clips of each cell. Negative is good: the generation sits closer to the ground-truth target than to the clip it was given.