Bernini-Diffusers-v2 — reference-to-video
Drop in a few reference images (a subject, an outfit, a prop, a scene…), then describe the
video you want while pointing at them as image0, image1, … Bernini's Qwen2.5-VL planner reads
the references plus your instruction and plans a target visual embedding, which the Wan2.2-A14B
MoE renderer turns into a video.
Longer clips and more steps look better but cost more GPU time. The defaults (33 frames ≈ 2 s at 16 fps, 16 steps) take about 4 minutes; the authors' reference setting is 81 frames / 40 steps, which does not fit in a single ZeroGPU slot.
Official Bernini r2v examples