Current diffusion fashions allow high-quality video technology, however endure from gradual runtimes. The big transformer-based backbones utilized in these fashions are bottlenecked by spatiotemporal consideration. On this paper, we establish {that a} vital fraction of token-to-token connections persistently yield negligible scores throughout varied inputs, and their patterns usually repeat throughout queries. Thus, the eye computation in these circumstances may be skipped with little to no impact on the consequence. This remark continues to carry for connections amongst native token blocks. Motivated by this, we introduce CalibAtt, a training-free methodology that accelerates video technology by way of calibrated sparse consideration. CalibAtt performs an offline calibration move that identifies block-level sparsity and repetition patterns which are steady throughout inputs, and compiles these patterns into optimized consideration operations for every layer, head, and diffusion timestep. At inference time, we compute the chosen input-dependent connections densely, and skip the unselected ones in a hardware-efficient method. Intensive experiments on Wan 2.1 14B, Mochi 1, and few-step distilled fashions at varied resolutions present that CalibAtt achieves as much as 1.58× end-to-end speedup, outperforming present training-free strategies whereas sustaining video technology high quality and text-video alignment.
† Tel Aviv College** Work performed whereas at Apple

