Google DeepMindResearchMateriality 2

Google speeds up video diffusion attention on TPUs

Google developers implemented Sparse VideoGen to route attention heads to structured sparse masks for video diffusion. Combined with Splash Attention kernel optimizations, this achieved up to 1.69x end-to-end inference speedup for 1440p video generation on TPUs.

First seen 30 Sep, 16:24 UTC on Google for Developers1 sourceLast update 1h ago

What the sources say

Linked, never rewritten. Official means the lab itself.

How it unfolded

Every source in the order it appeared. Times in UTC.

Keep up with Google