Google speeds up video diffusion attention on TPUs
Google developers implemented Sparse VideoGen to route attention heads to structured sparse masks for video diffusion. Combined with Splash Attention kernel optimizations, this achieved up to 1.69x end-to-end inference speedup for 1440p video generation on TPUs.
First seen 30 Sep, 16:24 UTC on Google for Developers1 sourceLast update 1h ago
Google DeepMindResearch · 30 Sep
Research
What the sources say
Linked, never rewritten. Official means the lab itself.
Press0
No independent coverage yet.
How it unfolded
Every source in the order it appeared. Times in UTC.
Keep up with Google
Google on the ScoreIts timelineFollow to mark Google's stories across the site. No account needed.