FastVideo contributions
FastVideo is an open-source framework for accelerated video generation, covering inference and post-training (4.5k stars). I've worked on it at UCSD's Hao AI Lab since January 2026, mostly on inference performance: profile where the time actually goes, fix it, and gate every change on output quality. 16 of my 32 PRs are merged.
Inference performance
- Quantize decoded frames on the GPU before the device→host copy #1362 merged
−15.6% end-to-end on Cosmos 2.5 / A100. A “slow” post-decode stage was really absorbing a deferred fp32 transfer; the fix makes it 4× smaller. Independently cross-validated on Wan / H100.
- torch.compile, benchmarked and documented #1366 merged
−23.7% end-to-end on Wan2.1-1.3B / A100 (259.7 s → 198.1 s) at SSIM ≈ 1.0 — a speedup that already existed but was off by default and undocumented.
- Widen the sparse-attention (VSA) Triton autotune range #1706 merged
18–22% faster kernels. The best pipeline depth sat outside the range the autotuner was allowed to search.
- bf16 VAE decode for every Wan variant #1472 merged
~1.2–1.3× faster decode at MS-SSIM 0.9999. Decode precision is split from encode so causal and I2V models keep bit-identical trajectories.
- Opt-in step caching for the Wan DiT #1426 open
−18% wall time at SSIM 0.957, or −32% at 0.943, via cache-dit with TaylorSeer.
Compiler & kernels
- FlashAttention as a torch.library custom op #1373 merged
Makes attention a traceable node under torch.compile. Output-identical: SSIM 1.000 across all frames, compiled and eager.
- Real backward for FlashAttention-2's default and masked/varlen paths #1388 merged
Extends the custom op to training, so the backward pass compiles too. SSIM 1.000 on 49/49 frames.
- Remove a per-layer graph break in the layer-offload hook #1365 merged
Graph breaks 25 → 13 in a default-on path every model runs. SSIM 1.000 across 81/81 frames.
New hardware: NVIDIA DGX Spark (GB10)
- DGX Spark performance and tuning guide #1631 merged
Measured guidance with reproduction examples: few-step vs full-step models (~18×), bf16 VAE decode, and opt-in FP4 attention.
- Wire Cosmos 2.5 2B to its sampling preset #1468 merged
A missing registry entry silently disabled guidance, producing blurry video. Found on the Spark; sharpness ~92 → ~431, with a test for the gap.
- 4-bit (FP4) linear layers for LTX-2.3 on the Spark #1594 open
−23.6% denoise time at 1080p, 40.5 seconds saved per video (231 s → 192 s end to end).
- Build and enable FP4 attention on sm_121a #1598 open
−6% denoise, visually equivalent to bf16, on hardware where the kernel had been considered unsupported.
- Cast fp32 inputs to bf16 at the FP4 boundary #1488 merged
Replaces an assertion with a cast, unblocking the FP4 path on real pipelines.
Models & pipelines
- Cosmos 2.5 training pipeline (LoRA + full fine-tune) #1227 merged
A 2,000-step LoRA run completed cleanly on a single GPU (final loss 0.067).
-
Chains text-to-world and video continuation stages into one seamless clip — boundary error 2.7 vs 39 for an unrelated control.
- FLUX.1-dev port fixes #1321 merged
Native RoPE, parity tests, and an SSIM reference.
- Z-Image port hardening #1339 merged
Strict weight loading and bf16 encoder parity, with a per-layer diagnostic separating drift from bugs.
- Fall back when a large pinned-memory allocation fails #1759 merged
Degrades gracefully instead of crashing on long, high-resolution outputs.