FastVideo contributions

FastVideo is an open-source framework for accelerated video generation, covering inference and post-training (4.5k stars). I've worked on it at UCSD's Hao AI Lab since January 2026, mostly on inference performance: profile where the time actually goes, fix it, and gate every change on output quality. 16 of my 32 PRs are merged.

  • Quantize decoded frames on the GPU before the device→host copy #1362 merged

    −15.6% end-to-end on Cosmos 2.5 / A100. A “slow” post-decode stage was really absorbing a deferred fp32 transfer; the fix makes it 4× smaller. Independently cross-validated on Wan / H100.

  • torch.compile, benchmarked and documented #1366 merged

    −23.7% end-to-end on Wan2.1-1.3B / A100 (259.7 s → 198.1 s) at SSIM ≈ 1.0 — a speedup that already existed but was off by default and undocumented.

  • Widen the sparse-attention (VSA) Triton autotune range #1706 merged

    18–22% faster kernels. The best pipeline depth sat outside the range the autotuner was allowed to search.

  • bf16 VAE decode for every Wan variant #1472 merged

    ~1.2–1.3× faster decode at MS-SSIM 0.9999. Decode precision is split from encode so causal and I2V models keep bit-identical trajectories.

  • Opt-in step caching for the Wan DiT #1426 open

    −18% wall time at SSIM 0.957, or −32% at 0.943, via cache-dit with TaylorSeer.

  • FlashAttention as a torch.library custom op #1373 merged

    Makes attention a traceable node under torch.compile. Output-identical: SSIM 1.000 across all frames, compiled and eager.

  • Real backward for FlashAttention-2's default and masked/varlen paths #1388 merged

    Extends the custom op to training, so the backward pass compiles too. SSIM 1.000 on 49/49 frames.

  • Remove a per-layer graph break in the layer-offload hook #1365 merged

    Graph breaks 25 → 13 in a default-on path every model runs. SSIM 1.000 across 81/81 frames.

  • DGX Spark performance and tuning guide #1631 merged

    Measured guidance with reproduction examples: few-step vs full-step models (~18×), bf16 VAE decode, and opt-in FP4 attention.

  • Wire Cosmos 2.5 2B to its sampling preset #1468 merged

    A missing registry entry silently disabled guidance, producing blurry video. Found on the Spark; sharpness ~92 → ~431, with a test for the gap.

  • 4-bit (FP4) linear layers for LTX-2.3 on the Spark #1594 open

    −23.6% denoise time at 1080p, 40.5 seconds saved per video (231 s → 192 s end to end).

  • Build and enable FP4 attention on sm_121a #1598 open

    −6% denoise, visually equivalent to bf16, on hardware where the kernel had been considered unsupported.

  • Cast fp32 inputs to bf16 at the FP4 boundary #1488 merged

    Replaces an assertion with a cast, unblocking the FP4 path on real pipelines.

  • Cosmos 2.5 training pipeline (LoRA + full fine-tune) #1227 merged

    A 2,000-step LoRA run completed cleanly on a single GPU (final loss 0.067).

  • Cosmos Predict2.5 distilled video continuation #1767 #1768 open

    Chains text-to-world and video continuation stages into one seamless clip — boundary error 2.7 vs 39 for an unrelated control.

  • FLUX.1-dev port fixes #1321 merged

    Native RoPE, parity tests, and an SSIM reference.

  • Z-Image port hardening #1339 merged

    Strict weight loading and bf16 encoder parity, with a per-layer diagnostic separating drift from bugs.

  • Fall back when a large pinned-memory allocation fails #1759 merged

    Degrades gracefully instead of crashing on long, high-resolution outputs.

All 32 PRs on GitHub →