Blog
2026
- One scene, three futures: the GPU kept up, but the choices did not alwaysI batched three five-second video futures from one frame on a B200, built a choice-and-continue interface, and found that fast generation and controllable outcomes are very different things.
- 231s to 192s: 4-bit video diffusion on a $3,000 desktopQuantizing a video model's linear layers to 4-bit cut a 1080p generation by 39 seconds — and did nothing at all on the next two models I tried it on. The gap is the interesting part.
- From 15 minutes to 45 seconds: rebuilding an Amazon on-call tool in GoHow I rebuilt a production log triage system — binary search over a sorted index, a channel-based ingestion pool, a write-ahead log for crash recovery, and what I'd do differently at scale.
- Building Flare: LLM-Powered Incident Detection on Real Log DataHow I built an end-to-end log anomaly detection pipeline that combines classical ML with LLM summarization — and what I learned about evaluating LLM output in production-adjacent systems.