Inference optimization on FastVideo — EasyCache, a model-agnostic residual-skip cache that takes 23% off Wan2.1 wall time at SSIM 0.946; a serving-path split that cut a Cosmos 2.5 profile 26.2%; a torch.compile/FlashAttention traceability path for training; and a port to NVIDIA’s DGX Spark. 16 merged PRs so far — see all my PRs →
Work
- Jan2026 - PresentHao AI Lab, UC San DiegoContributor
- Jun2025 - Sep2025AmazonSoftware Development Engineer Intern
Built a distributed log indexing service (42M+ entries/hour) that cut incident triage from 15 minutes to under 45 seconds, and automated on-call SOPs with Step Functions.
- Apr2023 - Sep2024Aark GlobalSoftware Developer, AI/ML
Async document-ingestion pipeline at 18K+ pages/day, a read/write routing layer that held sub-100ms P95 through a live datastore migration, and an OCR-to-Elasticsearch search pipeline.
- Jun2022 - Mar2023ConcentrixData Engineer
Replaced sequential scrapers with Airflow-orchestrated Kafka streaming — 60% more throughput, data-freshness lag from 3 days to 6 hours.