Inference optimization on FastVideo — a caching module that takes 22–47% off Wan2.1, a torch.compile/FlashAttention traceability path for training, and a port to NVIDIA’s DGX Spark. Merged PRs →
Work
- Jan2026 - PresentHao AI Lab, UC San DiegoStudent Researcher
- Jun2025 - Sep2025AmazonSoftware Development Engineer Intern
Built a distributed log indexing service (42M+ entries/hour) that cut incident triage from 15 minutes to under 45 seconds, and automated on-call SOPs with Step Functions.
- Apr2023 - Sep2024Aark GlobalSoftware Developer, AI/ML
Async document-ingestion pipeline at 18K+ pages/day, a read/write routing layer that held sub-100ms P95 through a live datastore migration, and an OCR-to-Elasticsearch search pipeline.
- Jun2022 - Mar2023ConcentrixData Engineer
Replaced sequential scrapers with Airflow-orchestrated Kafka streaming — 60% more throughput, data-freshness lag from 3 days to 6 hours.