The serving system that introduced iteration-level scheduling — now known as continuous batching — plus selective batching for transformer inference (Gyeong-In Yu et al., Byung-Gon Chun's group; 36.9x throughput over FasterTransformer at equal latency). The foundational technique beneath vLLM, TGI, TensorRT-LLM, and effectively every modern LLM serving stack. Legacy anchor: Chun now builds serving infrastructure full-time at FriendliAI, with his SNU professorship on leave.

Paper

Venue OSDI 2022
inferenceinfrastructure