Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning
Shan Yu,
Yifan Qiao,
Mingyuan Ma,
Yangmin Li,
Shuo Yang,
Xinyuan Tong,
Yang Wang,
Zhiqiang Xie,
Yuwei An,
Shiyi Cao,
Ke Bao,
Deepak Vij,
Xiaoning Ding,
Yichen Wang,
Qingda Lu,
Zhong Wang,
Gao Gao,
Harry Xu#,
Junyi Shu#,
Jiarong Xing#,
and Ying Sheng#
In 20th USENIX Symposium on Operating Systems Design and Implementation (OSDI 26)
2026
Deployed in production by Tencent, Alibaba Cloud, Red Hat, and others at a scale of 10,000+ GPUs.
Artifacts Available
Inference providers must maintain availability for many LLMs, including low-volume but essential models, making resource efficiency increasingly important as token prices fall. Analysis of production traces reveals a dynamic bursty-group pattern in which sets of models become active together and shift over time; existing space- and time-sharing approaches lack principled mechanisms to adapt to this variability, forcing trade-offs between SLO adherence and efficiency. We observe that elastic memory allocation can unify spatial and temporal sharing. Based on this insight, we have developed Prism, a memory-centric LLM co-serving framework that applies memory ballooning to reclaim memory across models and support both forms of sharing under a single scheme. Prism’s balloon driver, referred to as kvcached, has been open-sourced at https://github.com/ovg-project/kvcached, and deployed in production environments across 10K+ GPUs.