vLLM

高吞吐量的大语言模型推理和服务框架,适合部署自有模型 API。

二维码

手机扫码访问

访问
网站快照
vLLM 网站截图

vLLM is a high-throughput and memory-efficient inference and serving engine for Large Language Models (LLMs). Deploy AI models faster with state-of-the-art performance. Easy, fast, and cost-efficient LLM serving for everyone.