vLLM
“高吞吐量的大语言模型推理和服务框架,适合部署自有模型 API。”
手机扫码访问
网站快照
vLLM is a high-throughput and memory-efficient inference and serving engine for Large Language Models (LLMs). Deploy AI models faster with state-of-the-art performance. Easy, fast, and cost-efficient LLM serving for everyone.
点评与评分已隐藏
该站点的网友评分及点评意见需要登录后才能查看与发表。