Skip to content

Type: inference engine

vLLM is an open-source, high-throughput inference and serving engine for large language models, known for its PagedAttention memory management. It is widely used to deploy open-weight models efficiently in production. GROUNDING tracks vLLM’s releases, supported models, and role in self-hosted AI serving.

Recent Updates