vLLM

High-throughput, memory-efficient inference and serving for large language models.

Visit website

Open-source (Apache 2.0) engine for serving LLMs — PagedAttention, continuous batching, OpenAI-compatible API on your own GPUs. Free.

vLLM homepage — high-throughput LLM serving

vLLM homepage — high-throughput LLM serving

Overview

What is vLLM?

High-throughput, memory-efficient inference and serving for large language models.

vLLM is an open-source library for high-throughput, memory-efficient inference and serving of large language models, originally developed in UC Berkeley's Sky Computing Lab and now maintained by a 2,000+ contributor community.

It runs open-weight models on your own GPU infrastructure through an OpenAI-compatible API server, a Python library, or offline batch inference. Key techniques include PagedAttention for KV-cache memory efficiency, continuous batching, chunked prefill, and prefix caching — plus quantization (FP8, GPTQ/AWQ, GGUF) and distributed inference. Apache 2.0 licensed; no hosted SaaS or paid tier.

Platforms and languages

vLLM Availability

Platforms

Linux

Languages

English

Capabilities

vLLM Key Features

PagedAttention

Memory-efficient attention KV-cache management that reduces fragmentation during generation.

Continuous batching

Batch incoming requests continuously, with chunked prefill and prefix caching for high throughput.

OpenAI-compatible server

Drop-in API server plus Anthropic Messages API and gRPC support.

Quantization support

FP8, MXFP8/MXFP4, NVFP4, INT8/INT4, GPTQ/AWQ, GGUF, and more.

Broad hardware support

NVIDIA and AMD GPUs, CPUs, and plugins for TPUs, Intel Gaudi, Huawei Ascend, and Apple Silicon.

Best for

Who uses vLLM?

ML engineers and teams

Serve open-weight LLMs at high throughput on own GPU infrastructure

Plans and access

vLLM Pricing

Free

No free trial listed

Common questions

vLLM FAQs

What is vLLM?

A fast, open-source library for LLM inference and serving, originating at UC Berkeley's Sky Computing Lab.

Is vLLM free?

Yes — Apache 2.0 licensed, fully free and open source, with no paid tiers.

What do I need to run vLLM?

Linux, Python 3.9–3.12, and a GPU with compute capability 7.0+; CPU-only builds are also available.

Which models does vLLM support?

200+ model architectures on Hugging Face, including Llama, Qwen, Gemma, MoE models like Mixtral and DeepSeek-V3, and multimodal models.

Reviews

See what the community thinks and share your experience.

—

Based on 0 ratings

Rating distribution

0
0
0
0
0

Leave a review

Sign in to rate

Community reviews

No reviews

No written reviews yet.

Explore by category

AI Development Infrastructure

View all AI Development Infrastructure websites
Ollama preview
Ollama

Free, open-source runtime for running open LLMs locally on macOS, Windows, and Linux — CLI, OpenAI-compatible API, GGUF import. Local runs unlimited.

AI Development Infrastructure

Explore by category

AI Apis Sdks

View all AI Apis Sdks websites
热核算力 preview
热核算力

面向开发者的付费 GPT API 聚合服务,支持 GPT GO、PLUS、PRO 分组与按量使用

AI Apis Sdks
Space Bunny preview
Space Bunny

A long-context multimodal AI playground and developer API for reasoning across text, images, video, code, and large documents.

AI Apis Sdks
Jev Model preview
Jev Model

Jev Model turns real-world context into fast, structured decisions with probabilities for software teams.

AI Apis Sdks
cl0q preview
cl0q

Open search engine for domain research & OSINT — 38.5M domains scanned

AI Apis Sdks

Explore by category

Llms Foundation Models

View all Llms Foundation Models websites
Ollama preview
Ollama

Free, open-source runtime for running open LLMs locally on macOS, Windows, and Linux — CLI, OpenAI-compatible API, GGUF import. Local runs unlimited.

Llms Foundation Models
Hugging Face preview
Hugging Face

The Hub for open machine learning: 2M+ models, 1.5M datasets and Spaces apps, plus inference APIs. Free for public use; paid plans from $9/month.

Llms Foundation Models
LM Studio preview
LM Studio

Desktop app to run open LLMs locally on macOS, Windows, and Linux — chat UI, OpenAI-compatible local server, offline transcription. Free plan.

Llms Foundation Models
Clef preview
Clef

Cloudflare's open-source 27B multimodal decision model: state plus typed questions in, probabilities for every option out.

Llms Foundation Models