Sign inSign up
vLLM OpenAI

dhi.io/vllm-openai

vLLM OpenAI

CIS
FIPS
STIG
linux/amd64
linux/arm64

High-throughput LLM serving with an OpenAI-compatible API

How to use this image

All examples in this guide use the public image. If you've mirrored the repository for your own use (for example, to your Docker Hub namespace), update your commands to reference the mirrored image instead of the public one.

For example:

  • Public image: dhi.io/vllm-openai:<tag>
  • Mirrored image: <your-namespace>/dhi-vllm-openai:<tag>

For the examples, you must first use docker login dhi.io to authenticate to the registry to pull the images.

Start a vLLM OpenAI server

The image requires an NVIDIA GPU, the NVIDIA Container Toolkit, and a model supported by vLLM. The entrypoint is vllm serve, so provide the model identifier followed by any vLLM server arguments⁠.

Create named volumes for the Hugging Face and vLLM caches:

$ docker volume create vllm-huggingface-cache
$ docker volume create vllm-cache

Start the server with a small public model:

$ docker run --rm --name vllm-openai \
  --gpus all \
  --ipc=host \
  -p 8000:8000 \
  -v vllm-huggingface-cache:/home/nonroot/.cache/huggingface \
  -v vllm-cache:/home/nonroot/.cache/vllm \
  dhi.io/vllm-openai:<tag> \
  Qwen/Qwen2.5-0.5B-Instruct \
  --host 0.0.0.0 \
  --port 8000

Model initialization can take several minutes. When the server is ready, list the models exposed by its OpenAI-compatible API:

$ curl http://localhost:8000/v1/models

For a gated model, pass a Hugging Face access token:

$ docker run --rm --name vllm-openai \
  --gpus all \
  --ipc=host \
  -p 8000:8000 \
  -e HF_TOKEN \
  -v vllm-huggingface-cache:/home/nonroot/.cache/huggingface \
  -v vllm-cache:/home/nonroot/.cache/vllm \
  dhi.io/vllm-openai:<tag> \
  <model-id> \
  --host 0.0.0.0 \
  --port 8000
Cache environment variables

XDG_CACHE_HOME is /home/nonroot/.cache, and that directory is owned by nonroot. A library can create a new cache directory there without an image change. The image also pins these locations:

  • HF_HOME=/home/nonroot/.cache/huggingface
  • VLLM_CACHE_ROOT=/home/nonroot/.cache/vllm
  • TRITON_CACHE_DIR=/home/nonroot/.cache/triton
  • NUMBA_CACHE_DIR=/home/nonroot/.cache/numba
  • /home/nonroot/.cache/flashinfer for FlashInfer kernels (XDG_CACHE_HOME)

Mount persistent volumes at these paths when repeated model downloads or kernel compilation would otherwise slow startup.

Common vLLM OpenAI use cases

Send a text-completion request

After starting the server, send a request to the OpenAI-compatible completions endpoint. The model value must match the model identifier passed when starting the container.

$ curl http://localhost:8000/v1/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen2.5-0.5B-Instruct",
    "prompt": "Docker Hardened Images help",
    "max_tokens": 32,
    "temperature": 0
  }'
Require a key for OpenAI-compatible endpoints

Pass --api-key when starting the server:

$ docker run --rm --name vllm-openai \
  --gpus all \
  --ipc=host \
  -p 8000:8000 \
  -v vllm-huggingface-cache:/home/nonroot/.cache/huggingface \
  -v vllm-cache:/home/nonroot/.cache/vllm \
  dhi.io/vllm-openai:<tag> \
  Qwen/Qwen2.5-0.5B-Instruct \
  --host 0.0.0.0 \
  --port 8000 \
  --api-key change-me

Include the key in requests:

$ curl http://localhost:8000/v1/models \
  -H "Authorization: Bearer change-me"

This option protects the /v1, /v2, and /inference path prefixes, but it doesn't protect every server endpoint. Use a trusted reverse proxy or another network access control when exposing vLLM outside a trusted network.

Non-hardened images vs. Docker Hardened Images

The standard upstream vllm/vllm-openai target runs as root, while this runtime image runs as the nonroot user (uid 65532). The -dev and -fips-dev variants run as root. /home/nonroot/.cache is owned by nonroot, so new cache directories can be created there. Mount model and cache volumes at the configured cache paths instead of root-owned paths.

The package-backed environment is stored at /usr/lib/vllm-openai (also reachable as /opt/venv). The vllm CLI on PATH (/usr/bin/vllm) is correct for serving. For custom Python scripts, use /usr/lib/vllm-openai/bin/python or /opt/venv/bin/python: bare /usr/bin/python3 sees the system PyTorch package but not the vLLM install in the venv. The upstream /vllm-workspace examples and benchmark source tree aren't included.

This image is package-backed and uses CUDA-enabled PyTorch for NVIDIA GPUs. It doesn't provide a CPU-only serving variant.

Image variants

Docker Hardened Images come in different variants depending on their intended use. Image variants are identified by their tag.

  • Runtime variants are designed to run your application in production. These images are intended to be used either directly or as the FROM image in the final stage of a multi-stage build. These images typically:

    • Run as a nonroot user
    • Do not include a shell or a package manager
    • Contain only the minimal set of libraries needed to run the app
  • Build-time variants typically include dev in the tag name and are intended for use in the first stage of a multi-stage Dockerfile. These images typically:

    • Run as the root user
    • Include a shell and package manager
    • Are used to build or compile applications
  • FIPS variants include fips in the variant name and tag. They come in both runtime and build-time variants. These variants use cryptographic modules that have been validated under FIPS 140, a U.S. government standard for secure cryptographic operations. For example, usage of MD5 fails in FIPS variants.

To view the image variants and get more information about them, select the Tags tab for this repository, and then select a tag.

Migrate to a Docker Hardened Image

To migrate your application to a Docker Hardened Image, you must update your Dockerfile. At minimum, you must update the base image in your existing Dockerfile to a Docker Hardened Image. This and a few other common changes are listed in the following table of migration notes.

ItemMigration note
Base imageReplace your base images in your Dockerfile with a Docker Hardened Image.
Package managementNon-dev images, intended for runtime, don't contain package managers. Use package managers only in images with a dev tag.
Non-root userBy default, non-dev images, intended for runtime, run as the nonroot user. Ensure that necessary files and directories are accessible to the nonroot user.
Multi-stage buildUtilize images with a dev tag for build stages and non-dev images for runtime. For binary executables, use a static image for runtime.
TLS certificatesDocker Hardened Images contain standard TLS certificates by default. There is no need to install TLS certificates.
PortsNon-dev hardened images run as the nonroot user by default. As a result, applications in these images can't bind to privileged ports (below 1024) when running in Kubernetes or in Docker Engine versions older than 20.10. To avoid issues, configure your application to listen on port 1025 or higher inside the container.
Entry pointDocker Hardened Images may have different entry points than images such as Docker Official Images. Inspect entry points for Docker Hardened Images and update your Dockerfile if necessary.
No shellBy default, non-dev images, intended for runtime, don't contain a shell. Use dev images in build stages to run shell commands and then copy artifacts to the runtime stage.

The following steps outline the general migration process.

  1. Find hardened images for your app.

    A hardened image may have several variants. Inspect the image tags and find the image variant that meets your needs.

  2. Update the base image in your Dockerfile.

    Update the base image in your application's Dockerfile to the hardened image you found in the previous step. For framework images, this is typically going to be an image tagged as dev because it has the tools needed to install packages and dependencies.

  3. For multi-stage Dockerfiles, update the runtime image in your Dockerfile.

    To ensure that your final image is as minimal as possible, you should use a multi-stage build. All stages in your Dockerfile should use a hardened image. While intermediary stages will typically use images tagged as dev, your final runtime stage should use a non-dev image variant.

  4. Install additional packages

    Docker Hardened Images contain minimal packages in order to reduce the potential attack surface. You may need to install additional packages in your Dockerfile. Inspect the image variants to identify which packages are already installed.

    Only images tagged as dev typically have package managers. You should use a multi-stage Dockerfile to install the packages. Install the packages in the build stage that uses a dev image. Then, if needed, copy any necessary artifacts to the runtime stage that uses a non-dev image.

    For Alpine-based images, you can use apk to install packages. For Debian-based images, you can use apt-get to install packages.

Troubleshooting migration

The following are common issues that you may encounter during migration.

General debugging

The hardened images intended for runtime don't contain a shell nor any tools for debugging. The recommended method for debugging applications built with Docker Hardened Images is to use Docker Debug⁠ to attach to these containers. Docker Debug provides a shell, common debugging tools, and lets you install other tools in an ephemeral, writable layer that only exists during the debugging session.

Permissions

By default image variants intended for runtime, run as the nonroot user. Ensure that necessary files and directories are accessible to the nonroot user. You may need to copy files to different directories or change permissions so your application running as the nonroot user can access them.

Privileged ports

Non-dev hardened images run as the nonroot user by default. As a result, applications in these images can't bind to privileged ports (below 1024) when running in Kubernetes or in Docker Engine versions older than 20.10. To avoid issues, configure your application to listen on port 1025 or higher inside the container, even if you map it to a lower port on the host. For example, docker run -p 80:8080 my-image will work because the port inside the container is 8080, and docker run -p 80:81 my-image won't work because the port inside the container is 81.

No shell

By default, image variants intended for runtime don't contain a shell. Use dev images in build stages to run shell commands and then copy any necessary artifacts into the runtime stage. In addition, use Docker Debug to debug containers with no shell.

Entry point

Docker Hardened Images may have different entry points than images such as Docker Official Images. Use docker inspect to inspect entry points for Docker Hardened Images and update your Dockerfile if necessary.