jnachi
Learning Hub
Python Development9 min readAdvanced

Deploying & Scaling Python AI Services with Docker, GPU Workers & vLLM

Containerize and scale Python AI workloads: Multi-stage Docker builds, high-throughput self-hosted inference with vLLM PagedAttention, and Kubernetes GPU scaling.

Works with:DockerKubernetesvLLMNVIDIA TritonOllama / AWQ

Key Takeaways

  • Multi-stage Docker builds separate build dependencies from runtime containers, creating lightweight, hardened production images
  • vLLM uses PagedAttention memory management to achieve up to 24x higher serving throughput for open-source models (Llama 3, Mistral, Qwen) than naive HuggingFace pipelines
  • Quantization (AWQ, GPTQ, GGUF) compresses 16-bit model weights down to 4-bit, enabling 70B models to run on cost-effective GPUs without quality degradation
  • Kubernetes Horizontal Pod Autoscalers (HPA) scale AI services based on custom metrics (concurrent requests, queue depth, GPU memory utilization)

The Diagnostic Context

Running an AI prototype in a Jupyter notebook is easy; serving low-latency inference to millions of concurrent users under strict SLA constraints requires industrial-grade containerization, GPU memory optimization, and resilient deployment topologies.

The Core Technique

High-Performance Open-Source Model Serving: vLLM

DIAGRAM / WORKFLOW
graph TD
    subgraph ClientLayer["Client Requests"]
        C1["User 1 (Fast Stream)"]
        C2["User 2 (Long Prompt)"]
        C3["User 3 (Short Query)"]
    end

    subgraph vLLM_Engine["vLLM High-Throughput Inference Engine"]
        Scheduler["Continuous Batching Engine (No idle GPU cycles)"]
        PagedAttention["PagedAttention KV-Cache Memory Management<br/>(Zero fragmented GPU VRAM)"]
        Scheduler --- PagedAttention
    end

    subgraph GPU_Hardware["NVIDIA GPU (A100 / H100 / L40S)"]
        VRAM["Model Weights (Llama 3 70B AWQ Quantized: 38GB VRAM)"]
        TensorCores["Tensor Core Matrix Computation"]
    end

    C1 --> Scheduler
    C2 --> Scheduler
    C3 --> Scheduler

    PagedAttention --> VRAM
    PagedAttention --> TensorCores

Production Multi-Stage Dockerfile for Python AI Microservices

DOCKERFILE
# Stage 1: Build Dependencies
FROM python:3.12-slim AS builder

WORKDIR /build
ENV UV_SYSTEM_PYTHON=1

# Install uv for ultra-fast dependency installation:
COPY --from=ghcr.io/astral-sh/uv:latest /uv /bin/uv
COPY pyproject.toml requirements.txt ./

RUN uv pip install --no-cache -r requirements.txt --target /build/packages

# Stage 2: Hardened Runtime Container
FROM python:3.12-slim AS runtime

WORKDIR /app
ENV PYTHONUNBUFFERED=1     PYTHONDONTWRITEBYTECODE=1     PYTHONPATH=/app/packages

# Create non-root user for enterprise container security:
RUN useradd -u 10001 -m aiuser

COPY --from=builder /build/packages /app/packages
COPY ./src /app/src

USER aiuser
EXPOSE 8000

CMD ["python", "-m", "uvicorn", "src.main:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "4"]
5-Minute Activation Challenge

Try This Right Now

Write a `docker run` command using the `--gpus all` flag to run a containerized vLLM instance serving `meta-llama/Meta-Llama-3-8B-Instruct` with an OpenAI-compatible REST endpoint on port 8000.

Tip: Knowledge only becomes capability once you run the prompt yourself.

Comprehension Check

Test Your Instincts (3 Questions)

1

What breakthrough memory management algorithm does vLLM introduce to eliminate KV-cache fragmentation and boost LLM serving throughput by up to 24x?

2

Why are multi-stage Docker builds recommended for enterprise Python AI container deployments?

3

What does 4-bit weight quantization (e.g., AWQ or GPTQ) accomplish for large language models?