Self-hosted vLLM inference stack with an OpenAI-compatible API, Docker Compose templates, Caddy reverse proxy, and NVIDIA GPU thermal guard.