Serve with Docker
Distilled from docs/docker.md. Commands run from the Laya repository root. This site does not host the container. The workbench cURL targets 127.0.0.1:8000/v1/systemone after you start laya-serve yourself.
CPU quickstart
Allow about 8 GB of RAM and 10 GB of disk, with Docker Compose v2 or newer. The first run downloads a public checkpoint and can take several minutes. Weights stay in a named volume.
docker compose run --build --rm laya
# later, weights already in the named volume:
docker compose run --rm layaGPU: a compatible NVIDIA driver and the NVIDIA Container Toolkit. The GPU image uses PyTorch CUDA 12.8 wheels. LAYA_DEVICE defaults to cuda in that override, and LAYA_GPU_ID selects the host GPU. The doc pins LAYA_TORCH_VERSION at 2.14.0 for the Compose build.
docker compose -f compose.yaml -f compose.cuda.yaml run --build --rm layaThe image sets TORCH_DISABLE_NATIVE_JIT=1. The doc says PyTorch 2.14 otherwise compiles Triton kernels on first inference and needs a C compiler the slim image does not have, so the container looks healthy and then fails every request (#365).
HTTP
compose.http.yaml adds laya-serve. The port is published on 127.0.0.1 only. There is no authentication until LAYA_API_KEY is set. Set a key before LAYA_BIND_ADDRESS=0.0.0.0, and put TLS in front for remote clients. /health stays unauthenticated.
docker compose -f compose.yaml -f compose.http.yaml up --build laya-serve
curl -s localhost:8000/health
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' \
--data @examples/docker/request.jsonAdd -f compose.cuda.yaml for a GPU server. The CUDA override has to name the serve service, because overrides for the one-shot service do not reach it.
LAYA_PRELOADdefaults to0in this Compose service so the first boot does not download all three checkpoints. The package default is1. Set1for a long-running process.LAYA_MAX_LOADEDdefaults to 2.LAYA_AUTO_TASKdefaults to 0, sotyped-decisionsis not reached automatically.LAYA_MAX_CONCURRENTdefaults to 16. Further requests get 503.LAYA_MAX_TOKEN_BUDGETdefaults to 8192.- The image runs as UID 10001. A LoRA adapter alone is not a complete checkpoint.
- A local
LAYA_MODEL_PATHresponse has no Routerroutingmetadata. See Router vs one checkpoint. - Fine-tuning stays outside the image. The Kaggle notebook is the loop the doc points at. Delete the cache with
docker compose down --volumesonly when you mean to drop the weights.
Without Docker, Get started has pip install "laya[serve]". The workbench export is the same POST /v1/systemone body.