Jeff Docker¶
A production-oriented, CPU-first Docker distribution for Jeff — a self-hosted, TypeSafe jev-compatible structured-classification API powered by GLiFormer.
You send a piece of text (state) and a set of questions (score,
choice, noul); Jeff answers with probabilities, a choice or a score.
It is not a chat LLM: it is a fast, deterministic classifier you can run
on a normal CPU, at home or on a server.
Why this project¶
- One command to run: prebuilt multi-arch image on GHCR
(
linux/amd64,linux/arm64), no Python toolchain on your machine. - CPU-only PyTorch: the image ships CPU wheels instead of the default CUDA stack: about 0.4 GB to pull instead of 3.2 GB (measured on amd64).
- Model downloaded once: weights land in a persistent volume and are reused on every restart, redeploy or upgrade.
- Secure by default: the container refuses to start without an API key, runs as a non-root user, with no Linux capabilities and no host port exposure in Dokploy.
- Dokploy-first: dedicated Compose file and a step-by-step guide, including Proxmox LXC.
- Tested: shellcheck, hadolint, bats unit tests, Compose validation and an end-to-end smoke test that runs the real API on every change (see Testing and CI).
- GPU optional: AMD ROCm and NVIDIA CUDA documented as experimental overlays, never required.
How it works¶
- On start, the entrypoint checks
JEFF_API_KEYS(mandatory unless you explicitly enable no-auth mode for local tests). - If the model is not yet in the
/modelsvolume, it downloads it from Hugging Face once and writes a marker file; later starts skip the download. - If
JEFF_BACKEND=onnx, Jeff exports the ONNX encoder into the same volume on first start and reuses it afterwards. - Jeff starts on port
8000with dynamic batching, per-key rate limits and a/healthzendpoint used by the DockerHEALTHCHECK.
Where to start¶
New to Docker?
Follow the Quick start for beginners: every step is explained, from installing Docker to your first API call.
| I want to... | Go to |
|---|---|
| Try it on my laptop in 5 minutes | Quick start |
| Install Docker on macOS, Windows or Linux | Platform setup |
| Use the prebuilt image and pin versions | Prebuilt image |
| Run it with Docker Compose on a server | Docker Compose |
| Deploy on Dokploy (incl. Proxmox LXC) | Dokploy |
| Run it without Compose | Plain Docker |
| Call the API from curl, Python, TypeScript, Home Assistant | Using the API |
| Tune every option | Environment variables |
| Choose or pin a model, run offline | Models |
| Make it faster on CPU | ONNX backend and Performance |
| Back up, upgrade, roll back | Persistence and Upgrades |
| Fix a problem | Troubleshooting and FAQ |
Support policy¶
| Runtime | Status |
|---|---|
CPU, linux/amd64 (Linux, Proxmox LXC, Dokploy, Windows/WSL2, Intel Mac) |
Supported |
CPU, linux/arm64 (Apple Silicon via Docker Desktop, ARM servers, Raspberry Pi 5) |
Supported |
| NVIDIA CUDA | Experimental |
| AMD ROCm, integrated Radeon GPUs, GPU passthrough in LXC | Experimental |
| Apple MPS | Not available inside Docker (Linux VM); run Jeff natively on macOS instead |
Repository¶
Source code, issues and pull requests: github.com/tommasomarchionni/jeff-docker. Model behavior and API questions belong upstream: github.com/logan-markewich/jeff.