FAQ¶
Is Jeff a chatbot / LLM like Llama?
No. It is a classifier: given a text and questions with predefined
answers (score, choice, noul), it returns probabilities. It does
not generate free text.
Do I need a GPU?
No. CPU is the supported runtime. On a modern 8-core CPU, typical requests take hundreds of milliseconds. GPUs are experimental.
Can I use the Radeon 780M / integrated AMD GPU?
Only experimentally, with ROCm, and the benefit for a 400 M model is uncertain. See AMD ROCm. CPU is recommended.
Does it work on Apple Silicon?
Yes, the linux/arm64 image runs natively under Docker Desktop, on CPU.
The Apple GPU (MPS) is not reachable from containers.
How much RAM do I need?
Roughly 3–6 GB for the large model depending on batch size (estimate,
check docker stats), plus OS/other services. 8 GB is
the minimum host, 16 GB comfortable. The base model needs less.
Why is the first start so slow?
It downloads ~1.7 GB of weights. Later starts reuse the volume and take seconds.
Why is the model not inside the image?
Smaller image, faster pulls, independent model updates, and you can choose/pin any checkpoint without rebuilding.
Is it compatible with the TypeSafe / jev SDK?
The wire format is compatible: set the SDK base URL to your instance and use one of your keys. Quality differs from the hosted jev service.
Can I run several models at once?
Run several instances, each with its own JEFF_MODEL_*, port and volume
prefix (see Docker Compose).
How do I add or revoke a client?
Edit JEFF_API_KEYS (comma-separated) and redeploy.
Does it send data anywhere?
No requests leave your host. The only outbound connections are the model download from Hugging Face (telemetry disabled) and image pulls from GHCR.
Why does /v1/models say gliformer-large-v1 while I run the base model?
The reported name comes from JEFF_MODEL_NAME (upstream default). Set
JEFF_MODEL_NAME=gliformer-base-v1.
latest or a version tag?
Version tags in production (JEFF_DOCKER_TAG=0.2.0), latest for tests.
Can I run it on Kubernetes?
Yes: one Deployment with the image, a PVC on /models, a Secret for
JEFF_API_KEYS, readiness/liveness probes on /healthz with a generous
initial delay. Not officially documented yet; PRs welcome.
Where do I report wrong answers from the model?
Upstream: logan-markewich/jeff. Packaging and deployment issues: this repository.