Environment variables
All configuration happens through environment variables (.env for
Compose, the Environment tab in Dokploy, -e for docker run). Empty
values are dropped by the entrypoint so upstream defaults apply.
jeff-docker (image and entrypoint)
Variable
Default
Description
JEFF_API_KEYS
— required
Comma-separated bearer keys. Container exits if empty
JEFF_ALLOW_NO_AUTH
false
Allow empty keys. Local tests only
JEFF_MODEL_ID
knowledgator/gliformer-large-v1
Hugging Face repo to download
JEFF_MODEL_PATH
/models/gliformer-large-v1
Where the weights live (inside the /models volume)
JEFF_MODEL_REVISION
main
HF branch, tag or commit; pin a commit for reproducibility
JEFF_DOWNLOAD_MODEL
true
Download when missing. false = fail if the directory is empty
JEFF_FORCE_MODEL_DOWNLOAD
false
Download again on next start (then set back to false)
HF_TOKEN
—
Hugging Face token for private/gated repos or higher rate limits
Compose / container
Variable
Default
Used by
Description
JEFF_DOCKER_TAG
latest
prebuilt, dokploy
GHCR tag to run
JEFF_REF
main
local build
Upstream Jeff ref
UV_EXTRAS
onnx
local build
Upstream extras
TORCH_INDEX_URL
CPU index
local build
PyTorch wheel index
JEFF_CONTAINER_NAME
jeff
local, prebuilt
Container name
JEFF_VOLUME_PREFIX
jeff
local, prebuilt
Volume names <prefix>-models, <prefix>-data
JEFF_BIND
127.0.0.1
local, prebuilt
Host interface for the port
JEFF_HOST_PORT
8000
local, prebuilt
Host port
JEFF_CPU_LIMIT
0 (no limit)
all
CPU limit (cpus). Must be ≤ CPUs visible to Docker (nproc), otherwise the deploy fails with "range of CPUs is from 0.01 to N"
JEFF_MEM_LIMIT
10g
all
Memory limit (mem_limit)
Jeff runtime (upstream)
Variable
Compose default
Upstream default
Description
JEFF_DEVICE
cpu
auto (cuda → mps → cpu)
Inference device
JEFF_BACKEND
torch
torch
torch or onnx (see ONNX )
JEFF_QUANT
—
fp32
ONNX precision: fp32 or int8
JEFF_ONNX_PATH
—
<model>/onnx/encoder[.int8].onnx
Custom ONNX encoder path
JEFF_DTYPE
—
bf16 on CUDA, float32 on CPU
Torch dtype
JEFF_THREADS
4
all cores
Inference threads (also sets OMP_NUM_THREADS, MKL_NUM_THREADS)
JEFF_WARMUP
true
off
Run a warm-up inference at start (true/1)
JEFF_COMPILE / JEFF_COMPILE_MODE
—
off
torch.compile (GPU, advanced)
JEFF_ATTN
—
auto
Attention kernel: auto, flash, eager
JEFF_PAD_MULTIPLE
—
0
Pad sequence length to a multiple (GPU tuning)
JEFF_MODEL_NAME
—
gliformer-large-v1
Name reported by /v1/models and responses
JEFF_MODEL_ALIASES
—
jev-latest,jev
Accepted model values
Batching, limits and protection
Variable
Compose default
Upstream default
Description
JEFF_MAX_BATCH
8
16
Max requests per inference batch
JEFF_MAX_WAIT_MS
10
5
Time to wait for a batch to fill
JEFF_MAX_QUEUE
—
256
Queue size; beyond it requests get 529
JEFF_RATE_LIMIT_RPS
4
0 (off)
Per-key requests per second
JEFF_RATE_LIMIT_BURST
8
20
Per-key burst
JEFF_MAX_QUESTIONS
—
64
Max questions per request (422 above)
JEFF_MAX_LABELS
—
64
Max labels per question
JEFF_MAX_STATE_CHARS
—
20000
Max state length
Prompt and calibration (advanced)
Variable
Upstream default
Description
JEFF_NOUL_MODE
yes_no
single, single_named, yes_no. Use single with gliformer-base-v1
JEFF_ISOLATE
nouls
none, nouls, all — separate encoder passes per question
JEFF_STATE_FORMAT
kv
How a JSON state is rendered: kv, json, values
JEFF_TEMPERATURE
3.2
Probability calibration; 1.0 disables scaling
Fixed by the image
Variable
Value
JEFF_HOST / JEFF_PORT
0.0.0.0 / 8000
HF_HOME
/data/huggingface
XDG_CACHE_HOME
/data/cache
HF_HUB_DISABLE_TELEMETRY
1
Profiles