Skip to content

ONNX backend (faster CPU inference)

Jeff can run the encoder through ONNX Runtime, optionally quantized to int8. The RNN and classification heads stay in PyTorch, so this is a partial acceleration, but the encoder is the expensive part.

The image already contains ONNX Runtime (UV_EXTRAS=onnx): no rebuild needed.

Enable it

JEFF_BACKEND=onnx
JEFF_QUANT=int8      # or fp32
JEFF_THREADS=6

Redeploy. On the first start Jeff exports the encoder into <JEFF_MODEL_PATH>/onnx/ (≈10–60 s and ~0.7–1.6 GB extra disk), then reuses it:

INFO jeff.onnx: .../onnx/encoder.int8.onnx missing; exporting the encoder now
INFO jeff.onnx: onnx encoder ready: .../encoder.int8.onnx quant=int8 providers=['CPUExecutionProvider']

What to expect

Measured by this project on gliformer-base-v1, 2 vCPU (Intel Xeon 2.9 GHz), 3 questions on a short text, warm server:

Backend Latency per request
torch fp32 ~190 ms
onnx int8 ~110 ms

int8 changes probabilities slightly (quantization): validate on your own data before switching production traffic. Re-run the benchmark on your hardware.

Exporting manually (optional)

docker compose exec jeff python /app/scripts/export_onnx.py /models/gliformer-large-v1 --int8

Going back

Set JEFF_BACKEND=torch and redeploy. The ONNX files can stay; delete them with docker compose exec jeff rm -rf /models/gliformer-large-v1/onnx.