Using the API¶
Jeff exposes a small HTTP API compatible with the TypeSafe jev wire
format. All /v1/* endpoints need Authorization: Bearer <JEFF_API_KEYS
entry>.
| Method | Path | Auth | Purpose |
|---|---|---|---|
POST |
/v1/systemone |
yes | Answer classification questions about a text |
GET |
/v1/models |
yes | List model name and aliases (jev-latest, jev) |
GET |
/healthz |
no | Liveness/readiness (used by the Docker healthcheck) |
GET |
/stats |
no | Batcher counters and active configuration |
Errors: 401 missing/invalid key, 422 validation error or request
limit exceeded, 429 rate limit (retry-after-ms header), 529
queue full. Responses carry x-typesafe-request-id, x-jeff-server-ms and
x-jeff-batcher-ms headers.
Request anatomy¶
{
"state": "The text to classify (string, or a JSON object)",
"model": "jev-latest",
"questions": {
"<your-question-id>": { "type": "score | choice | noul", "instructions": "...", "criteria": ... }
}
}
score — ordinal scale¶
criteria is an ordered list of levels (low → high). The answer has a
continuous score (0 … n-1), per-level probabilities, confidence and a
legend.
"severity": {"type": "score", "instructions": "How severe is this?",
"criteria": ["cosmetic", "degraded", "blocking"]}
choice — pick one label¶
criteria is an object: label → description (or null).
"area": {"type": "choice", "instructions": "Which area is affected?",
"criteria": {"frontend": "UI and browser", "backend": "API and server", "infra": null}}
noul — yes/no probability¶
No criteria needed; returns noul between 0 and 1.
"regression": {"type": "noul", "instructions": "Is this a regression?"}
Example response¶
{
"model": "gliformer-large-v1",
"answers": {
"severity": {"type": "score", "score": 1.27, "confidence": 0.08,
"legend": {"0": "cosmetic", "1": "degraded", "2": "blocking"},
"probabilities": {"0": 0.30, "1": 0.31, "2": 0.39}},
"area": {"type": "choice", "choice": "frontend", "confidence": 0.56,
"probabilities": {"frontend": 0.71, "backend": 0.15, "infra": 0.14}},
"regression": {"type": "noul", "noul": 0.15}
},
"usage": {"input_tokens": 77, "output_tokens": 18}
}
Clients¶
curl -s https://jeff.example.com/v1/systemone \
-H "Authorization: Bearer $JEFF_KEY" -H "Content-Type: application/json" \
-d '{"state":"Payment page times out","model":"jev-latest",
"questions":{"urgent":{"type":"noul","instructions":"Is this urgent?"}}}' | jq
import os, httpx
client = httpx.Client(
base_url="http://127.0.0.1:8000",
headers={"Authorization": f"Bearer {os.environ['JEFF_KEY']}"},
timeout=30,
)
r = client.post("/v1/systemone", json={
"state": "Payment page times out",
"model": "jev-latest",
"questions": {
"severity": {"type": "score", "criteria": ["low", "medium", "high"]},
},
})
r.raise_for_status()
print(r.json()["answers"]["severity"]["score"])
type ScoreAnswer = { type: "score"; score: number; confidence: number;
probabilities: Record<string, number> };
const res = await fetch("http://127.0.0.1:8000/v1/systemone", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.JEFF_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
state: "Payment page times out",
model: "jev-latest",
questions: { severity: { type: "score", criteria: ["low", "medium", "high"] } },
}),
});
if (!res.ok) throw new Error(`Jeff ${res.status}: ${await res.text()}`);
const { answers } = (await res.json()) as { answers: { severity: ScoreAnswer } };
console.log(answers.severity.score);
# configuration.yaml
rest_command:
jeff_classify:
url: "http://jeff.home.lan/v1/systemone"
method: POST
headers:
Authorization: !secret jeff_bearer # "Bearer <key>"
Content-Type: application/json
payload: >
{"state": {{ text | tojson }}, "model": "jev-latest",
"questions": {"urgent": {"type": "noul", "instructions": "Is this urgent?"}}}
timeout: 15
Call it from an automation with response_variable and branch on
response.content.answers.urgent.noul > 0.7.
The wire format is compatible with the official jev SDK: point the SDK
base URL to your Jeff instance and use one of your JEFF_API_KEYS as the
API key. Model behavior differs from the hosted jev service (see below).
Differences from hosted jev¶
From upstream: probabilities are temperature-scaled sigmoids
(JEFF_TEMPERATURE, default 3.2); noul questions get separate encoder
passes while choice/score share one (set JEFF_ISOLATE=all for full
independence at extra cost); token counts are not comparable to jev billing.
Upstream also reports lower accuracy than jev on harder reasoning tasks.
Good practices¶
- Be explicit in
instructionsand label descriptions: they are part of the prompt. - Keep
stateshort (limit:JEFF_MAX_STATE_CHARS, default 20 000): latency grows with length. - Group questions about the same text in one request rather than sending many requests.
- Handle 429/529 with retry and backoff (honor
retry-after-ms). - Use one key per client (comma-separated in
JEFF_API_KEYS) so you can revoke them independently.