Piko API
Base URL & auth
Piko is an HTTP service in front of your own OpenAI-compatible engine. It never holds a second model: the weights that answer your chat are the weights that decide. The service binds one host and port and speaks two routes over the same process.
http://127.0.0.1:8787
When the server is started with an API key, the POST routes become bearer-guarded. The refusal uses the Jev `401` shape, so an SDK parses it without changes:
Authorization: Bearer <your-key>
Quickstart
# 1. point the service at your engine, then start it
export NOTJEV_BASE_URL=http://127.0.0.1:8080
export NOTJEV_MODEL=gemma-4-12b
notjev serve --port 8787 &
# 2. wait for health
until curl -sf localhost:8787/health >/dev/null; do sleep 0.2; done
# 3. read a decision
curl -s localhost:8787/v1/decide -H 'content-type: application/json' -d '{
"state": "The minister announced a plan on Tuesday.",
"questions": [
{ "id": "kind", "question": "event type", "options": ["ANNOUNCE", "MEET", "VOTE"] }
]
}'
POST /v1/decide
The native route. One state, N questions, typed answers — a verdict, a margin and an
abstention per question. All N questions share the state, so the engine's prefix cache makes the
tail nearly free; the readout is still one token per question.
/v1/decideRequest body
| Field | Type | Description |
|---|---|---|
state | string | Required. Everything the model may see — a document, a rendered game board, a conversation. |
questions | array | Required. One object per question: { id, question, options }, { id, question, noul } or { id, question, score }. |
theta | number | Abstention threshold for the whole request, in [0, 1). Defaults to the server's value. |
concurrency | number | How many questions to fan out at once. Defaults to the question count. |
raw | boolean | Include the heavy fields (entries, prompt, the upstream body). |
Example request
curl -s localhost:8787/v1/decide -H 'content-type: application/json' -d '{
"state": "The minister announced a plan on Tuesday.\n",
"theta": 0.5,
"questions": [
{ "id": "kind", "question": "event type", "options": ["ANNOUNCE", "MEET", "VOTE"] },
{ "id": "policy", "question": "is this about policy?", "noul": { "yes": "YES", "no": "NO" } },
{ "id": "depth", "question": "how concrete is it?", "score": 5 }
]
}'
Example response
{
"model": "gemma-4-12b",
"theta": 0.5,
"ms": 74,
"results": [
{
"id": "kind", "kind": "choice", "choice": "ANNOUNCE", "value": "ANNOUNCE",
"p1": 0.990, "p2": 0.008, "margin": 0.982, "band": "certain", "prior": 0.95,
"coverage": 0.991, "degraded": false, "undecided": false, "theta": 0.5,
"probabilities": [0.990, 0.008, 0.002],
"options": ["ANNOUNCE", "MEET", "VOTE"],
"ms": 25, "model": "gemma-4-12b"
}
]
}
coverage is 0 and degraded: true. Piko returns
choice: null — the uniform has margin 0, and at theta = 0 it would
otherwise pick the first option “with full confidence in nothing”.
POST /v1/systemone
The Jev wire contract — the one-endpoint contract of TypeSafe's Jev and the OpenJev
ecosystem. A typesafe-sdk (or anything that speaks the contract) pointed at this server
works unchanged, against any engine you configure. Answers come back grouped by type.
/v1/systemoneRequest body
| Field | Type | Description |
|---|---|---|
state | string | Required. The context, as a string. |
questions | object | Required. Named questions: { name: { type, instructions, criteria } }. |
theta | number | Abstention, in [0, 1). A Piko extension — the wire contract has no abstention. |
concurrency | number | Fan-out width. |
The three type values
choice— needscriteria: { NAME: "description", … }. At most 26 options; more is refused422, never truncated.noul— optionalcriteria: { true: "…", false: "…" }. ReturnsP(true).score— needscriteria: [level0, level1, …], 2–26 named levels. Returns the 0-indexed expectation.
Example request
curl -s localhost:8788/v1/systemone -H 'content-type: application/json' -d '{
"state": "The minister announced a plan on Tuesday.",
"questions": {
"policy": { "type": "noul", "instructions": "is this about policy?" },
"kind": { "type": "choice", "instructions": "event type",
"criteria": { "ANNOUNCE": "someone makes something public",
"MEET": "two parties meet",
"VOTE": "a ballot is held" } }
}
}'
Example response
{
"model": "gemma-4-12b",
"answers": {
"nouls": {
"policy": { "type": "noul", "noul": 0.971, "confidence": 0.877,
"notjev": { "margin": 0.942, "band": "certain", "coverage": 0.993, "degraded": false,
"undecided": false, "theta": 0.5 } }
},
"choices": {
"kind": { "type": "choice", "choice": "ANNOUNCE", "confidence": 0.842,
"probabilities": { "ANNOUNCE": 0.99, "MEET": 0.008, "VOTE": 0.002 },
"notjev": { "margin": 0.982, "band": "certain", "coverage": 0.991 } }
},
"scores": {}
},
"usage": { "input_tokens": 142, "output_tokens": 2 }
}
422 with a named reason (only the first 26 letters can be presented), and an
upstream failure is a 502 for the whole request — never half-filled answers an SDK
would not notice.
GET /v1/models
Lists the configured engine and the aliases an SDK may ask for by default. They all resolve to the same engine.
/v1/models{
"object": "list",
"data": [
{ "object": "model", "id": "gemma-4-12b" },
{ "object": "model", "id": "notjev-latest" },
{ "object": "model", "id": "jev-latest" },
{ "object": "model", "id": "jev-preview" }
]
}
GET /health
Liveness and the configuration the process is running with — the first thing to check before blaming a verdict.
/health{ "ok": true, "baseUrl": "http://127.0.0.1:8080", "model": "gemma-4-12b", "theta": 0.5 }
Question forms
There are three shapes of question, and only three. In each case you write the possible answers down first — Piko can only ever return one of them, so there is no way for it to invent a value you did not offer. That is the property that makes an answer safe to act on.
| Form | Plain name | You send | You get back |
|---|---|---|---|
| choice | Pick one from a list | options: ["A1","B1", …] (≤ 26) | choice — your code, in your order |
| noul | True or false | noul: true or noul: { yes, no } | value — a boolean |
| score | A score on a scale | score: 5 or score: { min, max } | value — the top grade; expectation — the mean over the whole scale |
A true / false question named noul may also be sent as
"type": "boolean" — the two are the same form.
{ id, description }. A description renders
id: description in the menu and changes the prompt, therefore the
measurement — the published numbers use bare codes.
The decision object
What a single question returns. Fields marked extension live in the
notjev block on the wire contract — they are the instrument the contract has no room
for.
| Field | Type | Meaning |
|---|---|---|
choice | string | null | The code, or null — null is the absence of a verdict, not a fallback. |
value | any | The re-typed answer: the code, a boolean, an integer. |
expectation | number | null | score only: the mean grade under the whole distribution. |
top | string | What the model would have said without the margin — readable, never applied. |
p1, p2, margin | number | The odds of the top two options, and the gap between them, after rescaling to your menu. extension |
band | string | low < 0.5 ≤ med < 0.75 ≤ high < 0.9 ≤ certain. extension |
prior | number | The middle of the band — what stays true across engines, unlike the raw float. |
coverage | number | How much of the model's answer landed on one of your options, before rescaling. Low = it was thinking about something else. extension |
exactMass, spacedMass | number | The two ways the letter was written ("A" vs " A"), counted apart so the matching rule can be audited. |
confidence | number | 1 − H(p)/ln K over the rescaled menu: 1 certain, 0 a coin toss between all options. |
degraded | boolean | true when no option letter appeared. Then choice is always null. |
undecided | boolean | margin < theta, or degraded. extension |
probabilities | object | Per option, summing to 1. |
ms, usage, model | — | The call itself. |
prompt, request, raw | — | The exact string, the exact body, the server response — only with raw: true. |
Errors
Errors are named: the caller never has to guess from a stack. The 422
shape is FastAPI's, because that is what Jev clients parse.
| Status | Shape | When |
|---|---|---|
| 400 | { error: { code, message } } | Malformed body, a missing question, a question form the readout cannot honour. |
| 401 | { detail: { error_type, message } } | The server carries an API key and the request did not present it. |
| 422 | { detail: [ { loc, msg, type } ] } | Wire-contract validation: a bad type, a menu of more than 26 options, a malformed criteria. |
| 502 | { detail: { error_type, message } } | The engine failed. All-or-nothing: the whole request fails, never a half-filled answer. |
| 413 | { error: { code } } | Body over the limit (8 MB by default). |
Abstention & theta
An answer is only returned when margin ≥ theta. Under that bar, choice
is null and the question stays pending — ask it again when things have
changed, rather than accepting a guess. theta is the one dial worth tuning: raise it
and you get fewer answers that are right more often.
| theta | Behaviour |
|---|---|
| 0 | Always an answer. On the entity-matching bench, every new entity was wrongly merged. |
| 0.5 | Recommended default. Of the answers it gave, 92.7% were right, and it correctly held back on 71.6% of the ones it should not have answered. |
| 0.8+ | Conservative. Use where a wrong answer costs far more than asking again. |
Limits
Everything the service will not do, stated plainly. A refusal here is deliberate and always named in the response, never a silent guess.
| Limit | Value | What happens at the edge |
|---|---|---|
Options per choice | 26 (the letters A–Z) | 422, refused whole. Never truncated. |
| Questions per request | 26 | 422 on the playground API; the answer is all-or-nothing. |
score levels | 2–26 named levels | Fewer than two is 422. |
| Request body | 8 MB | 413. Media rides inline, so size it before you send it. |
| Context window | 8,192 tokens on the card here | Set by the engine (--max-model-len), not by Piko. A longer state is the engine's error, surfaced as 502. |
Media in /v1/decide | string state only | The native route takes a text state. For an image or a recording use the playground API (array of parts) or the gateway's /notjev/context, then decide against the snapshot. |
| Degraded readouts | no option letter in the mass | coverage: 0, degraded: true, choice: null. |
| Abstention | theta in [0, 1) | Below the bar, choice: null and undecided: true. Nothing is written down. |
| Output | 1 token per question | It never writes prose. There is nothing to parse and nothing to hallucinate. |
| Text-only state | a string | state must be a string, or an array of parts when media is present. Anything else is 422. |
Known limitations
- The numbers are provenance, not a promise. Measured latency and cost come from one card; your engine, your model and your prompt length will differ.
- Letter prior. A model may favour early letters (
AoverM). The repository ships a correction (NOTJEV_LETTER_PRIOR); it is off unless you set it. - Multimodal is quantizer-dependent. The vision path has failed silently on some quantized exports (every image read the same). Verify one image before trusting a run.
- Determinism. Above a high confidence, answers are near-stable across
machines; store the
band, not the raw float.
Clients · cURL
curl -s localhost:8787/v1/decide \
-H 'content-type: application/json' \
-H 'authorization: Bearer $NOTJEV_API_KEY' \
-d '{"state":"…","questions":[{"id":"q","question":"…","options":["A","B"]}]}'
Clients · JavaScript
Works in Node ≥ 20, using the fetch that Node already ships.
const { createClient } = require('notjev');
const jev = createClient(); // NOTJEV_BASE_URL, NOTJEV_MODEL, NOTJEV_API_KEY, NOTJEV_THETA
const a = await jev.decide({
state : 'A: "Sarah Knafo"\nB: "Sarah Knafot" (seen in an audio transcript)\n',
question: 'verdict',
options : ['SAME', 'OTHER'],
theta : 0.5,
});
console.log(a.choice, a.p1.toFixed(3), a.band, a.coverage.toFixed(3));
// null 0.500 0.000 med 0.983 <- exactly torn on this pair: no verdict at theta 0.5
const b = await jev.noul('The minister announced a plan.\n', 'is this about policy?',
{ yes: 'YES', no: 'NO' });
console.log(b.value, b.p1.toFixed(3)); // true 0.970
const s = await jev.score('The minister announced a plan.\n', 'how concrete is it?', 5);
console.log(s.value, s.expectation.toFixed(2)); // 3 3.41
Talk to the service over HTTP instead
const r = await fetch('http://127.0.0.1:8787/v1/decide', {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({
state: state,
questions: [{ id: 'kind', question: 'event type', options: ['ANNOUNCE', 'MEET', 'VOTE'] }],
}),
});
const { results } = await r.json();
Clients · Python
Nothing Python-specific is required — it is JSON over HTTP, so requests or
httpx is enough:
import httpx
r = httpx.post("http://127.0.0.1:8787/v1/decide", json={
"state": "The minister announced a plan on Tuesday.",
"theta": 0.5,
"questions": [
{"id": "kind", "question": "event type", "options": ["ANNOUNCE", "MEET", "VOTE"]},
{"id": "draft", "question": "is this a draft?", "noul": {"yes": "YES", "no": "NO"}},
],
})
for row in r.json()["results"]:
print(row["id"], row["choice"], row["margin"], row["undecided"])
Running the server
notjev serve [--port 8787] [--host 127.0.0.1]
[--base-url URL] [--model ID] [--api-key KEY] [--theta 0.5]
Two more serving surfaces exist for decisions over a conversation — a gateway and an MCP server —
configured with the same NOTJEV_BASE_URL / NOTJEV_MODEL:
notjev gateway --port 8789 & # Chat Completions proxy + /notjev/decide, /notjev/context/*
notjev mcp # stdio: notjev_decide, notjev_context_put, notjev_context_drop
Engines & environment variables
| Variable | Used for |
|---|---|
NOTJEV_BASE_URL | The OpenAI-compatible engine (vLLM, llama.cpp, Ollama, OpenAI). |
NOTJEV_MODEL | The model name the engine serves. |
NOTJEV_API_KEY | Bearer guards the POST routes when set. |
NOTJEV_THETA | Default abstention threshold. |
NOTJEV_LETTER_PRIOR | A JSON array dividing out the letter prior (see the harness). Malformed is refused, never guessed. |
Engine quickstarts
| Engine | Setup |
|---|---|
| vLLM | vllm serve <model> --max-model-len 32768 --max-num-seqs 4 — logprobs are on by default. |
| llama.cpp | llama-server -m model.gguf --port 8080 — chat logprobs since PR #10783. |
| Ollama | ollama ≥ 0.12.11 on /v1/chat/completions. Use an instruct model, not a thinking one. |
| OpenAI | Remove the template switch (--no-template-kwargs), it rejects unknown body fields. |
As a service
notjev serve --port 8788 & # /v1/decide AND /v1/systemone, one process
curl -s localhost:8788/v1/systemone -H 'content-type: application/json' -d '{ … }'
Playground API
The Piko web app is a thin client over a small JSON API. It is handy for demos and internal tools, and it never holds a key of its own — it forwards to your engine.
/api/healthEngine reachability, the served models, the configured theta.
{ "ok": true, "theta": 0.5,
"engine": { "baseUrl": "http://127.0.0.1:8080", "model": "gemma-4-12b",
"reachable": true, "served": ["gemma-4-12b"] } }
/api/examplesThe curated example catalog { packs, examples: [{ id, pack, title, state, questions }] }.
/api/promptRender the exact prompt for each question, sending nothing to the engine.
// in
{ "state": "…", "questions": [{ "id": "q", "type": "choice", "instructions": "…", "options": ["A","B"] }] }
// out
{ "ok": true, "prompts": { "q": "Choose the correct option. Reply with only its letter.\n\nContext:\n…" } }
/api/decideRead the decisions on your engine. Same body as /api/prompt, plus an optional
theta and concurrency; returns the /v1/decide rows plus totals.
{ "ok": true, "model": "…", "theta": 0.5, "ms": 74, "perQuestionMs": 24.6,
"results": [ { "id": "q", "choice": "A", "p1": 0.99, "margin": 0.98, "band": "certain",
"coverage": 0.99, "undecided": false, "probabilities": { "A": 0.99, "B": 0.01 } } ] }
Piko sits on the project's readout library — no second model, no hosted key, no per-call cost. The benchmark numbers quoted here come from a production judge and are provenance, not a promise about your questions.