Engine room

live · 30s

The serving infrastructure behind this site’s assistant, deliberately overengineered and left visible. Every number below is real and current; every answer links to its own trace.

All lanes nominal: the canary is green, the queue is idle, 302 answers served so far.

Request pipeline · all time

since 2026-09-05

  1. Requests

    670

    arrived at the edge

  2. Admission

    304

    rate + token budgets

    320 shed

  3. Cache

    889

    answered without a model

    31 dedup reattached

  4. Guardrail

    10

    scoped refusals

    97 failed open

  5. Generate

    302

    durable workflow, ladder armed

    62 degraded to raw KB

  6. Verify

    302

    citations resolve

Time to first token

p50 593ms · p95 8.20s · 304 samples

target 2.5s0ms–409ms: 83 answers409ms–819ms: 125 answers819ms–1.2s: 8 answers1.2s–1.6s: 3 answers1.6s–2.0s: 6 answers2.5s–2.9s: 8 answers2.9s–3.3s: 37 answers3.3s–3.7s: 10 answers3.7s–4.1s: 1 answer4.1s–4.5s: 1 answer5.3s–5.7s: 2 answers5.7s–6.1s: 2 answers6.1s–6.5s: 1 answer7.0s–7.4s: 1 answer8.2s–8.6s: 1 answer8.6s–9.0s: 4 answers9.0s–9.4s: 4 answers9.4s–9.8s: 1 answer9.8s–10.2s: 1 answer10.2s–10.6s: 1 answer12.3s–12.7s: 1 answer12.7s–13.1s: 1 answer13.1s–13.5s: 1 answer13.9s–14.3s: 1 answerp50 593msp95 8.2s014.7s

Goodput

100.0% · 302 ok / 0 failed

Full answer

p50 3.97s · p95 16.25s

End-to-end through guardrail, retrieval, generation, and verification, measured inside the workflow.

Ledger · all time

302 answers served for $0.05 in model spend ($0.00016 per answer), 341,691 tokens through the gateway, 920 requests absorbed by the cache and dedup layers before spending anything.

Recent answers

timeoutcomettfttotalcited
09-11 08:40glm-5.3-flash8.20s15.98strace →
09-11 08:40glm-5.3-flash9.18s16.59strace →
09-11 08:34glm-5.3-flash10.25s21.91strace →
09-09 18:41glm-5.3-flash12.43s15.35strace →
09-07 05:00glm-5.3-flash13.47s23.10strace →
09-06 17:52glm-5.3-flash8.88s18.95strace →
09-06 16:48refused: off topic7.19s7.51strace →
09-06 16:48glm-5.3-flash2.93s3.36strace →
09-06 16:48glm-5.3-flash2.81s3.46strace →
09-06 16:48glm-5.3-flash2.88s5.23strace →

Register

primary model
zai/glm-5.3-flash
model override
none
prompt hash
45b79cdd7ce5
policy hash
6bb17bbe0908
breaker · primary
closed
breaker · fallback
closed
eval gate
27/32 · 2026-09-06
canary
up · 2.4s · 00:29 UTC
daily token ceiling
0 / 2,000,000 (0.0%)
kill switch
off

Counters accumulate for the site’s lifetime; latency distributions cover the last 500 answers; traces are kept for seven days and never contain visitor message content. The architecture behind this page is the subject of the writing.