diff --git a/README.md b/README.md index 2e2f86e..a82d098 100644 --- a/README.md +++ b/README.md @@ -239,7 +239,7 @@ api/ FastAPI, classifier, agent, contracts, guardrails ingest/ Docling → chunk → embed → pg-ai, plus the Phase 9 upload worker web/ React + Vite operator UI -eval/ 78-case test set and the scorecard runner +eval/ the eval set and the scorecard runner scripts/ deploy.sh, verify.sh ``` diff --git a/spec/REBUILD.md b/spec/REBUILD.md index a7fae93..30f4921 100644 --- a/spec/REBUILD.md +++ b/spec/REBUILD.md @@ -181,7 +181,7 @@ build loads on a control-room PC and then fails every question on DNS. Only python eval/run_eval.py --api https://api.yokogawa.tech ``` -62 engineer-reviewable cases in [`eval/testset.jsonl`](../eval/testset.jsonl), every +Every case in [`eval/testset.jsonl`](../eval/testset.jsonl) is engineer-reviewable, and every data-dependent one with a **pinned time window** — `imh` is live, and an unpinned question gives a different answer each run. diff --git a/status/current-state.html b/status/current-state.html index 08841e1..079cec2 100644 --- a/status/current-state.html +++ b/status/current-state.html @@ -600,7 +600,7 @@ footer { margin-top: 72px; padding-top: 22px; border-top: 1px solid var(--hair); 08

The exam

Never run -
75 engineer-reviewable questions, every data-dependent one with a pinned time +
Engineer-reviewable questions, every data-dependent one with a pinned time window. Nothing blocks running it — the model account arrived on 27 August — and it has never been run as a gate. No scorecard exists. Two cases fail by design until the historian lands.
diff --git a/status/workflow-map.html b/status/workflow-map.html index 2193b24..33c4c86 100644 --- a/status/workflow-map.html +++ b/status/workflow-map.html @@ -651,7 +651,7 @@ shows the control-room machine opening the page on 28 August, the day the rule w three questions on 31 August, each answered. 8 -The exam — 78 engineer-checked questions run end to end and scored +The exam — every engineer-checked question run end to end and scored Runnable, and never run Nothing is blocking this and it has still not been done. The model account arrived on 27 August, so both AI steps can be scored; no scorecard has ever been produced. The set covers 28 @@ -714,7 +714,7 @@ returns that document’s heading, purpose and prerequisites and withholds t The rulebook already refused instruction language and still does; this removes the temptation rather than relying on catching it.
  • What it still does not prove: whether the labelling and wording are good enough. -They work; they are not yet scored. The 78-question exam can now be run and has not been. And no +They work; they are not yet scored. The exam can now be run and has not been. And no operator has yet sat at a control-room PC and got an answer — the one check that cannot be done from here.
  • Found today, and worth repeating: the most obvious question at this station —