From 45034eb4d1dd603b65b2cd6768c138bc5008d4d5 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 16:22:23 +1000 Subject: [PATCH] Correct two counts that disagreed with the files they describe The exam: eval/testset.jsonl holds 78 cases (A10 H28 L8 N5 P10 R10 T5 U2). workflow-map.html called it a 75-question exam in the two places it states the current total. Corrected. The two sentences describing how the exam "grew from 67 questions to 75" are left alone - that step is historically correct: 67 plus the eight live-model failures is 75, and the Phase 5 findings H26, H27 and H31 took it to 78 afterwards. The env files: .env.example said TWO 0600 files under ~/ai/ and listed pg-ai.env and api.env, but its own line 105 refers to langfuse.env, and README.md, scripts/deploy.sh and compose/langfuse-compose.yml all use three. The header was simply wrong; langfuse.env is now named in it. Co-Authored-By: Claude Opus 5 --- .env.example | 7 ++++--- status/workflow-map.html | 4 ++-- 2 files changed, 6 insertions(+), 5 deletions(-) diff --git a/.env.example b/.env.example index d579c8a..6a02975 100644 --- a/.env.example +++ b/.env.example @@ -1,9 +1,10 @@ # ============================================================================= # .env.example — every key, no values. Committed deliberately. # -# On lin001 these live as TWO 0600 files under ~/ai/, never in Git: -# ~/ai/pg-ai.env the POSTGRES_* / AGENT_DB_* block -# ~/ai/api.env everything else +# On lin001 these live as THREE 0600 files under ~/ai/, never in Git: +# ~/ai/pg-ai.env the POSTGRES_* / AGENT_DB_* block +# ~/ai/langfuse.env the Langfuse SALT / NEXTAUTH_SECRET block (see below) +# ~/ai/api.env everything else # Follow the ~/authelia/authelia.env precedent: chmod 0600, owned by azureuser. # ============================================================================= diff --git a/status/workflow-map.html b/status/workflow-map.html index d4ff079..2193b24 100644 --- a/status/workflow-map.html +++ b/status/workflow-map.html @@ -651,7 +651,7 @@ shows the control-room machine opening the page on 28 August, the day the rule w three questions on 31 August, each answered. 8 -The exam — 75 engineer-checked questions run end to end and scored +The exam — 78 engineer-checked questions run end to end and scored Runnable, and never run Nothing is blocking this and it has still not been done. The model account arrived on 27 August, so both AI steps can be scored; no scorecard has ever been produced. The set covers 28 @@ -714,7 +714,7 @@ returns that document’s heading, purpose and prerequisites and withholds t The rulebook already refused instruction language and still does; this removes the temptation rather than relying on catching it.
  • What it still does not prove: whether the labelling and wording are good enough. -They work; they are not yet scored. The 75-question exam can now be run and has not been. And no +They work; they are not yet scored. The 78-question exam can now be run and has not been. And no operator has yet sat at a control-room PC and got an answer — the one check that cannot be done from here.
  • Found today, and worth repeating: the most obvious question at this station —