yau-plant-assistant/CLAUDE.md
Claude 73e90cf710 Say what PS_* was, everywhere it is still mentioned
Seven mentions remain across the repository and none of them is data.
Each is a sentence recording that these names were wrong, what it cost,
and that they must not come back: the namespace tables in CLAUDE.md and
the build spec, the generator's docstring, the fixtures rewrite note, the
Phase 5 narrative in REBUILD.md and current-state.html, and the delivered
configuration's own README.

The build spec's namespace table carried the same error CLAUDE.md did -
PS_STN_WET_WELL_LEVEL listed as the CI Server point layer, sourced from a
file that was a PLC-side proposal. Both now name the real point layer,
WRPS_PLC:STN_LEVEL from wrps_modbus_point_df.qli, and both point at
db/seed/scada-source/ rather than at a path on one person's laptop.

The name-collision entry in the build spec is marked closed, and records
that it bit a second time during the rename itself.

These notes explain; they do not enforce. What actually caught the
collision was check_mapping() refusing to write, and what stops a retired
name being queried is the guard in equipment.resolve(). The prose is here
so the next person does not have to re-derive the story - not as a
control.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 19:00:57 +10:00

8.2 KiB
Raw Blame History

CLAUDE.md — rules to keep front of mind

Plant Operations Assistant for the Waterloo Road Pump Station (WRPS), deployed onto an existing, live, shared Docker host. Full detail: spec/BUILD-AI-CONTAINERS.md (the build spec) and spec/YAU_Linux_Host_Onboarding.md (the host brief). Where the two conflict, the host brief wins.

The three lines this system does not cross

  1. No instructions for safety-critical actions. Procedural questions get document identity, revision, effective date and verbatim prerequisites — never synthesised steps.
  2. No recommended setpoints or operating parameters. Advisory questions get evidence, ranges, outcomes, documented limits, then an explicit deferral. Never a number presented as the answer.
  3. No answers outside the evidence. Zero rows means "no records found", never a plausible figure.

These are code paths, not prompt instructions. A safety rule living only in a prompt is not implemented. The classifier runs first; each class has a Pydantic contract validated in Python after generation and before returning. Contract failure -> regenerate once -> error. Never return.

When class is uncertain, choose the more restrictive class. Procedural beats Reference. Advisory beats Historical. Partly-advisory is advisory.

Host rules — inherited, non-negotiable

  • Growing data goes on /datadisk, never /. Root is 62 GB and has hit 100% before.
  • No published host ports for anything we build. Join the external proxy network, reach it through Caddy. Some existing containers do publish (caddy, wireguard, mosquitto, chirpstack-gateway-bridge, openplc-runtime) — they carry non-HTTP protocols. We do not.
  • Never omit import authelia from a Caddyfile block — it silently makes a service public. One deliberate exception, applied 2026-08-28 at the customer's direction: ai.yokogawa.tech admits remote_ip 10.0.0.21 (cicore1) unauthenticated and 403s everything else. It does not generalise — do not copy it, and do not widen it. api.yokogawa.tech stays gated.
  • ~/authelia/configuration.yml is root-owned: sudo, back up as .bak-<purpose>-<date>. Restarting Authelia logs out every active user — announce it first.
  • AD group membership must be DIRECT. Nested membership silently fails.
  • pg-ai and cube are pinned — do not add them to Watchtower's update list.
  • openplc-runtime is live control for this demo. Never restart, update or reconfigure it as a side effect of AI work. Do not copy its published-port pattern, and never change its 10.0.0.17 binding to 0.0.0.0 — that binding is what keeps unauthenticated Modbus off the internet.
  • Never install on, write to or restart cicore1. Never exceed read-only on imh.
  • No secrets in Git, compose files, logs or error messages. 0600 env files under ~/ai/.

Verification

docker ps showing "Up" is not proof. curl -sI the public URL and read the container logs. Expect 302 to the auth portal from api, cube and lf — but 403 from ai, which is Authelia-free and IP-matched (see the exception above). A 302 on ai means the SCADA-only block was reverted; a 200 means the matcher is restricting nothing. verify.sh treats the 403 as the PASS. Prove each layer without the LLM first: prove Cube returns the right number by hand, prove retrieval finds the right procedure by hand, then wire the agent.

Working conventions

  • Prefer additive changes. Snapshot config before editing.
  • Do not invent schema. imh is pending — inspect it, or ask. Fixtures are marked as fixtures. Before connecting it, work through db/README-standin-historian.md — the stand-in's shape is reasoned from the SCADA config, but every SQL Server detail in it is a guess, and two of the twelve listed assumptions fail silently rather than erroring.
  • Store UTC. Convert to SITE_TIMEZONE exactly once, in Cube. Never do timezone maths in a prompt.
  • Fix failures in the classifier, Cube or ingestion — not by adding instructions to the prompt.
  • A document becomes citable only after a human confirms its number, revision and effective dateconfirm_header() at a terminal, or the Phase 9 review screen. Never add a path that ingests an unconfirmed header. The web-facing role may only make a document less citable (withdraw). The design puts publishing and restoring behind ai-docs-worker, which has no HTTP surface. That is not what is built. ai-docs-worker was never deployed: api/documents.py serves POST /documents/review/{id}/approve and POST /documents/restore directly on ai-api. The boundary still exists, but it moved — it is enforced by db/005's trigger and by which database role each function opens its connection with, not by deployment. Read the docstring at the top of api/documents.py before touching that path, and do not assume an HTTP request cannot publish.
  • When something fails, add the failing case to eval/testset.jsonl before fixing it.
  • A defect you are not fixing now goes in status/OPEN-ISSUES.md, not in a comment. Read its header first — it says what belongs there and what belongs in REQUESTS.md or in §14 of the build spec, which is closed by decision and must not be re-raised.
  • Small commits, one concern each. If a change alters an accepted phase, re-run that phase's gate.

The plant

Waterloo Road Pump Station: a three-pump wastewater station. Wet well WW-101 (07000 mm, 120 m³/m), duty/assist/assist pumps PU-301/302/303 on a common VSD speed reference, discharging through manifold MAN-301 against 22 m static lift. Spill weir at 6000 mm, LSHH-102 at 5500 mm. Control runs on openplc-runtime; Yokogawa CI Server on cicore1 polls it over Modbus TCP and historises the result.

Four namespaces name the same measurement. Know which one you are holding.

Layer Example Source of truth
Instrument tag LIT-101 WRPS/01-design-doc
PLC symbol + address %QW0 WRPS/04-plc/register-map.csv
CI Server point STN_LEVEL, on Modbus station WRPS_PLC WRPS/05-scada/modbus/wrps_modbus_point_df.qli
CI Server item AID.WRPS.STN.LEVEL WRPS/05-scada/modbus/wrps_item_df.qli

The point layer is written WRPS_PLC:STN_LEVEL in the .qli export: the Modbus station name, a colon, then the point name. There are 49 points and 49 items, one to one here, but that is a property of this small station and not a rule.

PS_* names are not a namespace. An earlier delivery of the point list carried PS_STN_WET_WELL_LEVEL and its like where the item name belongs. They were a proposal derived from the PLC register map that CI Server never adopted, and they were removed from this repository on 2026-09-01: the tag seed, the alarm bits, the fixtures, the Cube models and the exam are all keyed on the item now. Do not reintroduce them. The only legitimate PS_ names are the four Modbus poll groups — PS_STATUS_BITS, PS_PUBLISHED, PS_SETPOINTS, PS_SIM_CONTROL — which are live configuration and appear only in a poll_group column. The delivered CI Server configuration lives in db/seed/scada-source/; read its README before touching any of this.

The historian is keyed on the item, and only the item. Modbus carries register numbers, not names, so the point layer and the item layer are free to differ — and they do. Keying history on a point name is what produced all three Phase 5 findings. db/seed/tags.csv owns plant facts; db/seed/historian_items.csv owns the item-to-tag mapping and is generated by scripts/gen_historian_items.py — never hand-edited.

  • Equipment is asserted in exactly one place: tags.equipment_id. Nothing in the history carries an equipment column. CI Server's section tree stops at the station and the three pumps; it has no wet well, weir, manifold or switchboard.
  • The historian retains seven days. Zero rows outside that window means "the historian does not go back that far", never "nothing happened". metrics.HISTORY_RETENTION_DAYS.
  • Never hardcode a sample interval. Read scan_interval_seconds from the item. The analogue groups run at 5 s and 30 s; a measure that assumed 60 s was wrong by twelvefold.