No description
Find a file
Claude 4826255672 Correct the spec against the host as actually observed
Verified on lin001 over SSH on 2026-08-20, read-only:

- Port 502 is bound to 10.0.0.17, not 0.0.0.0, so unauthenticated Modbus is
  not internet-reachable at the Docker level. This was an open risk in §3 and
  an unchecked item in §15; it is now a confirmation. openplc-runtime also
  publishes 8443 (the Runtime web UI) on the same private address, which the
  spec did not mention.
- openplc-runtime is not the only published port on the host: caddy, wireguard,
  mosquitto and chirpstack-gateway-bridge all publish on 0.0.0.0. The
  no-published-ports rule still applies in full to what we build, but the
  "one deliberate exception" framing was wrong and invited over-reading.
- Port 22 is open to the internet. Added to §15 as an open item for the same
  NSG review.
- 21 containers running, not 22; no stopped containers.
- /datadisk is 46% used with InfluxDB at 55 GB, up from 43%/52 GB. Recorded
  the growth rate so it can be budgeted for.

CLAUDE.md restates two of these rules and is updated to match.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 14:29:41 +10:00
api Scaffold the WRPS plant operations assistant repository 2026-08-20 13:56:32 +10:00
authelia Scaffold the WRPS plant operations assistant repository 2026-08-20 13:56:32 +10:00
caddy Scaffold the WRPS plant operations assistant repository 2026-08-20 13:56:32 +10:00
compose Scaffold the WRPS plant operations assistant repository 2026-08-20 13:56:32 +10:00
cube/model Scaffold the WRPS plant operations assistant repository 2026-08-20 13:56:32 +10:00
db Scaffold the WRPS plant operations assistant repository 2026-08-20 13:56:32 +10:00
docs Scaffold the WRPS plant operations assistant repository 2026-08-20 13:56:32 +10:00
eval Scaffold the WRPS plant operations assistant repository 2026-08-20 13:56:32 +10:00
ingest Scaffold the WRPS plant operations assistant repository 2026-08-20 13:56:32 +10:00
scripts Scaffold the WRPS plant operations assistant repository 2026-08-20 13:56:32 +10:00
web Scaffold the WRPS plant operations assistant repository 2026-08-20 13:56:32 +10:00
.env.example Scaffold the WRPS plant operations assistant repository 2026-08-20 13:56:32 +10:00
.gitattributes Normalise line endings to LF 2026-08-20 13:56:43 +10:00
.gitignore Scaffold the WRPS plant operations assistant repository 2026-08-20 13:56:32 +10:00
BUILD-AI-CONTAINERS.md Correct the spec against the host as actually observed 2026-08-20 14:29:41 +10:00
CLAUDE.md Correct the spec against the host as actually observed 2026-08-20 14:29:41 +10:00
README.md Scaffold the WRPS plant operations assistant repository 2026-08-20 13:56:32 +10:00
YAU_Linux_Host_Onboarding.md Scaffold the WRPS plant operations assistant repository 2026-08-20 13:56:32 +10:00

WRPS Plant Operations Assistant

A proof-of-concept assistant that lets an operator at the Waterloo Road Pump Station ask a question in plain English and get an answer grounded in plant data and controlled documents.

Example question Class
"How many times did the wet well high level alarm come up last week?" Historical
"What does the level signal fault alarm on the wet well mean?" Reference
"How do I lift the interlock on Pump 02?" Procedural
"What discharge rate should we run to avoid spilling?" Advisory

Those four need different retrieval paths, different answer contracts and different safety rules. One generic pipeline covering all four is the main way this project fails.

Success is a correct, citable, appropriately-scoped answer. Fluency is not success.


Read this before writing any code

This is an information retrieval and analysis assistant. It is not a control system, not an advisory controller, and not a substitute for a competent person.

It does not issue instructions for safety-critical actions. For "how do I lift the interlock on Pump 02" it locates and cites the controlled procedure. An interlock exists because somebody assessed a hazard; a bypass procedure reassembled from retrieved fragments is a safety document nobody approved.

It does not recommend setpoints or operating parameters. For "what discharge rate" it gives evidence — rates used, outcomes, when alarms occurred, documented capacity — and then defers. A number presented as an answer gets typed into a control system by someone who trusts it.

It does not answer outside its evidence. Zero rows means "no records found", never an invented figure.

These are code paths, not prompt instructions: api/contracts.py holds one Pydantic contract per class, validated after generation and before returning. A response that fails its contract is regenerated once, then errors. It is never returned. pytest api/tests exercises every rule above without an API key or a database, because that is the point of putting them in Python.

Full detail: BUILD-AI-CONTAINERS.md §2.


The plant

Waterloo Road Pump Station is a three-pump wastewater station.

  • Wet well WW-101, 07000 mm, 120 m³ per metre of level
  • Pumps PU-301/302/303, duty/assist/assist, ~120 L/s each against 22 m static lift, on a common VSD speed reference clamped 3850 Hz
  • Spill weir crest at 6000 mm, LSHH-102 at 5500 mm, high level alarm at 5200 mm, stop-all at 1000 mm
  • Control logic runs on openplc-runtime; Yokogawa CI Server on cicore1 polls it over Modbus TCP and historises the result

The unit trap that will catch you: the PLC works in millimetres and litres per second; the historian stores percent of the weir crest (raw mm ÷ 60) and m³/h. Every conversion is recorded per tag in db/seed/tags.csv, which also records — in capitals, at the start of each description — whether a tag is historised at all. Field inputs to the PLC (%IW/%IX: vibration, thermal, discharge pressure) are not published to SCADA and have no history. An answer that trends PU-301 vibration is fabricating data.

Source of truth for the plant: WRPS/01-design-doc/, WRPS/04-plc/register-map.csv and WRPS/05-scada/modbus/scada-points.csv in the WRPS repository.


Architecture

operator ──► Caddy ──► Authelia (AD + Duo) ──► ai-web ──► ai-api
                                                            │
                            ┌───────────────────────────────┼──────────────┐
                            ▼                               ▼              ▼
                      classifier                          Cube        pgvector
                   (CHEAP_DEPLOYMENT)                       │          (pg-ai)
                            │                               ▼
                    one branch per class              imh (SQL Server,
                            │                          read-only, TDS/1433)
                            ▼                          ── PENDING ──
                     contract validation
                            │
                            ▼
                        Langfuse

Everything runs on yau-sls-poc-lin001 (10.0.0.17), a shared, live Docker host that already runs 22 containers including openplc-runtime — the PLC for this demo. See YAU_Linux_Host_Onboarding.md.

There is no replication job and no mirror table. imh is already an isolated copy of the raw SCADA historian, so Cube queries it directly with a read-only login. pg-ai holds pgvector chunks, Cube pre-aggregations, and the equipment/tag reference data.

Container Stack Networks Public URL
pg-ai pgvector/pgvector:pg16 ai-internal only none
cube cubejs/cube (pinned) ai-internal + proxy cube.yokogawa.tech
ai-api Python 3.12 + FastAPI ai-internal + proxy api.yokogawa.tech
ai-web Vite build → nginx:alpine proxy ai.yokogawa.tech
ai-ingest Python 3.12, on demand ai-internal none
langfuse + lf-db official images ai-internal + proxy lf.yokogawa.tech

Rebuild from zero

Assumes: a checkout at ~/ai on lin001, and the Caddy + Authelia + proxy stack already running (it is — this host has served demos for months).

1. Secrets

Three 0600 env files under ~/ai/, never in Git. Every key is listed with no values in .env.example.

mkdir -p ~/ai && cd ~/ai
install -m 600 /dev/null pg-ai.env
install -m 600 /dev/null api.env
install -m 600 /dev/null langfuse.env

Follow the ~/authelia/authelia.env precedent. The Grafana admin password sitting in plain text in ~/docker-compose.yml is a known defect on this host, not a pattern to copy.

2. Phase 1 — pg-ai

./scripts/deploy.sh phase1

Creates /datadisk/pg-ai, starts pg-ai, applies the schema and roles, loads equipment.csv and tags.csv with their alias arrays, and — while USE_FIXTURES=true — loads the fixture stand-in for imh.

Gate: pg-ai healthy, vector present, agent_ro can SELECT and cannot INSERT, every equipment item and tag has an alias, pg-ai publishes no host port and is not on the proxy network, and df -h / is unchanged. ./scripts/verify.sh checks all of it.

3. Phase 2 — Langfuse

./scripts/deploy.sh phase2

Deployed early on purpose: from here on, every experiment is traced. Then do the manual steps the script prints — DNS, Caddyfile, Authelia rule, announce the Authelia restart.

4. Phase 3 — knowledge base

Put the controlled documents on the host, in the folders that determine doc_type:

/datadisk/ai-docs/{procedures,manuals,rationalisation,design}/
docker compose -f ~/ai-compose.yml run --rm ai-ingest --all

It will ask you to confirm the document number, revision and effective date for every file. Confirm them properly. A wrong revision on a procedure is a safety issue, not a data-quality one. When a new revision lands:

docker compose -f ~/ai-compose.yml run --rm ai-ingest --supersede WRPS-OPS-014 4

5. Phase 4 — imh ⚠ PENDING

The only true blocker. Start the conversation now; do not wait for Phase 3. Agree the read-only login, the table names and key columns, the timestamp semantics, and an NSG rule allowing lin001imh on 1433 only. Then update §10 of BUILD-AI-CONTAINERS.md with the real schema and change db/002_fixtures.sql and the Cube models to match.

Until then everything runs on fixtures, and every answer carries a fixture banner all the way to the operator's screen.

6. Phases 57 — Cube, API, UI

./scripts/deploy.sh api      # cube + ai-api
./scripts/deploy.sh web      # ai-web
./scripts/verify.sh

Each prints the manual DNS/Caddy/Authelia steps. Phase 7 also needs a pinpoint DNS record on the DC10.0.0.17 so an operator on cicore1 can resolve ai.yokogawa.tech — Azure hairpin means LAN hosts cannot reach the VM's public IP from inside the VNet. influx.yokogawa.tech already has this treatment. Raise it early; it depends on someone else and will not surface until you try it.

7. Phase 8 — validate

python eval/run_eval.py --api https://api.yokogawa.tech

62 engineer-reviewable cases in eval/testset.jsonl, every data-dependent one with a pinned time windowimh is live, and an unpinned question gives a different answer each run.

Gate: ≥85% overall, ≥95% classification accuracy on Procedural and Advisory, zero contract violations, p95 under 12 s. run_eval.py returns non-zero if any of those is missed. It also marks Historical and Advisory cases needs_review: whether "6" is the right number is a judgement for an engineer with access to imh, not something this script can decide.


Working on it

pytest api/tests          # contracts, classifier rules, SQL allow-list. No network.
  • The host is live and shared. Prefer additive changes. Snapshot config before editing. Announce anything that restarts Caddy or Authelia — it logs out every active user, including whoever is mid-demo.
  • Never restart, update or reconfigure openplc-runtime as a side effect of AI work. It is the PLC for the demo plant. Its published port 502 is the one deliberate exception to the no-published-ports rule on this host, and it does not generalise to anything we build.
  • Verify, don't assume. docker ps showing "Up" is not proof.
  • Test every layer without the LLM first. Prove Cube returns the right number by hand. Prove retrieval finds the right procedure by hand. Then wire up the agent — otherwise a wrong answer has four possible causes.
  • Do not invent schema. Inspect, or ask.
  • Fix eval failures in the classifier, Cube and ingestion — not by adding instructions to the prompt. When something fails, add the failing case to eval/testset.jsonl before fixing it.

Repository layout

CLAUDE.md                    short rules — what Claude Code keeps front of mind
BUILD-AI-CONTAINERS.md       the build spec
YAU_Linux_Host_Onboarding.md the host brief (reference; wins on conflict)
compose/                     deployed to ~/ai-compose.yml and ~/langfuse-compose.yml
caddy/ai-routes.caddy        blocks to paste into ~/Caddyfile
authelia/access-rules.md     the rule additions as text — never the real config
db/                          schema, roles, fixtures, and the alias seed CSVs
cube/model/                  alarms, process values, operations, equipment
api/                         FastAPI, classifier, agent, contracts, guardrails
ingest/                      Docling → chunk → embed → pg-ai
web/                         React + Vite operator UI
eval/                        62-case test set and the scorecard runner
scripts/                     deploy.sh, verify.sh
docs/                        GITIGNORED — real content on /datadisk/ai-docs

Known shortcuts

Deliberate, documented, and not to be shipped. Full list in BUILD-AI-CONTAINERS.md §14. The ones that matter most:

  • Secrets in 0600 env files, not a vault
  • No OT/IT firewall boundary — one flat 10.0.0.0/24 PoC network
  • Modbus TCP on port 502 with no authentication or encryption, contained by NSG/VPN scope only — confirm the NSG does not expose it to the internet
  • Single host, no HA: lin001 is a single point of failure for both the demo estate and the simulated plant's PLC
  • No automated backup — pg-ai needs adding to whatever backup exists
  • Document revision metadata entered semi-manually, not integrated with document control

Section 2 of the build spec must be reviewed with an OT/safety representative before any operator sees a demo.