yau-plant-assistant/README.md
Claude 34d2ccc576 Scaffold the WRPS plant operations assistant repository
Build spec and host brief carried in from C:\Claude and WRPS/02-env; the
plant model (equipment, tags, alarm bitmask, enums, unit conversions) is
derived from WRPS/04-plc/register-map.csv, WRPS/05-scada/modbus/scada-points.csv
and WRPS-CTL-003.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 13:56:32 +10:00

289 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# WRPS Plant Operations Assistant
A proof-of-concept assistant that lets an operator at the **Waterloo Road Pump
Station** ask a question in plain English and get an answer grounded in plant
data and controlled documents.
| Example question | Class |
|---|---|
| "How many times did the wet well high level alarm come up last week?" | **Historical** |
| "What does the level signal fault alarm on the wet well mean?" | **Reference** |
| "How do I lift the interlock on Pump 02?" | **Procedural** |
| "What discharge rate should we run to avoid spilling?" | **Advisory** |
Those four need different retrieval paths, different answer contracts and
different safety rules. One generic pipeline covering all four is the main way
this project fails.
**Success is a correct, citable, appropriately-scoped answer. Fluency is not
success.**
---
## Read this before writing any code
This is an **information retrieval and analysis assistant**. It is not a
control system, not an advisory controller, and not a substitute for a
competent person.
**It does not issue instructions for safety-critical actions.** For *"how do I
lift the interlock on Pump 02"* it locates and cites the controlled procedure.
An interlock exists because somebody assessed a hazard; a bypass procedure
reassembled from retrieved fragments is a safety document nobody approved.
**It does not recommend setpoints or operating parameters.** For *"what
discharge rate"* it gives evidence — rates used, outcomes, when alarms
occurred, documented capacity — and then defers. A number presented as an
answer gets typed into a control system by someone who trusts it.
**It does not answer outside its evidence.** Zero rows means "no records
found", never an invented figure.
These are **code paths, not prompt instructions**: [`api/contracts.py`](api/contracts.py)
holds one Pydantic contract per class, validated after generation and before
returning. A response that fails its contract is regenerated once, then errors.
It is never returned. `pytest api/tests` exercises every rule above without an
API key or a database, because that is the point of putting them in Python.
Full detail: [`BUILD-AI-CONTAINERS.md`](BUILD-AI-CONTAINERS.md) §2.
---
## The plant
Waterloo Road Pump Station is a three-pump wastewater station.
- Wet well `WW-101`, 07000 mm, **120 m³ per metre** of level
- Pumps `PU-301/302/303`, duty/assist/assist, ~120 L/s each against 22 m static
lift, on a **common VSD speed reference** clamped 3850 Hz
- Spill weir crest at **6000 mm**, `LSHH-102` at 5500 mm, high level alarm at
5200 mm, stop-all at 1000 mm
- Control logic runs on `openplc-runtime`; Yokogawa CI Server on `cicore1`
polls it over Modbus TCP and historises the result
**The unit trap that will catch you:** the PLC works in millimetres and litres
per second; the historian stores **percent of the weir crest** (raw mm ÷ 60) and
**m³/h**. Every conversion is recorded per tag in
[`db/seed/tags.csv`](db/seed/tags.csv), which also records — in capitals, at the
start of each description — whether a tag is **historised at all**. Field inputs
to the PLC (`%IW`/`%IX`: vibration, thermal, discharge pressure) are not
published to SCADA and have no history. An answer that trends PU-301 vibration
is fabricating data.
Source of truth for the plant: `WRPS/01-design-doc/`, `WRPS/04-plc/register-map.csv`
and `WRPS/05-scada/modbus/scada-points.csv` in the WRPS repository.
---
## Architecture
```
operator ──► Caddy ──► Authelia (AD + Duo) ──► ai-web ──► ai-api
┌───────────────────────────────┼──────────────┐
▼ ▼ ▼
classifier Cube pgvector
(CHEAP_DEPLOYMENT) │ (pg-ai)
│ ▼
one branch per class imh (SQL Server,
│ read-only, TDS/1433)
▼ ── PENDING ──
contract validation
Langfuse
```
Everything runs on `yau-sls-poc-lin001` (`10.0.0.17`), a **shared, live** Docker
host that already runs 22 containers including `openplc-runtime` — the PLC for
this demo. See [`YAU_Linux_Host_Onboarding.md`](YAU_Linux_Host_Onboarding.md).
**There is no replication job and no mirror table.** `imh` is already an
isolated copy of the raw SCADA historian, so Cube queries it directly with a
read-only login. `pg-ai` holds pgvector chunks, Cube pre-aggregations, and the
equipment/tag reference data.
| Container | Stack | Networks | Public URL |
|---|---|---|---|
| `pg-ai` | `pgvector/pgvector:pg16` | `ai-internal` only | none |
| `cube` | `cubejs/cube` (pinned) | `ai-internal` + `proxy` | `cube.yokogawa.tech` |
| `ai-api` | Python 3.12 + FastAPI | `ai-internal` + `proxy` | `api.yokogawa.tech` |
| `ai-web` | Vite build → `nginx:alpine` | `proxy` | `ai.yokogawa.tech` |
| `ai-ingest` | Python 3.12, on demand | `ai-internal` | none |
| `langfuse` + `lf-db` | official images | `ai-internal` + `proxy` | `lf.yokogawa.tech` |
---
## Rebuild from zero
Assumes: a checkout at `~/ai` on `lin001`, and the Caddy + Authelia + `proxy`
stack already running (it is — this host has served demos for months).
### 1. Secrets
Three `0600` env files under `~/ai/`, never in Git. Every key is listed with no
values in [`.env.example`](.env.example).
```bash
mkdir -p ~/ai && cd ~/ai
install -m 600 /dev/null pg-ai.env
install -m 600 /dev/null api.env
install -m 600 /dev/null langfuse.env
```
Follow the `~/authelia/authelia.env` precedent. The Grafana admin password
sitting in plain text in `~/docker-compose.yml` is a known defect on this host,
not a pattern to copy.
### 2. Phase 1 — `pg-ai`
```bash
./scripts/deploy.sh phase1
```
Creates `/datadisk/pg-ai`, starts `pg-ai`, applies the schema and roles, loads
`equipment.csv` and `tags.csv` with their alias arrays, and — while
`USE_FIXTURES=true` — loads the fixture stand-in for `imh`.
**Gate:** `pg-ai` healthy, `vector` present, `agent_ro` can SELECT and cannot
INSERT, every equipment item and tag has an alias, `pg-ai` publishes no host
port and is not on the `proxy` network, and `df -h /` is unchanged.
`./scripts/verify.sh` checks all of it.
### 3. Phase 2 — Langfuse
```bash
./scripts/deploy.sh phase2
```
Deployed early on purpose: from here on, every experiment is traced. Then do
the manual steps the script prints — DNS, Caddyfile, Authelia rule, announce
the Authelia restart.
### 4. Phase 3 — knowledge base
Put the controlled documents on the host, in the folders that determine
`doc_type`:
```
/datadisk/ai-docs/{procedures,manuals,rationalisation,design}/
```
```bash
docker compose -f ~/ai-compose.yml run --rm ai-ingest --all
```
It will ask you to confirm the document number, revision and effective date for
every file. **Confirm them properly.** A wrong revision on a procedure is a
safety issue, not a data-quality one. When a new revision lands:
```bash
docker compose -f ~/ai-compose.yml run --rm ai-ingest --supersede WRPS-OPS-014 4
```
### 5. Phase 4 — `imh` ⚠ PENDING
**The only true blocker.** Start the conversation now; do not wait for Phase 3.
Agree the read-only login, the table names and key columns, the timestamp
semantics, and an NSG rule allowing `lin001``imh` on 1433 only. Then update
§10 of `BUILD-AI-CONTAINERS.md` with the real schema and change
[`db/002_fixtures.sql`](db/002_fixtures.sql) and the Cube models to match.
Until then everything runs on fixtures, and every answer carries a fixture
banner all the way to the operator's screen.
### 6. Phases 57 — Cube, API, UI
```bash
./scripts/deploy.sh api # cube + ai-api
./scripts/deploy.sh web # ai-web
./scripts/verify.sh
```
Each prints the manual DNS/Caddy/Authelia steps. Phase 7 also needs a
**pinpoint DNS record on the DC**`10.0.0.17` so an operator on `cicore1` can
resolve `ai.yokogawa.tech` — Azure hairpin means LAN hosts cannot reach the
VM's public IP from inside the VNet. `influx.yokogawa.tech` already has this
treatment. **Raise it early**; it depends on someone else and will not surface
until you try it.
### 7. Phase 8 — validate
```bash
python eval/run_eval.py --api https://api.yokogawa.tech
```
62 engineer-reviewable cases in [`eval/testset.jsonl`](eval/testset.jsonl), every
data-dependent one with a **pinned time window**`imh` is live, and an
unpinned question gives a different answer each run.
Gate: ≥85% overall, ≥95% classification accuracy on Procedural and Advisory,
**zero** contract violations, p95 under 12 s. `run_eval.py` returns non-zero if
any of those is missed. It also marks Historical and Advisory cases
`needs_review`: whether "6" is the *right* number is a judgement for an
engineer with access to `imh`, not something this script can decide.
---
## Working on it
```bash
pytest api/tests # contracts, classifier rules, SQL allow-list. No network.
```
- **The host is live and shared.** Prefer additive changes. Snapshot config
before editing. **Announce anything that restarts Caddy or Authelia** — it
logs out every active user, including whoever is mid-demo.
- **Never restart, update or reconfigure `openplc-runtime`** as a side effect
of AI work. It is the PLC for the demo plant. Its published port 502 is the
one deliberate exception to the no-published-ports rule on this host, and it
does not generalise to anything we build.
- **Verify, don't assume.** `docker ps` showing "Up" is not proof.
- **Test every layer without the LLM first.** Prove Cube returns the right
number by hand. Prove retrieval finds the right procedure by hand. Then wire
up the agent — otherwise a wrong answer has four possible causes.
- **Do not invent schema.** Inspect, or ask.
- Fix eval failures in the classifier, Cube and ingestion — **not by adding
instructions to the prompt**. When something fails, add the failing case to
`eval/testset.jsonl` *before* fixing it.
---
## Repository layout
```
CLAUDE.md short rules — what Claude Code keeps front of mind
BUILD-AI-CONTAINERS.md the build spec
YAU_Linux_Host_Onboarding.md the host brief (reference; wins on conflict)
compose/ deployed to ~/ai-compose.yml and ~/langfuse-compose.yml
caddy/ai-routes.caddy blocks to paste into ~/Caddyfile
authelia/access-rules.md the rule additions as text — never the real config
db/ schema, roles, fixtures, and the alias seed CSVs
cube/model/ alarms, process values, operations, equipment
api/ FastAPI, classifier, agent, contracts, guardrails
ingest/ Docling → chunk → embed → pg-ai
web/ React + Vite operator UI
eval/ 62-case test set and the scorecard runner
scripts/ deploy.sh, verify.sh
docs/ GITIGNORED — real content on /datadisk/ai-docs
```
---
## Known shortcuts
Deliberate, documented, and not to be shipped. Full list in
[`BUILD-AI-CONTAINERS.md`](BUILD-AI-CONTAINERS.md) §14. The ones that matter most:
- Secrets in `0600` env files, not a vault
- No OT/IT firewall boundary — one flat `10.0.0.0/24` PoC network
- Modbus TCP on port 502 with no authentication or encryption, contained by
NSG/VPN scope only — **confirm the NSG does not expose it to the internet**
- Single host, no HA: `lin001` is a single point of failure for both the demo
estate and the simulated plant's PLC
- No automated backup — `pg-ai` needs adding to whatever backup exists
- Document revision metadata entered semi-manually, not integrated with
document control
**Section 2 of the build spec must be reviewed with an OT/safety representative
before any operator sees a demo.**