README.md was 481 lines, 60% of which was a nine-phase deployment runbook. A new team member had to scroll past every deploy command to reach what the folders are. The runbook moves out whole to spec/REBUILD.md - content unchanged apart from heading levels and seven link targets repointed for the new depth. README.md is now 258 lines. Two additions: "Start here" gives a reading order - README, then the workflow map, then the current state, then CLAUDE.md - and a table routing to the runbook, the build spec, the host brief or REQUESTS.md by what the reader came to do. It also names the three accesses needed before starting, none of which are in this repository. The architecture diagram is redrawn because the old one was wrong. It showed Authelia in front of ai-web. Since 28 August that hostname carries no Authelia at all - it is matched on source IP. The README said so correctly, 170 lines below the picture. The new diagram shows both parties arriving at the same Caddy and then branching: ai.yokogawa.tech to the @scada remote_ip matcher, the other three through import authelia. The operator is inside the Caddy/Authelia environment; they take a branch Authelia does not sit on. The answer path behind ai-api is drawn separately, because it is the same whichever door the question came in through. Structure follows caddy/ai-routes.caddy:190-203. The layout block gains REBUILD.md and the eval count is corrected to 78 in passing, that line being part of the block rewritten here. Every markdown link in all nine documentation files was checked after the move: none broken. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
14 KiB
Rebuild from zero
The full deployment runbook, moved out of
README.mdso that file can stay an orientation document. Nine phases, in order, each with a gate. Design detail for every phase is inBUILD-AI-CONTAINERS.md; the host it deploys onto is described inYAU_Linux_Host_Onboarding.md, which wins wherever the two conflict.
Paths in this file are written from the repository root, which is ~/ai on lin001.
Assumes: a checkout at ~/ai on lin001, and the Caddy + Authelia + proxy
stack already running (it is — this host has served demos for months).
1. Secrets
Three 0600 env files under ~/ai/, never in Git. Every key is listed with no
values in .env.example.
mkdir -p ~/ai && cd ~/ai
install -m 600 /dev/null pg-ai.env
install -m 600 /dev/null api.env
install -m 600 /dev/null langfuse.env
Follow the ~/authelia/authelia.env precedent. The Grafana admin password
sitting in plain text in ~/docker-compose.yml is a known defect on this host,
not a pattern to copy.
2. Phase 1 — pg-ai
./scripts/deploy.sh phase1
Creates /datadisk/pg-ai, starts pg-ai, applies the schema and roles, loads
equipment.csv and tags.csv with their alias arrays, and — while
USE_FIXTURES=true — loads the fixture stand-in for imh.
Three roles come out of this, and the split matters: agent_ro for the answer
path (SELECT only, everywhere), cube_rw for pre-aggregations, and ingest_rw
— the only role that writes doc_chunks. Set all three passwords in the 0600
env files; INGEST_DB_USER / INGEST_DB_PASSWORD are what Phase 3 connects
with, and Phase 3 refuses to start without them.
Gate: pg-ai healthy, vector present, agent_ro can SELECT and cannot
INSERT, ingest_rw can write doc_chunks and nothing else, every equipment item and tag has an alias, pg-ai publishes no host
port and is not on the proxy network, and df -h / is unchanged.
./scripts/verify.sh checks all of it.
3. Phase 2 — Langfuse
./scripts/deploy.sh phase2
Deployed early on purpose: from here on, every experiment is traced. Then do the manual steps the script prints — DNS, Caddyfile, Authelia rule, announce the Authelia restart.
4. Phase 3 — knowledge base
Put the controlled documents on the host, in the folders that determine
doc_type:
/datadisk/ai-docs/{procedures,manuals,rationalisation,design}/
docker compose -f ~/ai-compose.yml run --rm ai-ingest --all
It will ask you to confirm the document number, revision and effective date for every file. Confirm them properly. A wrong revision on a procedure is a safety issue, not a data-quality one. When a new revision lands:
docker compose -f ~/ai-compose.yml run --rm ai-ingest --supersede WRPS-OPS-014 4
That sticks. Re-ingesting a superseded document brings it back superseded, and
--all skips it — the flag is not reset by replacing chunks. To undo one:
docker compose -f ~/ai-compose.yml run --rm ai-ingest --restore WRPS-OPS-014 3
This is the SSH path, and it stays. Phase 9 adds the same thing as a screen, so that an operator who is issued a new document when the PLC logic changes does not have to find someone with a host login. It does not remove the header confirmation or the supersede decision — it puts them in front of a named person and records the answer. See step 8 below.
5. Phase 4 — imh ⚠ PENDING
The only true blocker. Start the conversation now; do not wait for Phase 3.
Agree the read-only login, the table names and key columns, the timestamp
semantics, and an NSG rule allowing lin001 → imh on 1433 only. Then update
§10 of BUILD-AI-CONTAINERS.md with the real schema and change
db/002_fixtures.sql and the Cube models to match.
Until then everything runs on fixtures, and every answer carries a fixture banner all the way to the operator's screen.
The three Phase 5 findings are fixed (2026-08-31), against the SCADA
configuration rather than against the fixtures. All three came from one
substitution: the stand-in historian was keyed on CI Server point names
(PS_STN_WET_WELL_LEVEL) when the historian is keyed on CI Server item
names (AID.WRPS.STN.LEVEL) — two layers apart, not one. Nothing is aliased
across that gap now: public.historian_items holds the mapping, generated from
WRPS/05-scada/modbus by scripts/gen_historian_items.py,
and an item that resolves to neither a tag nor a written reason is a build
error rather than a silent "no records found".
Two consequences worth knowing before you read a number off this system:
- The historian keeps seven days. Every WRPS history group is
LIFE_TIME "1 weeks", and the fixtures now match, so a question about last month fails here exactly as it would againstimh. Zero rows outside retention is reported as a retention limit, never as "nothing happened". - Equipment is asserted in exactly one place,
tags.equipment_id. The history carries no equipment column, because CI Server's section tree stops at the station and the three pumps and has no wet well to put there.
Full detail in BUILD-AI-CONTAINERS.md Phase 4,
"Resolved 2026-08-31". Eval cases H26, H27 and H31 cover them, and H29
covers the retention limit.
6. Phases 5–7 — Cube, API, UI
./scripts/deploy.sh api # cube + ai-api
./scripts/deploy.sh web # ai-web
./scripts/verify.sh
Each prints the manual DNS/Caddy/Authelia steps.
DNS is done as of 27 August 2026: ai, api and cube all resolve to
20.211.144.151, and the DC carries the pinpoint record ai.yokogawa.tech →
10.0.0.17 that the Azure hairpin requires. The Caddyfile blocks and Authelia
rules for the three names still have to be applied on the host — until they
are, the names resolve and nothing answers.
The operator console is unauthenticated, and only from the SCADA machine.
Applied 28 August 2026 at the customer's direction. The ai.yokogawa.tech Caddy
block admits remote_ip 10.0.0.21 (yau-poc-cicore1, static) and returns 403
to everything else — LAN, VPN and internet alike. An operator at the console
should not complete a Duo push to ask a question; nobody outside the plant
should reach the assistant at all.
Read this before relying on it. It is an IP allowlist on a flat network with
no OT/IT boundary: anything that can take 10.0.0.21 inherits unauthenticated
access. It is a demo affordance, not a security control, and it is listed as
such in BUILD-AI-CONTAINERS.md §14. It also makes
Langfuse traces anonymous — there is no longer a record of who asked what — and
it puts the assistant out of browser reach over the VPN, so engineers need an
SSH tunnel. api.yokogawa.tech is unchanged and still fully gated; Phase 9
document publishing depends on that and must stay there. Rollback is a Caddy
reload — the Authelia rule was left in place on purpose. See
caddy/ai-routes.caddy.
This is in scope for the §2 OT/safety review, which is still outstanding.
api.yokogawa.tech has no pinpoint record and does not resolve inside the
VNet. That is why the Phase 7 Caddy block routes /ask under
ai.yokogawa.tech to ai-api and the page is same-origin: a cross-origin
build loads on a control-room PC and then fails every question on DNS. Only
/ask is routed there — see the warning in
caddy/ai-routes.caddy before widening it.
7. Phase 8 — validate
python eval/run_eval.py --api https://api.yokogawa.tech
62 engineer-reviewable cases in eval/testset.jsonl, every
data-dependent one with a pinned time window — imh is live, and an
unpinned question gives a different answer each run.
Gate: ≥85% overall, ≥95% classification accuracy on Procedural and Advisory,
zero contract violations, p95 under 12 s. run_eval.py returns non-zero if
any of those is missed. It also marks Historical and Advisory cases
needs_review: whether "6" is the right number is a judgement for an
engineer with access to imh, not something this script can decide.
8. Phase 9 — operator document upload
BUILT AND LIVE 2026-08-28, ahead of Phase 8, at the customer's direction.
The screens are at https://api.yokogawa.tech/documents — served by ai-api,
not ai-web, because ai.yokogawa.tech is now SCADA-only and carries no
identity at all. Upload → convert → review → approve, plus withdraw and
restore. The pool screen was explicitly descoped.
Files are converted to text with pypdf / python-docx / openpyxl and the converted text is shown to the reviewer before approval — the raw file is never what the assistant reads, and a bad conversion is meant to be caught by eye. Scanned documents cannot be read and are refused rather than stored empty.
Two things to know before trusting it. Identity is self-asserted: the
publisher is a typed name checked against a one-entry list (admin) with no
password, so anyone who reaches the page can claim it. Rows are marked
demo:<name> / DEMO-UNVERIFIED so they stay distinguishable from
authenticated ones later. And the two ingest paths must not be used on the
same document — files published through the UI stay in the inbox and
ai-ingest --all cannot see them. Full list of divergences in
BUILD-AI-CONTAINERS.md §14.
The design below is what §16 specifies, and remains the target.
After Phase 8 passes, not before. When the PLC logic or the SCADA program changes, a new document is issued and the assistant is wrong about the plant until it is ingested. Today that needs SSH to a live shared host. Phase 9 puts it behind the UI:
upload ─► pre-scan ─► review ─► approve ─► published
(anyone (worker, (a named (header (chunks
with 2FA) Docling) publisher) confirmed) citable)
Nothing is citable until a named person in AI_DocPublishers has confirmed the
document number, revision and effective date, and decided what it supersedes —
the same questions ingest.py asks at a terminal, asked on a screen and, unlike
the terminal, recorded. The database refuses an approved row without them.
psql -h pg-ai -U postgres -d plant -f db/004_doc_uploads.sql
psql -h pg-ai -U postgres -d plant -f db/005_doc_actions.sql
psql -h pg-ai -U postgres -d plant -f db/006_doc_pool.sql
sudo install -d -o 10002 -g 10002 /datadisk/ai-docs-inbox # check df -h first
sudo install -d -o 10002 -g 10002 /datadisk/ai-docs-withdrawn
docker compose -f ~/ai-compose.yml up -d ai-docs-worker
Taking documents out is the other half, and today there is no way to do it:
--supersede needs a revision to keep, so a cancelled procedure or a manual
for equipment that has been removed cannot be withdrawn at all. Phase 9 adds
withdraw (immediate, reversible, keeps the chunks and the audit trail — this
is what "remove it" almost always means), restore, and purge (deletes,
irreversible, off unless ALLOW_PURGE=true and the publisher types the document
number). All three need the publisher group and a written reason, and all three
are recorded in doc_actions, which nothing can delete from.
Withdrawal also moves the file out of /datadisk/ai-docs into an archive, so
that folder keeps meaning "the documents this plant runs on". That is
housekeeping, not the safety mechanism: ingest.py reads a document's
withdrawal state before replacing its chunks and carries it through, and --all
skips withdrawn documents — so re-ingesting cannot resurrect one, whatever
folder it is in. --restore DOC_NUMBER REVISION is the way back, and it refuses
while another revision of the same document is live.
Then the manual steps: the copy_headers change on the api.yokogawa.tech
Caddy block, the ^/docs/.* Authelia rule above the general one, and
AI_DocPublishers in AD with direct membership.
Design and gate: BUILD-AI-CONTAINERS.md §16 and
Phase 9. Two gate items matter most. A user who is authenticated but not a
publisher must get a 403 from the API, tested by calling api.yokogawa.tech
directly — the button being hidden proves nothing. And after withdrawing a
document, ai-ingest --all must not bring it back; run it and check, because
that is the failure that puts a withdrawn procedure back in front of an
operator.
Choosing what is in the pool is a third, separate thing, and it is separate
on purpose. superseded says this document is withdrawn or replaced — a claim
about the document, with a reason and an audit row behind it. pool_enabled
says this document is part of the set we are running with — a claim about the
corpus, and no comment on whether the document is current. Retrieval requires
both, so putting a withdrawn document back in the pool does not make it
citable. A superuser curates the pool to keep it current; the same screen saves
named profiles.
For demos, POST /ask takes an optional pool_profile that narrows
retrieval for that one request and changes nothing stored — so "three
documents versus forty-seven" needs nothing undone afterwards on a host other
people are using. Every answer from a reduced pool carries a banner with the
document count, in the same place and for the same reason as the fixture-data
banner: an answer from a trimmed corpus is otherwise indistinguishable from a
complete one. The demo worth showing is not "more documents, better answers" —
it is that with the evidence removed the assistant says no governing procedure
in the active document set, rather than degrading into something plausible.
Read the HNSW note at the top of db/006_doc_pool.sql
before rehearsing that demo. The index covers every embedding and filters
afterwards, so a heavily trimmed pool can appear to collapse retrieval entirely.
What none of it does: update tags.csv, the Cube models or any
setpoint. A new design document changes what the assistant can cite; the
numbers behind Historical and Advisory answers still come from reference data
that is changed in Git and deployed. The review screen asks the reviewer to
confirm they know that, because a document going live while the tag metadata
behind it has not is a gap that is only visible at that moment.