yau-plant-assistant/cube/model/alarms.yml
Claude 8aba1f7f5c Rebuild the stand-in historian on CI Server item names
The three open Phase 5 findings were one defect: the stand-in was keyed on
CI Server POINT names (PS_STN_WET_WELL_LEVEL) when the historian is keyed on
CI Server ITEM names (AID.WRPS.STN.LEVEL). Modbus carries register numbers,
not names, so those two layers are free to differ - and do. Reconciling
against the register map, as planned, would only have proved the first three
namespaces agreed with each other.

Rebuilt from WRPS/05-scada/modbus, so item names, sample rates, retention and
timestamp semantics come from the machine rather than from a guess.

(a) Level tag does not join. PS_STN_WET_WELL_LEVEL becomes a tag row in its
    own right; LIT-101 is marked NOT HISTORISED - a field input on %IW0 that
    never reaches SCADA. It was the only seed row carrying two addresses.
    public.historian_items holds the item-to-tag mapping, generated by
    scripts/gen_historian_items.py and enforced non-empty at generate, at
    deploy and at verify.

(b) first_alarm/last_alarm returned UTC. Converted inside the measure, so it
    stays in Cube and happens once. Aggregate first, convert after - the other
    order picks the wrong row across a DST fall-back. Returned as a formatted
    string with a companion site_timezone measure. Storage being UTC is now
    confirmed, not assumed: all 49 points carry TIME_ZONE "Date+time GMT" and
    every history group CORRECT_DAYLIGHT=0. This answers Phase 4 task 4.

(c) High level alarm filed against the wrong equipment. Both sides were right
    about different things; the defect was asserting equipment twice. The
    history now carries no equipment column at all - faithful, since CI
    Server's section tree stops at the station and three pumps. Equipment is
    reached bit -> tag -> equipment via public.alarm_bits.

Alarms are derived, not stored: CI Server's ALARM_HISTORY group is empty
because every item imports with alarming off. Decomposing the alarm word needs
no configuration that does not exist.

Three things the SCADA config changed that were never filed as faults:
  - retention is 7 days, not 30. The advisory path was reporting a month of
    evidence drawn from a week of data
  - the analogue rate is 5 s, not 60. Two measures multiplied sample counts by
    a hardcoded 60 - a twelvefold overstatement that read as plausible
  - the deadband warning in process_values.yml was wrong and was steering
    people away from the correct measure

db/002_fixtures.sql now asserts its own counts at load and cross-checks the
alarm derivation against two independent signals. Those prove the pipeline,
not the plant.

db/README-standin-historian.md documents removal: the seam between generation
and contract, and twelve assumptions about imh that are NOT confirmed. Two of
them fail silently.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 11:42:02 +10:00

302 lines
13 KiB
YAML

# =============================================================================
# alarms.yml — alarm and event history.
#
# REMOVING THE STAND-IN: db/README-standin-historian.md. The view below is
# CONTRACT - repoint it at imh, do not edit this file to absorb a difference.
#
# SOURCE: fixture.alarm_history while USE_FIXTURES=true. That is a VIEW, not a
# table, and it is derived rather than stored — see below. When imh is live the
# view is repointed at imh and nothing in this file changes. That is the point
# of it.
#
# WHERE ALARMS COME FROM, AND WHY IT IS A DERIVATION
# --------------------------------------------------
# Every alarm at this station is a bit of the PLC alarm word, historised as the
# CI Server item AID.WRPS.STN.ALARM_WORD. The bit map is reference data in
# public.alarm_bits, and it is what carries an alarm to the equipment it
# belongs to: a bitmask packs several units' alarms into one item, so the item
# alone cannot say which pump a seal leak is about, but the bit can.
#
# CI Server's own ALARM_HISTORY group is configured on the server and is EMPTY:
# every WRPS item imports with alarming off and limits at 0, which
# 05-scada/modbus/README.md records as deliberate and still outstanding
# engineering judgement. Deriving from the alarm word needs no configuration
# that does not exist. If alarm limits are later configured on the items, this
# file does not change either — fixture.alarm_history does.
#
# THREE DEFINITIONS THAT DECIDE WHETHER THE ANSWERS ARE RIGHT. They are here in
# comments because the person checking the number needs to read them, and they
# are not obvious from the measure names.
#
# 1. AN ALARM IS A TRANSITION INTO THE ACTIVE STATE.
# state = 'ACTIVE' only — a 0 -> 1 on one bit. RTN is the return-to-normal
# of the activation that preceded it. Counting every row roughly doubles
# every answer. "6 times last week" must mean six activations.
#
# 2. "LAST WEEK" IS A ROLLING 7 x 24 h WINDOW IN SITE_TIMEZONE.
# Not the previous calendar week, not 7 calendar days. Storage is UTC and
# the conversion happens here, once. If someone means the calendar week they
# have to say so, and the answer must state the window it used.
#
# A ROLLING WEEK IS NOW ALSO THE WHOLE OF THE AVAILABLE HISTORY. Every WRPS
# history group carries LIFE_TIME "1 weeks", so a question about anything
# older returns no rows — correctly, and it must be reported as "outside
# retention", never as "no alarms occurred".
#
# 3. CHATTERING IS 3 OR MORE ACTIVATIONS OF THE SAME TAG WITHIN 60 MINUTES.
# An arbitrary threshold, chosen to match the site's alarm rationalisation
# convention. It is stated in the answer whenever chattering is reported,
# because a different threshold gives a different story.
#
# THERE IS NO equipment_id IN THE SOURCE, AND THERE MUST NOT BE ONE.
# Equipment is reached bit -> tag -> equipment, through public.tags, which is
# the single place equipment is asserted. A denormalised equipment column in
# the history disagreeing with the tag seed is what made the station's most
# obvious question unanswerable; there is now nothing left to disagree.
# =============================================================================
cubes:
- name: alarms
sql_table: fixture.alarm_history # -> imh-backed view at Phase 4
description: >
Alarm and event history for the Waterloo Road Pump Station, derived from
bit transitions of the PLC alarm word. One row per state transition.
Activations only are counted as alarms.
joins:
# tag first, then equipment through it. Both hops are defined once, here
# and in equipment.yml, so "how many wet well alarms" resolves without
# anything in this file naming a piece of equipment.
- name: tags
sql: "{CUBE}.tag_id = {tags}.tag_id"
relationship: many_to_one
dimensions:
- name: alarm_id
sql: alarm_id
type: number
primary_key: true
- name: event_time
sql: event_time
type: time
description: >
Transition time. Stored UTC — every WRPS Modbus point carries
TIME_ZONE "Date+time GMT" and every history group CORRECT_DAYLIGHT=0
— and presented in SITE_TIMEZONE. Cube converts time DIMENSIONS
automatically; see first_alarm below for why measures are different.
- name: item_name
sql: item_name
type: string
description: >
The CI Server item the alarm was derived from — always
AID.WRPS.STN.ALARM_WORD. Kept so an engineer can go from an answer
back to the raw history in one step.
- name: bit
sql: bit
type: number
description: Which bit of the alarm word. See public.alarm_bits.
- name: tag_id
sql: tag_id
type: string
description: >
The tag the alarm is ABOUT, from the bit map — MSE-333 for a PU-303
seal leak, not the bitmask item. This is the join to equipment.
- name: alarm_type
sql: alarm_type
type: string
description: >
HIGH_LEVEL, HIGH_HIGH_LEVEL, LOW_LOW_LEVEL, SPILL, PUMP_TRIP,
SEAL_LEAK, HIGH_VIBRATION, LEVEL_SIGNAL_FAULT, MAINS_FAILURE,
SETPOINT_REJECTED. One per bit of %QW17 — see db/seed/alarm_bits.csv.
- name: state
sql: state
type: string
description: ACTIVE or RTN. Only ACTIVE counts as an alarm.
- name: priority
sql: priority
type: number
description: >
1 highest, 3 lowest, from the bit map. Priority 1 is SPILL,
PUMP_TRIP, HIGH_HIGH_LEVEL, LOW_LOW_LEVEL, MAINS_FAILURE and
LEVEL_SIGNAL_FAULT — losing the level signal on a well that can spill
is a priority 1 condition. Reference data, not a constant in this
file, so this comment cannot drift out of step with the data again.
- name: value
sql: value
type: number
description: >
Process value at the transition, for the alarms that have one — a
level reading for a high level alarm, NULL for a pump trip. Do not
present NULL as zero.
- name: alarm_text
sql: alarm_text
type: string
- name: is_fixture
sql: is_fixture
type: boolean
description: >
TRUE means this row was derived from db/002_fixtures.sql and is
generated test data, not plant history. The API surfaces this to the
operator.
measures:
- name: alarm_count
type: count
filters:
- sql: "{CUBE}.state = 'ACTIVE'"
description: >
Number of alarm ACTIVATIONS. Definition 1 above. This is the measure
behind "how many times did X alarm come up".
- name: transition_count
type: count
description: >
Every row including RTN. Diagnostics only - do not answer an operator
question with this.
- name: distinct_tags
sql: tag_id
type: count_distinct
filters:
- sql: "{CUBE}.state = 'ACTIVE'"
description: How many different tags alarmed in the window.
# -----------------------------------------------------------------------
# first_alarm / last_alarm — Phase 5 finding (b), FIXED.
#
# These used to be plain min/max measures over a timestamp and came back
# in UTC. Cube converts time DIMENSIONS to the query timezone but not
# min/max MEASURES, so on the Sydney day bucket 2026-08-01 the measure
# returned 2026-07-31T20:00:35 — the right instant, ten hours and one
# calendar day out, beside a bucket label that WAS in site time.
#
# The conversion now happens here, inside the measure, which keeps it
# inside Cube and exactly once. The aggregate is taken FIRST and
# converted after — MIN(x) AT TIME ZONE z, not MIN(x AT TIME ZONE z) —
# because the second form takes the minimum of local clock readings and
# picks the wrong row across a daylight-saving fall-back, where one local
# hour occurs twice.
#
# Returned as a formatted string rather than a timestamp, deliberately:
# a bare timestamp with no offset is exactly what made the old defect
# invisible. The answer must state the timezone alongside it, and
# site_timezone below is what it states.
#
# THE ZONE NAME IS HARDCODED HERE, WHICH THE BUILD SPEC WOULD RATHER IT
# WERE NOT. Cube can read env vars in a YAML model through Jinja
# (`{{ env_var('SITE_TIMEZONE') }}`), which would take it from api.env
# like everything else. That was not shipped because a model that fails
# to compile takes every query down with it and Jinja support could not
# be tested against the pinned v1.1.7 before writing this. scripts/
# verify.sh asserts this literal matches SITE_TIMEZONE in api.env, so the
# two cannot drift silently. Switch it to env_var once someone can
# confirm the model still compiles on lin001.
# -----------------------------------------------------------------------
- name: first_alarm
sql: >
to_char(MIN({CUBE}.event_time) FILTER (WHERE {CUBE}.state = 'ACTIVE')
AT TIME ZONE 'Australia/Sydney', 'YYYY-MM-DD HH24:MI:SS')
type: string
description: >
Earliest activation in the window, in SITE_TIMEZONE. Quote it with
the timezone - see site_timezone.
- name: last_alarm
sql: >
to_char(MAX({CUBE}.event_time) FILTER (WHERE {CUBE}.state = 'ACTIVE')
AT TIME ZONE 'Australia/Sydney', 'YYYY-MM-DD HH24:MI:SS')
type: string
description: >
Latest activation in the window, in SITE_TIMEZONE. Quote it with the
timezone - see site_timezone.
- name: site_timezone
sql: "MAX('Australia/Sydney')"
type: string
description: >
The zone first_alarm and last_alarm are expressed in. A clock time
without its zone is what finding (b) shipped; this exists so the
answer never has to assume one.
- name: priority_1_count
type: count
filters:
- sql: "{CUBE}.state = 'ACTIVE' AND {CUBE}.priority = 1"
description: >
Priority 1 activations - trips, spills, high high level, dry run,
mains failure and level signal fault.
pre_aggregations:
# Keeps "count alarms last week" fast without repeatedly scanning imh.
# Materialised into Cube Store. first_alarm, last_alarm and
# site_timezone are deliberately NOT here: they are string measures over
# a non-additive aggregate and cannot be rolled up from an hourly
# partition. A query asking for them falls through to the source, which
# is correct and cheap at this volume.
- name: alarms_by_hour
measures: [alarm_count, distinct_tags, priority_1_count]
dimensions: [alarm_type, tag_id, bit]
time_dimension: event_time
granularity: hour
partition_granularity: month
refresh_key:
# 24h, not minutes: the fixtures are static, so a shorter
# interval rebuilds a byte-identical result on a 2-vCPU host
# and was a standing CPU load for no gain. Tune this back
# deliberately when imh makes the data genuinely live.
every: 24 hours
build_range_start:
# Seven days, matching the historian's own retention. Building 180
# days of partitions over a source that only ever holds seven is
# work that produces empty partitions.
sql: "SELECT now() - interval '8 days'"
build_range_end:
sql: "SELECT now()"
views:
- name: alarm_activity
description: >
Alarm activations joined to the tag they are about and the equipment that
tag belongs to, so a question about "Pump 02" or "the wet well" can be
answered without the caller knowing which bit of the alarm word carries
it.
cubes:
- join_path: alarms
includes:
- event_time
- alarm_type
- tag_id
- bit
- state
- priority
- value
- alarm_text
- is_fixture
- alarm_count
- distinct_tags
- priority_1_count
- first_alarm
- last_alarm
- site_timezone
- join_path: alarms.tags
prefix: true
includes:
- display_name
- signal_type
- engineering_unit
- join_path: alarms.tags.equipment
prefix: true
includes:
- equipment_id
- display_name
- equipment_type