Phase 5 deployed to lin001 for the first time, and the cube service as
committed could not answer a single query. Three separate faults, none of them
visible without deploying it.
- CUBEJS_EXT_DB_TYPE: postgres is not supported in Cube v1. Cube Store is the
only external pre-aggregation store, and it is also the default cache and
queue driver, so naming Postgres failed EVERY query - not just
pre-aggregated ones - with "It`s not possible to use Cube Store as
queue/cache driver without using it as external". So cubestore is now a
service: pinned in lockstep with cube, ai-internal only, no ports, data on
/datadisk because it grows and / is 62 GB.
CUBEJS_CACHE_AND_QUEUE_DRIVER: memory is NOT a way out. It does not fall
back - it hangs /readyz and every query indefinitely, logging nothing at
level warn. That cost longer to diagnose than the original error.
This is a deviation from the build spec, which says pre-aggregations
materialise into pg-ai schema cube_preagg. They cannot, on this version.
cube_preagg and its grants in 003_roles.sql stay, unused, so that nothing
else has to change if a later Cube restores Postgres as an external store.
- CUBEJS_REFRESH_WORKER was never set, so nothing built the pre-aggregations.
A query matching a rollup does not fall back to the source: it fails with
"No pre-aggregation partitions were built yet". max_value, min_value and
sample_count were dead on arrival while avg_value worked, which reads as a
per-measure bug and is not one.
- The healthcheck ran wget, which is not in the image (nor is curl). A
perfectly healthy cube reported unhealthy on every deploy, training the
reader to ignore the one signal that would show a real fault. It uses node,
which the image does have.
Verified on lin001: cube healthy, no published host ports, ai-internal and
proxy only, and a query matching alarms_by_hour now returns external: true -
served from Cube Store rather than scanning the source.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
259 lines
11 KiB
YAML
259 lines
11 KiB
YAML
# =============================================================================
|
|
# ai-compose.yml -> deployed to ~/ai-compose.yml on yau-sls-poc-lin001
|
|
#
|
|
# House style, inherited from ~/docker-compose.yml (host brief section 7):
|
|
# - restart: unless-stopped on everything
|
|
# - log rotation 10 MB x 3 on everything
|
|
# - NO published host ports: reach services through Caddy on the proxy network
|
|
# - secrets in 0600 env files under ~/ai/, never here and never in Git
|
|
#
|
|
# Orphan-container warnings are expected (shared Compose project name) - ignore.
|
|
#
|
|
# docker compose -f ~/ai-compose.yml up -d
|
|
# =============================================================================
|
|
|
|
services:
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# pg-ai - pgvector, reference data, Cube pre-aggregations.
|
|
# Deliberately NOT on proxy: no UI, nothing outside the AI stack reaches it.
|
|
# Pinned image - do NOT add to Watchtower's update list.
|
|
# ---------------------------------------------------------------------------
|
|
pg-ai:
|
|
image: pgvector/pgvector:pg16
|
|
container_name: pg-ai
|
|
restart: unless-stopped
|
|
networks: [ai-internal]
|
|
env_file:
|
|
- /home/azureuser/ai/pg-ai.env # 0600, not in Git
|
|
environment:
|
|
POSTGRES_DB: plant
|
|
POSTGRES_USER: postgres
|
|
PGDATA: /var/lib/postgresql/data/pgdata
|
|
volumes:
|
|
- /datadisk/pg-ai:/var/lib/postgresql/data
|
|
healthcheck:
|
|
test: ["CMD-SHELL", "pg_isready -U postgres -d plant"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
logging:
|
|
driver: json-file
|
|
options: { max-size: "10m", max-file: "3" }
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# cubestore - Cube's own store. Queue, cache and pre-aggregations. Deployed
|
|
# because Cube v1 does not run without it, not because we wanted another
|
|
# container. Version pinned in lockstep with cube: a mismatched pair is a
|
|
# documented Cube failure mode. Do NOT add either to Watchtower's list.
|
|
#
|
|
# Data goes on /datadisk. It grows with the pre-aggregations, and / is 62 GB
|
|
# and has hit 100% on this host before.
|
|
# sudo install -d -o 1000 -g 1000 /datadisk/cubestore
|
|
# No ports, ai-internal only: nothing outside the AI stack talks to it, and
|
|
# it has no authentication of its own.
|
|
# ---------------------------------------------------------------------------
|
|
cubestore:
|
|
image: cubejs/cubestore:v1.1.7
|
|
container_name: cubestore
|
|
restart: unless-stopped
|
|
networks: [ai-internal]
|
|
environment:
|
|
CUBESTORE_DATA_DIR: /cube/data
|
|
volumes:
|
|
- /datadisk/cubestore:/cube/data
|
|
logging:
|
|
driver: json-file
|
|
options: { max-size: "10m", max-file: "3" }
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# cube - semantic layer. Reads imh over TDS/1433 with the read-only login, or
|
|
# the fixture tables in pg-ai while USE_FIXTURES=true. Pinned - not in
|
|
# Watchtower's list.
|
|
#
|
|
# PRE-AGGREGATIONS LIVE IN CUBE STORE, NOT pg-ai. This is a deviation from
|
|
# the build spec, forced by the pinned version and found on lin001 deploying
|
|
# Phase 5, not in review. Cube v1 will not materialise pre-aggregations into
|
|
# Postgres - Cube Store is the only supported external store - so naming
|
|
# Postgres as CUBEJS_EXT_DB_TYPE fails every query with
|
|
# "It`s not possible to use Cube Store as queue/cache driver without using
|
|
# it as external"
|
|
# Cube Store is also the queue and cache driver, and it is not optional:
|
|
# CUBEJS_CACHE_AND_QUEUE_DRIVER=memory does not fall back, it hangs /readyz
|
|
# and every query forever, with nothing in the log at level warn.
|
|
# So `cubestore` below is a hard dependency of cube, not an optimisation.
|
|
# cube_preagg in pg-ai and its grants in db/003_roles.sql are now unused;
|
|
# they are left in place rather than dropped, because nothing else changes if
|
|
# a future Cube version restores Postgres as an external store.
|
|
# ---------------------------------------------------------------------------
|
|
cube:
|
|
image: cubejs/cube:v1.1.7
|
|
container_name: cube
|
|
restart: unless-stopped
|
|
depends_on:
|
|
pg-ai:
|
|
condition: service_healthy
|
|
cubestore:
|
|
condition: service_started
|
|
networks: [ai-internal, proxy]
|
|
env_file:
|
|
- /home/azureuser/ai/api.env # 0600, not in Git
|
|
environment:
|
|
CUBEJS_DEV_MODE: "false"
|
|
CUBEJS_LOG_LEVEL: warn
|
|
# Queue, cache and pre-aggregation store. Mandatory - see the note above.
|
|
CUBEJS_CUBESTORE_HOST: cubestore
|
|
CUBEJS_CUBESTORE_PORT: "3030"
|
|
# Without this NOTHING builds the pre-aggregations, and every query that
|
|
# matches one fails with "No pre-aggregation partitions were built yet"
|
|
# rather than falling back to the source. Single node, so the API
|
|
# instance is also the refresh worker.
|
|
CUBEJS_REFRESH_WORKER: "true"
|
|
# CUBEJS_DB_* and CUBEJS_API_SECRET come from api.env.
|
|
volumes:
|
|
- /home/azureuser/ai/cube/model:/cube/conf/model:ro
|
|
healthcheck:
|
|
# The image has neither wget nor curl - the original wget healthcheck
|
|
# marked a perfectly healthy cube unhealthy on every deploy. node is
|
|
# what it does have.
|
|
test: ["CMD", "node", "-e", "require('http').get('http://localhost:4000/readyz', r => process.exit(r.statusCode === 200 ? 0 : 1)).on('error', () => process.exit(1))"]
|
|
interval: 30s
|
|
timeout: 5s
|
|
retries: 3
|
|
logging:
|
|
driver: json-file
|
|
options: { max-size: "10m", max-file: "3" }
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# ai-api - FastAPI. Classifier, agent, contracts, guardrails.
|
|
# ---------------------------------------------------------------------------
|
|
ai-api:
|
|
build:
|
|
context: /home/azureuser/ai/api
|
|
dockerfile: Dockerfile
|
|
image: yau/ai-api:local
|
|
container_name: ai-api
|
|
restart: unless-stopped
|
|
depends_on:
|
|
pg-ai:
|
|
condition: service_healthy
|
|
networks: [ai-internal, proxy]
|
|
env_file:
|
|
- /home/azureuser/ai/api.env # 0600, not in Git
|
|
# Phase 9 - the upload inbox, and the ONLY writable path this container
|
|
# gets. Deliberately not /datadisk/ai-docs: a file that has been uploaded
|
|
# but not yet approved must not be visible to `ai-ingest --all`.
|
|
#
|
|
# COMMENTED OUT UNTIL PHASE 9, on the same principle as caddy/ai-routes.caddy:
|
|
# add each piece at the phase that needs it. Docker creates a missing bind
|
|
# source as a ROOT-OWNED directory, and ai-api does not run as root - so
|
|
# deploying this before the directory exists gives you an inbox the API
|
|
# cannot write to, on the growing disk of a live shared host.
|
|
# Create it first, then uncomment:
|
|
# sudo install -d -o 10002 -g 10002 /datadisk/ai-docs-inbox
|
|
# volumes:
|
|
# - /datadisk/ai-docs-inbox:/inbox
|
|
healthcheck:
|
|
test: ["CMD", "python", "-m", "app_healthcheck"]
|
|
interval: 30s
|
|
timeout: 5s
|
|
retries: 3
|
|
logging:
|
|
driver: json-file
|
|
options: { max-size: "10m", max-file: "3" }
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# ai-web - React/Vite build served by nginx. proxy only; the browser talks to
|
|
# the API through its public hostname, so it needs nothing on ai-internal.
|
|
# ---------------------------------------------------------------------------
|
|
ai-web:
|
|
build:
|
|
context: /home/azureuser/ai/web
|
|
dockerfile: Dockerfile
|
|
image: yau/ai-web:local
|
|
container_name: ai-web
|
|
restart: unless-stopped
|
|
networks: [proxy]
|
|
logging:
|
|
driver: json-file
|
|
options: { max-size: "10m", max-file: "3" }
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# ai-ingest - on demand, not a service. Docling -> chunk -> embed -> pg-ai.
|
|
# docker compose -f ~/ai-compose.yml run --rm ai-ingest --all
|
|
# The profile keeps it out of `up -d`.
|
|
# ---------------------------------------------------------------------------
|
|
ai-ingest:
|
|
build:
|
|
context: /home/azureuser/ai/ingest
|
|
dockerfile: Dockerfile
|
|
image: yau/ai-ingest:local
|
|
container_name: ai-ingest
|
|
profiles: [ingest]
|
|
restart: "no"
|
|
networks: [ai-internal]
|
|
env_file:
|
|
- /home/azureuser/ai/api.env # 0600, not in Git
|
|
volumes:
|
|
- /datadisk/ai-docs:/docs:ro
|
|
logging:
|
|
driver: json-file
|
|
options: { max-size: "10m", max-file: "3" }
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# ai-docs-worker - Phase 9. The ai-ingest IMAGE with worker.py as entrypoint,
|
|
# so an uploaded document is parsed and chunked by exactly the same code as a
|
|
# file ingested from the command line - by construction, not by discipline.
|
|
#
|
|
# Two jobs: pre-scan `uploaded` rows for a header proposal, and publish
|
|
# `approved` ones. It is the only container that can write doc_chunks
|
|
# (ingest_rw), and it has no HTTP surface and no place on the proxy network.
|
|
#
|
|
# /docs is READ-WRITE here, unlike the ai-ingest CLI service above, because
|
|
# publishing moves the approved file into the folder that determines its
|
|
# doc_type. That is the one write, and it happens only after a human has
|
|
# confirmed the header.
|
|
#
|
|
# Both mounts must be writable by the image's uid 10002 (ingestuser):
|
|
# sudo install -d -o 10002 -g 10002 /datadisk/ai-docs-inbox
|
|
# sudo chown -R 10002:10002 /datadisk/ai-docs
|
|
# ai-api writes the inbox as its own non-root uid - give the inbox group
|
|
# write and put both uids in the group rather than making it world-writable.
|
|
# ---------------------------------------------------------------------------
|
|
ai-docs-worker:
|
|
build:
|
|
context: /home/azureuser/ai/ingest
|
|
dockerfile: Dockerfile
|
|
image: yau/ai-ingest:local
|
|
container_name: ai-docs-worker
|
|
# Kept out of `up -d` by the profile, exactly as ai-ingest is. worker.py
|
|
# does not exist yet, so an unguarded service here would crash-loop on a
|
|
# live shared host. Drop the profile in the same commit that adds the file.
|
|
# docker compose -f ~/ai-compose.yml --profile worker up -d ai-docs-worker
|
|
profiles: [worker]
|
|
restart: unless-stopped
|
|
depends_on:
|
|
pg-ai:
|
|
condition: service_healthy
|
|
networks: [ai-internal]
|
|
env_file:
|
|
- /home/azureuser/ai/api.env # 0600, not in Git
|
|
entrypoint: ["python", "worker.py"]
|
|
command: []
|
|
volumes:
|
|
- /datadisk/ai-docs-inbox:/inbox
|
|
- /datadisk/ai-docs:/docs # rw - see above
|
|
# Withdrawn documents are MOVED here, not deleted. It is outside the four
|
|
# doc_type folders on purpose: ingest_file() inserts every chunk with
|
|
# superseded = FALSE, so a withdrawn file left in /docs comes back LIVE on
|
|
# the next `ai-ingest --all`.
|
|
- /datadisk/ai-docs-withdrawn:/withdrawn
|
|
logging:
|
|
driver: json-file
|
|
options: { max-size: "10m", max-file: "3" }
|
|
|
|
networks:
|
|
ai-internal:
|
|
driver: bridge
|
|
proxy:
|
|
external: true
|