diff --git a/CONTEXT.md b/CONTEXT.md index 88767533..f4404c71 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -14,6 +14,15 @@ > language, one screen, no identifiers in the prose. Same subjects, different readers; merging them > would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`** > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +## Rulings 2026-09-25 (evening, operator) + +- **`09` §3 decision 35 — part 10 is built now, docmost first** (D1 option A). The box converts a PostgreSQL + major as a guarded-update step; each other app needs its own two-venue proof; the engine-major gate stays for + every app without it. +- **`09` §3 decision 36 — a reinstall over kept data offers "use my kept data" / "start fresh"** (D2 option A), + and kept data is never a dead end: visible (read only), loadable, deletable by the household. +- **D3 is OPEN** (in `STATUS.md`): does the box ever delete kept data by itself? Until ruled, nothing does. + ## 2026-09-25 (midday) — Peti retired in fact; controller v0.272.0 - Peti's hub customer deleted through the cascade (journal #20); Storage Box sub-account 269130 (`u629488-sub2`) diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index cd332c60..06a953c6 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -402,6 +402,46 @@ R-636's louder repeated alarm. behaviour shipped in v0.271.0 is now the rule. **Why:** one night's delay for an app is cheap, and a resume would be one more mechanism acting with nobody watching. +### 2026-09-25 (evening) — two operator rulings + +35. **Part 10 is built now, one app first** (operator ruling 2026-09-25 evening, D1 option A). The box + converts a PostgreSQL major as a step of the guarded update. **docmost is the first app.** Each other app + needs its own proof on both venues (the test bench and a scratch box through the real guarded Update) + before the catalog may move it. The engine-major gate stays for every app without that proof. **Why:** + decision 16 said "each of the eleven apps is proven before the catalog may move it"; one app first makes + the first proof small enough to read. As built: §6.4 part 10. +36. **A reinstall over kept data offers a choice** (operator ruling 2026-09-25 evening, D2 option A): "use + my kept data" (the database from a backup, with the kept files) or "start fresh" (the kept files move to + a dated folder; nothing is deleted). **And, from the operator's follow-up question: kept data must never + be a dead end.** The household can see it (read only), load it into a later install, and delete it. + Support is for the rare case, not the normal one. **Not ruled (D3, in `STATUS.md`):** whether the box + ever deletes kept data by itself. Until it is ruled, **nothing deletes kept data automatically** — only a + household's explicit Delete. Where it lives and who deletes it: `07-backup-architecture.md` §"Kept data". + +### 2026-09-25 (evening) — two decisions taken by CC unattended, building part 10 + +37. **docmost converts PostgreSQL 16 → 18, not 16 → 17** — *decided by CC unattended 2026-09-25 — operator may + reverse.* *One sentence:* which major does the first conversion target? **Options:** (a) 17 — no mount + change; (b) 18 — the data volume's mount moves to `/var/lib/postgresql` in the same step. **Costs:** (a) every + box converts again later, and docmost's own upstream compose already ships `postgres:18` at + `/var/lib/postgresql`; (b) the step changes the mount point — measured: `postgres:18` REFUSES (exit 1) even an + EMPTY volume at `/var/lib/postgresql/data` — so the step's definition carries the mount, and the undo must put + the old mount back with the old bytes, which it does (it restores the old definition and every volume). + **Why (b):** one conversion is better than two when the app's upstream runs the newer major. Reversible: the + catalog could pin 17 instead before any customer box moves. `audits/night-2026-09-26/A/README.md` A1. +38. **The box loads from a `pg_dumpall` of the OLD engine, with no error tolerated** — *decided by CC unattended + 2026-09-25 — operator may reverse.* *One sentence:* what does the conversion load into the new engine? + **Options:** (a) the existing safety dump (`pg_dump` per database, `--no-owner --no-privileges`); (b) a + `pg_dumpall` taken from the old engine at conversion time, loaded with `ON_ERROR_STOP`. **Costs:** (a) is free, + and silently drops any role, grant or database setting beyond the bootstrap ones (for docmost both routes gave + identical databases — it has one role and one database; the other ten are not measured); (b) one more dump + (0.6 s for docmost) and a load that must step around exactly two objects the new engine's entrypoint makes (the + bootstrap role and database — measured: loaded raw, exactly two `already exists` errors). **Why (b):** decision + 16 says "save everything"; the two collisions are removed precisely (the entrypoint's databases dropped only + when they hold no table; `CREATE ROLE ;` skipped only for a role that exists — its `ALTER ROLE … PASSWORD` + still runs), so ANY other error stops the load and the undo runs. Proven live: an extension 18 lacks + (`adminpack`) stopped the load and the box undid it. `audits/night-2026-09-26/A/README.md` A2. + --- ## 3b. ANSWERED 2026-09-23 — the seven questions Slices 6 and 7 needed @@ -1175,7 +1215,7 @@ what the part can do to a household's data if it is wrong, not how likely that i | **7** | **SHIPPED — controller v0.271.0 (`9cf13a3`), proven on 9202 over four simulated nights 2026-09-24/25 and watched through the demo boxes' first real automatic night (`audits/DRILL-night-2026-09-25.md`).** Decisions 31–33 taken unattended. Gaps a scratch box cannot close: R-687; no resume after a restart: R-686. *(Before:)* **SPIKED 2026-09-24 (night, Part C) — the measured build brief is §6.4.2 below.** **The update leg in the chain + the automatic caller + the switch.** A leg that starts when the off-site leg has ENDED (on every exit path, skips included), one app at a time, one step per press, `app_update.unattended` default ON (decision 12), `stacks.update_window` removed, reads `UpdateRefusal.Reason`, remembers a failed step so it never re-presses it, and the full-system backup's gate waits for it until W+5h (decision 20). | 11, 12, 14, 15, 20 | **3** | 1, 2, 5 | medium — the only part that acts with nobody watching; everything above is what makes it safe | | **8** | **SHIPPED — controller v0.265.0 + hub v0.121.0, proven live on 9202 2026-09-23** (`audits/cleanup-2026-09-23/`; a kernel `oom_kill` counter, not the sticky flag — `08` §6.2). **R-636** — the same OOM key re-firing escalates instead of staying one `warning` for six hours. | R-636 | **1** | — | none | | **9** | **SHIPPED — controller v0.265.0, proven live on 9202 2026-09-23** (badge „Megállítva — visszaállítás szükséges" / "Stopped — restore needed", no Update button, 409 `held` unchanged). **R-625** — a held app stops offering an Update it will refuse. With the undo, holds become rarer; the lie on the page does not go away by itself. | R-625 | **0.5** | 1 | none | -| **10** | **PostgreSQL majors converted by the box.** A guarded-update step: `pg_dumpall` from the old engine, a NEW datadir (the old one kept aside, never deleted, until the check passes), load, check; then each of the eleven apps proven on the bench before its catalog move. | 16, R-463 | **2 + 3** | 1 (the same load discipline), 4 | **HIGH** — it rebuilds the datadir; bounded by keeping the old datadir aside | +| **10** | **SHIPPED FOR DOCMOST — controller v0.273.0 + catalog `6a4a5f0` (harness v4, the gate's proof clause), proven live on 9202 2026-09-25** (`audits/night-2026-09-26/D/`): one press converted docmost 16 → 18 in **9.9 s** of engine work (**42 s** end to end), the check equal (2 databases, 48 tables, 71 rows), `PG_VERSION` 18, the seed read back; the load failing (an extension 18 lacks) → **undone in 43.6 s**; the app unhealthy on 18 → **undone in 148 s** (90 s drill health timeout); the controller SIGKILLed one second after the volume was emptied → the restart **undid it in 40 s**; each ending on 16 with the seed read back. The old datadir's copy is kept until a backup is proven after the conversion. Decisions 35, 37, 38. Every other PostgreSQL app stays refused by the gate until its own two-venue proof. *(Before:)* **PostgreSQL majors converted by the box.** A guarded-update step: `pg_dumpall` from the old engine, a NEW datadir (the old one kept aside, never deleted, until the check passes), load, check; then each of the eleven apps proven on the bench before its catalog move. | 16, R-463 | **2 + 3** | 1 (the same load discipline), 4 | **HIGH** — it rebuilds the datadir; bounded by keeping the old datadir aside | | **11** | **Fleet view** — per compose service: installed ref, catalog ref, badge state in the report; the hub lists boxes behind. | 18, R-451 | 2 | — | none — **deferred by the ruling** until the fleet grows | **Recommended order: 1 → 2 + 3 → 4 → 5 → 6 → 7 → 8 → 9 → 10**, part 11 when the fleet grows. **Total diff --git a/documentation/audits/night-2026-09-26/A/A1-images.txt b/documentation/audits/night-2026-09-26/A/A1-images.txt new file mode 100644 index 00000000..86a2bf26 --- /dev/null +++ b/documentation/audits/night-2026-09-26/A/A1-images.txt @@ -0,0 +1,38 @@ +# A1 — what the postgres images declare (9202, 2026-09-25T10:20:36Z) +## postgres:16-alpine +digest=postgres@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea PGDATA/ENV=PG_MAJOR=16 PG_VERSION=16.15 PG_SHA256=c1575341fa7bd40f5274ea465b34390f4dc64cdd0770af327005caaeb9f6b7ed PGDATA=/var/lib/postgresql/data VOLUME={"/var/lib/postgresql/data":{}} +## postgres:17-alpine +digest=postgres@sha256:b0f9560a2de083e2cc7382e75f808c7381a32852a7ec49117deedb300e552b24 PGDATA/ENV=PG_MAJOR=17 PG_VERSION=17.11 PG_SHA256=dd27f2b3c59e73ed14aa3324901242bf69a032a6347805f274e6260322d42979 PGDATA=/var/lib/postgresql/data VOLUME={"/var/lib/postgresql/data":{}} +## postgres:18-alpine +digest=postgres@sha256:77f585114c32fbca283dc835b0596f4e52b51b4c6662d7810b2f4084f60a1873 PGDATA/ENV=PG_MAJOR=18 PG_VERSION=18.6 PG_SHA256=555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f PGDATA=/var/lib/postgresql/18/docker VOLUME={"/var/lib/postgresql":{}} + +# A1b — postgres:18-alpine with an EMPTY volume at /var/lib/postgresql/data (the templates' mount) +state: exited exit=1 +Error: in 18+, these Docker images are configured to store database data in a + format which is compatible with "pg_ctlcluster" (specifically, using + major-version-specific directory names). This better reflects how + PostgreSQL itself works, and how upgrades are to be performed. + + See also https://github.com/docker-library/postgres/pull/1259 + + Counter to that, there appears to be PostgreSQL data in: + /var/lib/postgresql/data (unused mount/volume) + + This is usually the result of upgrading the Docker image without + upgrading the underlying database using "pg_upgrade" (which requires both + versions). + + The suggested container configuration for 18+ is to place a single mount + at /var/lib/postgresql which will then place PostgreSQL data in a + subdirectory, allowing usage of "pg_upgrade --link" without mount point + boundary issues. + + See https://github.com/docker-library/postgres/issues/37 for a (long) + discussion around this process, and suggestions for how to do so. + +# A1c — control: postgres:18-alpine with an empty volume at /var/lib/postgresql (its own layout) +state: running +2026-09-25 10:21:17.475 UTC [1] LOG: listening on Unix socket "/var/run/postgresql/.s.PGSQL.5432" +2026-09-25 10:21:17.488 UTC [66] LOG: database system was shut down at 2026-09-25 10:21:17 UTC +2026-09-25 10:21:17.496 UTC [1] LOG: database system is ready to accept connections +18 diff --git a/documentation/audits/night-2026-09-26/A/A2-A3-A5-measure.txt b/documentation/audits/night-2026-09-26/A/A2-A3-A5-measure.txt new file mode 100644 index 00000000..37e81d3a --- /dev/null +++ b/documentation/audits/night-2026-09-26/A/A2-A3-A5-measure.txt @@ -0,0 +1,37 @@ +# A2/A3/A5 — docmost seeded on 9202, measured on a COPY of its datadir (2026-09-25T10:25:58Z) +datadir size: 65.5M +A3 counts (old engine): 0.49s, 48 tables, 62 rows total +db:docmost owner=docmost enc=UTF8 coll=en_US.utf8 +db:postgres owner=docmost enc=UTF8 coll=en_US.utf8 +role:docmost super=true login=true pw=true +ext:docmost:pg_trgm +ext:docmost:plpgsql +ext:docmost:unaccent +A2 pg_dumpall: 0.59s, 132184 B, tail: -- PostgreSQL database dump complete +A2 pg_dump --no-owner --no-privileges docmost: 0.27s, 123545 B +what pg_dumpall carries that the per-db dump does not: + 1 CREATE ROLE docmost; + 1 CREATE DATABASE docmost WITH TEMPLATE = template0 ENCODING = 'UTF8' LOCALE_PROVIDER = libc LOCALE = 'en_US.utf8'; + 1 ALTER TABLE public.workspaces OWNER TO docmost; + 1 ALTER TABLE public.workspace_invitations OWNER TO docmost; + 1 ALTER TABLE public.watchers OWNER TO docmost; + 1 ALTER TABLE public.users OWNER TO docmost; + 1 ALTER TABLE public.user_tokens OWNER TO docmost; + 1 ALTER TABLE public.user_sessions OWNER TO docmost; + 1 ALTER TABLE public.user_mfa OWNER TO docmost; + 1 ALTER TABLE public.templates OWNER TO docmost; + 1 ALTER TABLE public.spaces OWNER TO docmost; + 1 ALTER TABLE public.space_members OWNER TO docmost; +## route dumpall: new engine init 2.64s, load 1.76s, counts 0.48s, PG_VERSION=18, load errors=2 + 1 ERROR: database "docmost" already exists + 1 ERROR: role "docmost" already exists + dbs/owners/extensions/row counts: IDENTICAL + roles before: role:docmost super=true login=true pw=true + roles after: role:docmost super=true login=true pw=true + datadir after: 49.9M +## route pgdump: new engine init 2.64s, load 1.44s, counts 0.49s, PG_VERSION=18, load errors=0 + dbs/owners/extensions/row counts: IDENTICAL + roles before: role:docmost super=true login=true pw=true + roles after: role:docmost super=true login=true pw=true + datadir after: 49.8M +A5 space: copy of the old datadir = 68707765 B; dump = 132184 B; new datadir = 52271397 B diff --git a/documentation/audits/night-2026-09-26/A/A4-undo-after-empty.txt b/documentation/audits/night-2026-09-26/A/A4-undo-after-empty.txt new file mode 100644 index 00000000..f1fb0eae --- /dev/null +++ b/documentation/audits/night-2026-09-26/A/A4-undo-after-empty.txt @@ -0,0 +1,13 @@ +# A4 — the undo after the DB volume was EMPTIED, by hand with the product's own helper commands (undo.go Copy/Restore), 9202 2026-09-25T10:27:34Z +app stopped: 0 running +copy with marker: ok +volume emptied: 0 entries left +Restore (the product's command): ok +PG_VERSION after restore: 16 +2026-09-25 12:27:39.221 CEST [29] LOG: database system was shut down at 2026-09-25 12:27:35 CEST +2026-09-25 12:27:39.230 CEST [1] LOG: database system is ready to accept connections +copy removed +12:34:29 docmost: login as the seeded user http=200 ok=True +12:34:29 A4-after-undo-retry: read docmost=True +12:34:30 docmost: login as the seeded user http=200 ok=True +12:34:30 A4-after-undo-productstart: read docmost=True diff --git a/documentation/audits/night-2026-09-26/A/README.md b/documentation/audits/night-2026-09-26/A/README.md new file mode 100644 index 00000000..9669dff7 --- /dev/null +++ b/documentation/audits/night-2026-09-26/A/README.md @@ -0,0 +1,88 @@ +# Part A — the PostgreSQL conversion spike on docmost (9202, 2026-09-25 midday) + +Written BEFORE any build. docmost 0.96.0 / `postgres:16-alpine` installed on 9202 through the product +(`POST /api/stacks/docmost/deploy`), seeded through its own API (a workspace + an account), read back +(`data-readback.txt`: `A-read0`). Every measurement below ran on a COPY of its datadir. + +## A1 — the target major: 18 (CC-unattended decision, `09` §3 decision 37) + +**Measured (`A1-images.txt`):** + +- `postgres:16-alpine` and `17-alpine`: `PGDATA=/var/lib/postgresql/data`, `VOLUME /var/lib/postgresql/data`. +- `postgres:18-alpine` (18.6): `PGDATA=/var/lib/postgresql/18/docker`, `VOLUME /var/lib/postgresql`. +- **18 with an EMPTY volume at `/var/lib/postgresql/data` (where all eleven templates mount it) REFUSES: + exit 1** — *"in 18+, these Docker images are configured to store database data in a format which is + compatible with pg_ctlcluster … there appears to be PostgreSQL data in: /var/lib/postgresql/data (unused + mount/volume)"*. It refuses even when that mount is empty. Control: the same image with the volume at + `/var/lib/postgresql` initialises and serves, `PG_VERSION` 18. +- **docmost's upstream compose ships `image: postgres:18` with `db_data:/var/lib/postgresql`** + (github docmost/docmost `docker-compose.yml`, read 2026-09-25). Its docs name no other version. + +**So the brief's claim is TRUE**, and stronger than stated: 18 refuses at the old mount, even empty. A move to +18 therefore needs the mount moved to `/var/lib/postgresql` in the same step, which the step's own definition +carries (the template's compose / `steps/.yml`). + +**Decision 37 — docmost converts 16 → 18, not 16 → 17.** +*One sentence:* which major does the first conversion target? **Options:** (a) 17 — no mount change; +(b) 18 — the mount moves to `/var/lib/postgresql` in the same step. **Costs:** (a) is a second conversion +later, and upstream already ships 18, so every box converts twice; (b) the step changes the volume's mount +point, so an undo must put the OLD mount back with the old data — which the undo does anyway (it restores the +old definition and the volume's bytes). **Why (b):** one conversion is better than two when the app's own +upstream runs the newer major; the mount change rides the step's definition and the undo's existing +definition restore. Reversible (the catalog can pin 17 instead before any box moves). + +## A2 — what to load from: `pg_dumpall` from the old engine (decision 38) + +`A2-A3-A5-measure.txt`. On the seeded datadir (48 tables, 62 rows, 49 MB): + +| | size | time | what it carries | +|---|---|---|---| +| `pg_dumpall` (old engine) | 132 184 B | 0.59 s | roles WITH their password hashes, `CREATE DATABASE … LOCALE`, owners, grants | +| `pg_dump --no-owner --no-privileges docmost` (today's safety dump) | 123 545 B | 0.27 s | one database's schema + data; no roles, no owners, no database settings | + +Both loaded into a fresh 18 gave **identical** databases, owners, encodings, collations, extensions +(`pg_trgm`, `plpgsql`, `unaccent`) and row counts for docmost — because docmost has ONE role (the bootstrap +superuser the new engine's entrypoint recreates from the same env) and ONE database. + +**So the brief's claim is TRUE** (the safety dump is per-database `pg_dump` with `--no-owner --no-privileges` +— `appbackup/dbdump.go`), **and for docmost it would have been enough.** It would not be for an app with a +second role or a second database, which the other ten have not been measured for. + +**Decision 38 — the box loads from a `pg_dumpall` taken from the OLD engine at conversion time, and the load +tolerates NO error.** *One sentence:* what does the box load into the new engine? **Options:** (a) the +existing safety dump; (b) a `pg_dumpall` of the old engine, loaded with `ON_ERROR_STOP`. **Costs:** (a) is +free, and silently drops any role, grant or database setting beyond the bootstrap ones; (b) one more dump +(0.6 s here) and a load that must not trip on the two objects the new engine's entrypoint already made (the +bootstrap role and database: measured, `pg_dumpall` loaded raw gives exactly two `already exists` ERRORs). +**Why (b):** decision 16 says "save everything". The two known collisions are removed precisely — the +entrypoint's empty databases are dropped before the load (only when they hold no table), and the dump's +`CREATE ROLE ;` line is skipped for a role that already exists (its `ALTER ROLE … PASSWORD` still +runs) — so any OTHER error stops the load and the undo runs. Proven by a test and live. + +## A3 — the check + +Old engine, before it stops; new engine, after the load — the same query set (0.5 s each here): +per database: owner, encoding, collation; every role (superuser, login, has a password); every extension by +name; every table's exact row count (`count(*)` per table). After: plus `PG_VERSION` of the new datadir +(`$PGDATA/PG_VERSION`, which follows 17's and 18's different layouts). All IDENTICAL for docmost on both +load routes. + +## A4 — the undo after the volume was EMPTIED: proven by hand + +`A4-undo-after-empty.txt`: copy with the product's own helper command (+ marker) → volume emptied (0 +entries) → the product's `Restore` command → `PG_VERSION` 16 → the old engine "ready to accept +connections" → **the seeded account logged in through docmost's front door** (`A4-after-undo-productstart`). +**The brief's claim is TRUE.** + +*Recorded as it happened:* my first start after the restore was a bare `docker compose up -d` in the stack +dir, which has no `.env` (the controller passes the app's env itself). docmost started without its secrets, +crash-looped, and **the box stopped it after 7 restarts in 10 min (decision 28) — the product working**, not +a fault. Started again with the product's own Start (`POST /api/stacks/docmost/start`); the seed read back. + +## A5 — space + +The conversion adds, on top of today's undo copy (the old datadir, 68.7 MB here): the dump (132 KB — 0.2 % of +the datadir) and nothing for the new datadir (it is built IN the emptied volume: 52.3 MB). The box's check: +**the dump's bound is the DB volume's own size** (a logical dump of live rows is smaller than the datadir +holding them, measured 0.2 %), with a 25 % margin, on the filesystem that holds the stack directory (where the +dump is written), plus the fixed 2 GB floor the update already keeps. Refused before anything moves. diff --git a/documentation/audits/night-2026-09-26/B/B-redproofs.txt b/documentation/audits/night-2026-09-26/B/B-redproofs.txt new file mode 100644 index 00000000..0b323b0a --- /dev/null +++ b/documentation/audits/night-2026-09-26/B/B-redproofs.txt @@ -0,0 +1,55 @@ +### RED-PROOF B1 no mark → refused — mutation in pgconvert.go +--- FAIL: TestConvert_MajorWithoutMarkIsRefused (0.00s) + pgconvert_test.go:215: the preflight let it through: +--- FAIL: TestConvert_JobRefusesAMarkThatDoesNotMatch (0.01s) + pgconvert_test.go:235: phase "done" key "", want failed with the no-test sentence +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.020s +### RED-PROOF B2 marker validated again before emptying — mutation in pgconvert.go +--- FAIL: TestConvert_NeverEmptiedWithoutMarker (0.01s) + pgconvert_test.go:338: Empty was called although the copy had no marker +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.015s +FAIL +### RED-PROOF B3 counts compared — mutation in pgconvert.go +--- FAIL: TestConvert_CountsDifferAreUndone (0.01s) + pgconvert_test.go:271: ended "done" (), want undone +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.021s +FAIL +### RED-PROOF B3 PG_VERSION checked — mutation in pgconvert.go +--- FAIL: TestConvert_WrongVersionIsUndone (0.01s) + pgconvert_test.go:304: ended "done" (), want undone +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.019s +FAIL +### RED-PROOF B3 cut-off dump refused — mutation in pgconvert.go +--- FAIL: TestConvert_CutOffDumpNeverEmptiesTheVolume (0.01s) + pgconvert_test.go:315: ended "done" (), want undone +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.018s +FAIL +### RED-PROOF B4 restart during converting → undo — mutation in update.go +--- FAIL: TestConvert_RestartDuringConvertingIsUndone (0.06s) + pgconvert_test.go:399: RecoverUpdates resumed [] +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.062s +FAIL +### RED-PROOF B5 old datadir copy kept — mutation in update.go +--- FAIL: TestConvert_HappyPath (0.01s) + pgconvert_test.go:196: the kept copy is not recorded: +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.014s +FAIL +### RED-PROOF B5 release wired — mutation in main.go +--- FAIL: TestConvert_ReleaseIsWiredAtStartup (0.01s) + conversion_wiring_test.go:11: ReleaseConversionCopies is never called — a converted app's old datadir copy would stay on disk for ever +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.015s +FAIL +### RED-PROOF A5 space checked before anything moves — mutation in update.go +--- FAIL: TestConvert_NoSpaceIsRefusedWithNothingMoved (0.01s) + pgconvert_test.go:353: phase "done" key "" +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.014s +FAIL diff --git a/documentation/audits/night-2026-09-26/B/B-tests-green.txt b/documentation/audits/night-2026-09-26/B/B-tests-green.txt new file mode 100644 index 00000000..3bbaf2d5 --- /dev/null +++ b/documentation/audits/night-2026-09-26/B/B-tests-green.txt @@ -0,0 +1,18 @@ +--- PASS: TestConvert_HappyPath (0.01s) +--- PASS: TestConvert_MajorWithoutMarkIsRefused (0.01s) +--- PASS: TestConvert_JobRefusesAMarkThatDoesNotMatch (0.00s) +--- PASS: TestConvert_CountsDifferAreUndone (0.01s) +--- PASS: TestConvert_LoadFailureIsUndone (0.01s) +--- PASS: TestConvert_HealthFailureIsUndone (0.01s) +--- PASS: TestConvert_WrongVersionIsUndone (0.01s) +--- PASS: TestConvert_CutOffDumpNeverEmptiesTheVolume (0.01s) +--- PASS: TestConvert_NeverEmptiedWithoutMarker (0.01s) +--- PASS: TestConvert_NoSpaceIsRefusedWithNothingMoved (0.00s) +--- PASS: TestConvert_RestartDuringConvertingIsUndone (0.06s) +--- PASS: TestConvert_ReleaseAfterAProvenBackup (0.01s) +--- PASS: TestConvert_PostgresMajorFromTags (0.00s) +--- PASS: TestConvert_PostgresFamilyMatchesTheBackupSide (0.00s) +--- PASS: TestConvert_FilterSkipsOnlyExactLines (0.00s) +--- PASS: TestConvert_ValidateDumpAll (0.00s) +--- PASS: TestConvert_MarkDoesNotChangeOldLadderPrints (0.00s) +ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.189s diff --git a/documentation/audits/night-2026-09-26/C/C1-bench-create.txt b/documentation/audits/night-2026-09-26/C/C1-bench-create.txt new file mode 100644 index 00000000..3df96999 --- /dev/null +++ b/documentation/audits/night-2026-09-26/C/C1-bench-create.txt @@ -0,0 +1,13 @@ +local dir active 40453376 22728440 15637820 56.18% +nvme-scratch dir active 983379700 75082704 858270384 7.64% +template debian-13-standard_13.6-1_arm64.tar.zst +sync_wait: 34 An error occurred in another process (expected sequence number 7) +__lxc_start: 2288 Failed to spawn container "9401" +startup for container '9401' failed +purging CT 9401 from related configurations.. +destroyed wrong-arch 9401 +debian-13-standard_13.6-1_amd64.tar.zst +192.168.0.121/24 +Docker version 26.1.5+dfsg1, build a72d7cd +Docker Compose version 2.26.1-4 +net-ok diff --git a/documentation/audits/night-2026-09-26/C/C1-harness-unit.txt b/documentation/audits/night-2026-09-26/C/C1-harness-unit.txt new file mode 100644 index 00000000..80ef54b8 --- /dev/null +++ b/documentation/audits/night-2026-09-26/C/C1-harness-unit.txt @@ -0,0 +1,11 @@ + ok 18 moves the data mount to /var/lib/postgresql + ok 18 leaves the redis mount alone + ok 16 keeps /var/lib/postgresql/data + ok back from 18 to 16 restores /data + ok pg_conversion names the one moving service + ok within a major is no conversion + ok pgvector's pg16 tag reads 16 + ok an entry with a good mark is well-formed + ok a mark naming a service outside `to` is refused + ok a mark going DOWN is refused +pg conversion tests: 0 failed diff --git a/documentation/audits/night-2026-09-26/C/C2-decoys-green.txt b/documentation/audits/night-2026-09-26/C/C2-decoys-green.txt new file mode 100644 index 00000000..84c2426b --- /dev/null +++ b/documentation/audits/night-2026-09-26/C/C2-decoys-green.txt @@ -0,0 +1,8 @@ + ok FACT: docmost-postgres 16-alpine -> 17-alpine bundled with the app bump rc=1 (expected 1) + ok FACT: docmost-postgres postgres:16-alpine -> 17-alpine rc=1 (expected 1) + ok GENUINE: docmost-postgres 16 -> 18 ALONE with its proven two-venue MARKED entry rc=0 (expected 0) + ok DECOY: the entry is proven on both venues but carries NO conversion mark rc=1 (expected 1) + ok DECOY: the marked entry cites ONE venue only (no box_evidence) rc=1 (expected 1) + ok DECOY: the mark sits in a COMMENT, the entry has none rc=1 (expected 1) + ok DECOY: the marked entry, but the move is BUNDLED with the app's own bump rc=1 (expected 1) + ok FACT: adventurelog's postgis 16 -> 17 (the postgis family was never judged before) rc=1 (expected 1) diff --git a/documentation/audits/night-2026-09-26/C/C2-redproof.txt b/documentation/audits/night-2026-09-26/C/C2-redproof.txt new file mode 100644 index 00000000..01330786 --- /dev/null +++ b/documentation/audits/night-2026-09-26/C/C2-redproof.txt @@ -0,0 +1,7 @@ +# RED-PROOF of check-engine-major.py's PostgreSQL clause (2026-09-25, catalog tree before commit) +# mutation: `if why is None and not others:` -> `if not others:` (the ladder proof is not consulted) +# python3 scripts/test_gate_decoys.py, lines quoted verbatim: +FAIL: FACT: docmost-postgres 16-alpine -> 17-alpine bundled with the app bump: rc=0 expected 1; missing ['ENGINE-MAJOR GATE FAILED'] +FAIL: FACT: docmost-postgres postgres:16-alpine -> 17-alpine: rc=0 expected 1; missing ['R-463'] +FAIL: DECOY: the entry is proven on both venues but carries NO conversion mark: rc=0 expected 1; missing ['NOT PROVEN FOR THIS APP', 'engine_conversion None'] +# tree restored from the saved copy; the suite green again: C2-decoys-green.txt diff --git a/documentation/audits/night-2026-09-26/D/D0-deploy-rc1-9202.txt b/documentation/audits/night-2026-09-26/D/D0-deploy-rc1-9202.txt new file mode 100644 index 00000000..dd4b22d6 --- /dev/null +++ b/documentation/audits/night-2026-09-26/D/D0-deploy-rc1-9202.txt @@ -0,0 +1,5 @@ +# 9202 before: gitea.dooplex.hu/admin/felhom-controller:0.272.0 +gitea.dooplex.hu/admin/felhom-controller:0.273.0-rc1 +# 9202 after: gitea.dooplex.hu/admin/felhom-controller:0.273.0-rc1 healthy +2026/09/25 11:04:49 main.go:336: [INFO] felhom-controller 0.273.0-rc1 starting (customer: demo-hp, domain: enkisfelhom.hu) +2026/09/25 11:04:49 undo.go:675: [INFO] [stacks] applied-meta backfill (R-646): recorded 0 []; skipped 0 [] — not current with the catalog, their pinned version's .felhom.yml is no longer on the box diff --git a/documentation/audits/night-2026-09-26/D/D0-repoint-drill.txt b/documentation/audits/night-2026-09-26/D/D0-repoint-drill.txt new file mode 100644 index 00000000..f2c89786 --- /dev/null +++ b/documentation/audits/night-2026-09-26/D/D0-repoint-drill.txt @@ -0,0 +1,10 @@ +git: + branch: main + repo_url: https://gitea.dooplex.hu/admin/app-catalog-drill.git + sync_interval: 15m + token: + username: "admin" +hub: +update: + health_timeout: 90s + diff --git a/documentation/audits/night-2026-09-26/D/D1-happy-path.txt b/documentation/audits/night-2026-09-26/D/D1-happy-path.txt new file mode 100644 index 00000000..f04d3f45 --- /dev/null +++ b/documentation/audits/night-2026-09-26/D/D1-happy-path.txt @@ -0,0 +1,48 @@ +# D1-happy-path — 2026-09-25T13:12:42+0200; controller gitea.dooplex.hu/admin/felhom-controller:0.273.0-rc1 +before: PG_VERSION/image=['16', 'postgres:16-alpine@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} catalog={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:18-alpine', 'docmost-redis': 'redis:7-alpine'} +before: seed reads back through the front door = True + phase +0.0s safety-dump | Adatbázis pillanatkép… | err=None + phase +2.1s pulling | Új verzió letöltése… | err=None + phase +3.2s copying | Az adatok másolása a frissítés előtt… | err=None + phase +5.8s converting | Adatbázis átalakítása | err=None + phase +15.8s starting | Indítás az új verzióval… | err=None + phase +26.8s verifying | Működés ellenőrzése… | err=None + phase +42.1s done | Frissítve | err=None +result: {"accepted": true, "http": "202", "duration_s": 42.1, "final_phase": "done", "update_error": null, "hold_reason": null, "state": "running"} +after: PG_VERSION/image=['18', 'postgres:18-alpine@sha256:77f585114c32fbca283dc835b0596f4e52b51b4c6662d7810b2f4084f60a1873'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:18-alpine', 'docmost-redis': 'redis:7-alpine'} conversion_copy={'volume': 'docmost_docmost_postgres_data', 'copy': 'docmost_docmost_postgres_data.pre-update-20260925T111256Z', 'at': '2026-09-25T11:13:34Z', 'from': 16, 'to': 18} +after: seed reads back through the front door = True +controller lines (docmost / conversion): +2026/09/25 11:12:52 update.go:536: [INFO] [stacks] update docmost: accepted — guarded update started +2026/09/25 11:12:52 update.go:1365: [INFO] [stacks] update docmost: phase checking +2026/09/25 11:12:52 update.go:786: [INFO] [stacks] update docmost: ladder — the last step (2 of 2) — the catalog's current definition +2026/09/25 11:12:52 update.go:800: [INFO] [stacks] update docmost: this step CONVERTS docmost-postgres PostgreSQL 16 → 18 (the ladder entry's mark); database volume docmost_docmost_postgres_data +2026/09/25 11:12:52 update.go:814: [INFO] [stacks] update docmost: precondition met — Tier 1 (own recovery unit) copy from 2026-09-25T11:11:16Z (2m0s old, limit 24h0m0s) +2026/09/25 11:12:52 update.go:1365: [INFO] [stacks] update docmost: phase safety-dump +2026/09/25 11:12:53 update.go:846: [INFO] [stacks] update docmost: safety dump done (1 file(s)) [/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260925T111253Z-docmost-postgres.sql] +2026/09/25 11:12:54 undo.go:287: [INFO] [stacks] update docmost: the undo copy will hold 3 named volume(s), 66.4 MiB +2026/09/25 11:12:54 pgconvert.go:286: [INFO] [stacks] update docmost: conversion space — database volume docmost_docmost_postgres_data is 65 MB; the dump needs up to 82 MB + the 2 GB floor; 27.7 GB free beside the stack dir +2026/09/25 11:12:54 update.go:1365: [INFO] [stacks] update docmost: phase pinning +2026/09/25 11:12:54 pin.go:373: [INFO] [stacks] update docmost: pin advanced to /opt/docker/felhom-controller/data/catalog-cache/templates/docmost/docker-compose.yml (docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine) +2026/09/25 11:12:54 update.go:1365: [INFO] [stacks] update docmost: phase pulling +2026/09/25 11:12:55 update.go:1365: [INFO] [stacks] update docmost: phase copying +2026/09/25 11:12:56 update.go:1365: [INFO] [stacks] update docmost: phase copying +2026/09/25 11:12:57 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_postgres_data → docmost_docmost_postgres_data.pre-update-20260925T111256Z in 829ms +2026/09/25 11:12:57 update.go:1365: [INFO] [stacks] update docmost: phase copying +2026/09/25 11:12:58 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_redis_data → docmost_docmost_redis_data.pre-update-20260925T111256Z in 440ms +2026/09/25 11:12:58 update.go:1365: [INFO] [stacks] update docmost: phase copying +2026/09/25 11:12:58 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_storage → docmost_docmost_storage.pre-update-20260925T111256Z in 418ms +2026/09/25 11:12:58 update.go:1365: [INFO] [stacks] update docmost: phase converting +2026/09/25 11:12:58 pgconvert.go:369: [INFO] [stacks] update docmost: CONVERTING docmost-postgres PostgreSQL 16 → 18 (volume docmost_docmost_postgres_data, kept copy docmost_docmost_postgres_data.pre-update-20260925T111256Z) +2026/09/25 11:13:02 pgconvert.go:414: [INFO] [stacks] update docmost: pg_dumpall from 16 done in 566ms — 131 kB, completion line present; before: 2 database(s), 48 table(s), 71 row(s) +2026/09/25 11:13:02 update.go:1365: [INFO] [stacks] update docmost: phase converting +2026/09/25 11:13:03 pgconvert.go:430: [INFO] [stacks] update docmost: database volume docmost_docmost_postgres_data emptied (its copy docmost_docmost_postgres_data.pre-update-20260925T111256Z holds the old datadir, marker checked twice) +2026/09/25 11:13:07 pgconvert.go:459: [INFO] [stacks] update docmost: loaded into 18 in 1.753s (dropped the entrypoint's empty database(s) [docmost]; skipped CREATE ROLE for [docmost]) +2026/09/25 11:13:08 pgconvert.go:475: [INFO] [stacks] update docmost: CONVERTED 16 → 18 in 9.942s — the check is equal (2 database(s), 48 table(s), 71 row(s)), PG_VERSION 18 +2026/09/25 11:13:08 update.go:1365: [INFO] [stacks] update docmost: phase starting +2026/09/25 11:13:19 update.go:1365: [INFO] [stacks] update docmost: phase verifying +2026/09/25 11:13:34 update.go:997: [INFO] [stacks] update docmost: healthy after 15s (the app's health check passed) +2026/09/25 11:13:34 update.go:1015: [INFO] [stacks] update docmost: KEEPING the pre-conversion datadir copy docmost_docmost_postgres_data.pre-update-20260925T111256Z (PostgreSQL 16) until a backup of the converted app is proven +2026/09/25 11:13:34 update.go:1033: [INFO] [stacks] update docmost: DONE in 42s + +undo copies now: docmost_docmost_postgres_data.pre-update-20260925T111256Z +done diff --git a/documentation/audits/night-2026-09-26/D/D1b-backup-for-release.txt b/documentation/audits/night-2026-09-26/D/D1b-backup-for-release.txt new file mode 100644 index 00000000..ae07664e --- /dev/null +++ b/documentation/audits/night-2026-09-26/D/D1b-backup-for-release.txt @@ -0,0 +1,9 @@ +# 'Mentés most' on docmost after the conversion (the release of the kept copy waits for a PROVEN backup) 2026-09-25T13:13:59+0200 +13:13:59 [4] backup press SKIPPED for docmost (R-648: whole-box only; the update's backing-up phase backs up docmost alone) +None +-rw-r--r-- 1 root root 154767 2026-09-25T11:06:21 pre-restore-20260925T110621Z-docmost-postgres.sql +-rw-r--r-- 1 root root 154903 2026-09-25T11:08:02 pre-restore-20260925T110802Z-docmost-postgres.sql +-rw-r--r-- 1 root root 155357 2026-09-25T11:11:16 pre-restore-20260925T111115Z-docmost-postgres.sql +-rw-r--r-- 1 root root 155810 2026-09-25T11:12:53 pre-restore-20260925T111253Z-docmost-postgres.sql +docmost_docmost_postgres_data.pre-update-20260925T111256Z + diff --git a/documentation/audits/night-2026-09-26/D/Da-load-fails.txt b/documentation/audits/night-2026-09-26/D/Da-load-fails.txt new file mode 100644 index 00000000..e8075329 --- /dev/null +++ b/documentation/audits/night-2026-09-26/D/Da-load-fails.txt @@ -0,0 +1,47 @@ +# Da-load-fails — 2026-09-25T13:06:11+0200; controller gitea.dooplex.hu/admin/felhom-controller:0.273.0-rc1 +before: PG_VERSION/image=['16', 'postgres:16-alpine@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} catalog={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:18-alpine', 'docmost-redis': 'redis:7-alpine'} +before: seed reads back through the front door = True + phase +0.0s safety-dump | Adatbázis pillanatkép… | err=None + phase +2.1s pulling | Új verzió letöltése… | err=None + phase +2.7s copying | Az adatok másolása a frissítés előtt… | err=None + phase +5.8s converting | Adatbázis átalakítása | err=None + phase +13.7s undoing | Visszaállítás az előző változatra… | err=None + phase +43.6s undone | Visszaállítva az előző változatra | err=None +result: {"accepted": true, "http": "202", "duration_s": 43.6, "final_phase": "undone", "update_error": null, "hold_reason": null, "state": "running"} +after: PG_VERSION/image=['16', 'postgres:16-alpine@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} conversion_copy=None +after: seed reads back through the front door = True +controller lines (docmost / conversion): +2026/09/25 11:06:21 update.go:536: [INFO] [stacks] update docmost: accepted — guarded update started +2026/09/25 11:06:21 update.go:1365: [INFO] [stacks] update docmost: phase checking +2026/09/25 11:06:21 update.go:786: [INFO] [stacks] update docmost: ladder — the last step (2 of 2) — the catalog's current definition +2026/09/25 11:06:21 update.go:800: [INFO] [stacks] update docmost: this step CONVERTS docmost-postgres PostgreSQL 16 → 18 (the ladder entry's mark); database volume docmost_docmost_postgres_data +2026/09/25 11:06:21 update.go:814: [INFO] [stacks] update docmost: precondition met — Tier 1 (own recovery unit) copy from 2026-09-25T11:04:49Z (2m0s old, limit 24h0m0s) +2026/09/25 11:06:21 update.go:1365: [INFO] [stacks] update docmost: phase safety-dump +2026/09/25 11:06:21 update.go:846: [INFO] [stacks] update docmost: safety dump done (1 file(s)) [/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260925T110621Z-docmost-postgres.sql] +2026/09/25 11:06:22 undo.go:287: [INFO] [stacks] update docmost: the undo copy will hold 3 named volume(s), 66.3 MiB +2026/09/25 11:06:23 pgconvert.go:286: [INFO] [stacks] update docmost: conversion space — database volume docmost_docmost_postgres_data is 65 MB; the dump needs up to 82 MB + the 2 GB floor; 27.7 GB free beside the stack dir +2026/09/25 11:06:23 update.go:1365: [INFO] [stacks] update docmost: phase pinning +2026/09/25 11:06:23 pin.go:373: [INFO] [stacks] update docmost: pin advanced to /opt/docker/felhom-controller/data/catalog-cache/templates/docmost/docker-compose.yml (docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine) +2026/09/25 11:06:23 update.go:1365: [INFO] [stacks] update docmost: phase pulling +2026/09/25 11:06:23 update.go:1365: [INFO] [stacks] update docmost: phase copying +2026/09/25 11:06:24 update.go:1365: [INFO] [stacks] update docmost: phase copying +2026/09/25 11:06:25 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_postgres_data → docmost_docmost_postgres_data.pre-update-20260925T110624Z in 869ms +2026/09/25 11:06:25 update.go:1365: [INFO] [stacks] update docmost: phase copying +2026/09/25 11:06:26 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_redis_data → docmost_docmost_redis_data.pre-update-20260925T110624Z in 431ms +2026/09/25 11:06:26 update.go:1365: [INFO] [stacks] update docmost: phase copying +2026/09/25 11:06:26 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_storage → docmost_docmost_storage.pre-update-20260925T110624Z in 422ms +2026/09/25 11:06:26 update.go:1365: [INFO] [stacks] update docmost: phase converting +2026/09/25 11:06:26 pgconvert.go:369: [INFO] [stacks] update docmost: CONVERTING docmost-postgres PostgreSQL 16 → 18 (volume docmost_docmost_postgres_data, kept copy docmost_docmost_postgres_data.pre-update-20260925T110624Z) +2026/09/25 11:06:30 pgconvert.go:414: [INFO] [stacks] update docmost: pg_dumpall from 16 done in 572ms — 130 kB, completion line present; before: 2 database(s), 48 table(s), 65 row(s) +2026/09/25 11:06:30 update.go:1365: [INFO] [stacks] update docmost: phase converting +2026/09/25 11:06:31 pgconvert.go:430: [INFO] [stacks] update docmost: database volume docmost_docmost_postgres_data emptied (its copy docmost_docmost_postgres_data.pre-update-20260925T110624Z holds the old datadir, marker checked twice) +2026/09/25 11:06:34 update.go:1066: [ERROR] [stacks] update docmost FAILED after the new version was started: conversion failed: loading the dump into 18: exit status 3 (psql: ERROR: extension "adminpack" is not available) +2026/09/25 11:06:34 update.go:1058: [INFO] [stacks] update docmost: kept 11240 bytes of the app's own log at /opt/docker/stacks/docmost/hold-logs/20260925T110634Z/compose-logs.txt before stopping it (R-621) +2026/09/25 11:06:34 update.go:1365: [INFO] [stacks] update docmost: phase undoing +2026/09/25 11:06:34 undo.go:515: [WARN] [stacks] update docmost: UNDO — putting back the previous version and its 3 volume copy(ies) (reason: conversion failed: loading the dump into 18: exit status 3 (psql: ERROR: extension "adminpack" is not available)) +2026/09/25 11:07:04 unattended.go:490: [INFO] [stacks] update docmost: step docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine recorded as undone — the automatic leg will not press it again until the catalog's ladder changes (ladder 1cdde4765ba33137) +2026/09/25 11:07:04 undo.go:640: [INFO] [stacks] update docmost: event app_update_undone +2026/09/25 11:07:04 undo.go:568: [INFO] [stacks] update docmost: UNDONE in 30s — the previous version is running on the data from before the update (the app's health check passed) + +undo copies now: +done diff --git a/documentation/audits/night-2026-09-26/D/Da-setup-adminpack.txt b/documentation/audits/night-2026-09-26/D/Da-setup-adminpack.txt new file mode 100644 index 00000000..71bbd8af --- /dev/null +++ b/documentation/audits/night-2026-09-26/D/Da-setup-adminpack.txt @@ -0,0 +1,6 @@ +# case (a) setup: an object 16 has and 18 does not — the adminpack extension (removed in PostgreSQL 17) +CREATE EXTENSION +adminpack pg_trgm plpgsql unaccent +control: 18 has no adminpack: 0 +# after case (a): adminpack removed again +DROP EXTENSION diff --git a/documentation/audits/night-2026-09-26/D/Db-health-fails.txt b/documentation/audits/night-2026-09-26/D/Db-health-fails.txt new file mode 100644 index 00000000..4dd9eddd --- /dev/null +++ b/documentation/audits/night-2026-09-26/D/Db-health-fails.txt @@ -0,0 +1,53 @@ +# Db-health-fails — 2026-09-25T13:07:52+0200; controller gitea.dooplex.hu/admin/felhom-controller:0.273.0-rc1 +before: PG_VERSION/image=['16', 'postgres:16-alpine@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} catalog={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:18-alpine', 'docmost-redis': 'redis:7-alpine'} +before: seed reads back through the front door = True + phase +0.0s safety-dump | Adatbázis pillanatkép… | err=None + phase +2.1s pulling | Új verzió letöltése… | err=None + phase +2.6s copying | Az adatok másolása a frissítés előtt… | err=None + phase +5.8s converting | Adatbázis átalakítása | err=None + phase +15.2s starting | Indítás az új verzióval… | err=None + phase +26.3s verifying | Működés ellenőrzése… | err=None + phase +117.1s undoing | Visszaállítás az előző változatra… | err=None + phase +148.1s undone | Visszaállítva az előző változatra | err=None +result: {"accepted": true, "http": "202", "duration_s": 148.1, "final_phase": "undone", "update_error": null, "hold_reason": null, "state": "running"} +after: PG_VERSION/image=['16', 'postgres:16-alpine@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} conversion_copy=None +after: seed reads back through the front door = True +controller lines (docmost / conversion): +2026/09/25 11:08:02 update.go:536: [INFO] [stacks] update docmost: accepted — guarded update started +2026/09/25 11:08:02 update.go:1365: [INFO] [stacks] update docmost: phase checking +2026/09/25 11:08:02 update.go:786: [INFO] [stacks] update docmost: ladder — the last step (2 of 2) — the catalog's current definition +2026/09/25 11:08:02 update.go:800: [INFO] [stacks] update docmost: this step CONVERTS docmost-postgres PostgreSQL 16 → 18 (the ladder entry's mark); database volume docmost_docmost_postgres_data +2026/09/25 11:08:02 update.go:814: [INFO] [stacks] update docmost: precondition met — Tier 1 (own recovery unit) copy from 2026-09-25T11:06:21Z (2m0s old, limit 24h0m0s) +2026/09/25 11:08:02 update.go:1365: [INFO] [stacks] update docmost: phase safety-dump +2026/09/25 11:08:02 update.go:846: [INFO] [stacks] update docmost: safety dump done (1 file(s)) [/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260925T110802Z-docmost-postgres.sql] +2026/09/25 11:08:03 undo.go:287: [INFO] [stacks] update docmost: the undo copy will hold 3 named volume(s), 66.3 MiB +2026/09/25 11:08:04 pgconvert.go:286: [INFO] [stacks] update docmost: conversion space — database volume docmost_docmost_postgres_data is 65 MB; the dump needs up to 82 MB + the 2 GB floor; 27.7 GB free beside the stack dir +2026/09/25 11:08:04 update.go:1365: [INFO] [stacks] update docmost: phase pinning +2026/09/25 11:08:04 pin.go:373: [INFO] [stacks] update docmost: pin advanced to /opt/docker/felhom-controller/data/catalog-cache/templates/docmost/docker-compose.yml (docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine) +2026/09/25 11:08:04 update.go:1365: [INFO] [stacks] update docmost: phase pulling +2026/09/25 11:08:04 update.go:1365: [INFO] [stacks] update docmost: phase copying +2026/09/25 11:08:05 update.go:1365: [INFO] [stacks] update docmost: phase copying +2026/09/25 11:08:06 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_postgres_data → docmost_docmost_postgres_data.pre-update-20260925T110805Z in 839ms +2026/09/25 11:08:06 update.go:1365: [INFO] [stacks] update docmost: phase copying +2026/09/25 11:08:07 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_redis_data → docmost_docmost_redis_data.pre-update-20260925T110805Z in 435ms +2026/09/25 11:08:07 update.go:1365: [INFO] [stacks] update docmost: phase copying +2026/09/25 11:08:07 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_storage → docmost_docmost_storage.pre-update-20260925T110805Z in 441ms +2026/09/25 11:08:07 update.go:1365: [INFO] [stacks] update docmost: phase converting +2026/09/25 11:08:07 pgconvert.go:369: [INFO] [stacks] update docmost: CONVERTING docmost-postgres PostgreSQL 16 → 18 (volume docmost_docmost_postgres_data, kept copy docmost_docmost_postgres_data.pre-update-20260925T110805Z) +2026/09/25 11:08:10 pgconvert.go:414: [INFO] [stacks] update docmost: pg_dumpall from 16 done in 566ms — 130 kB, completion line present; before: 2 database(s), 48 table(s), 67 row(s) +2026/09/25 11:08:11 update.go:1365: [INFO] [stacks] update docmost: phase converting +2026/09/25 11:08:12 pgconvert.go:430: [INFO] [stacks] update docmost: database volume docmost_docmost_postgres_data emptied (its copy docmost_docmost_postgres_data.pre-update-20260925T110805Z holds the old datadir, marker checked twice) +2026/09/25 11:08:16 pgconvert.go:459: [INFO] [stacks] update docmost: loaded into 18 in 1.747s (dropped the entrypoint's empty database(s) [docmost]; skipped CREATE ROLE for [docmost]) +2026/09/25 11:08:17 pgconvert.go:475: [INFO] [stacks] update docmost: CONVERTED 16 → 18 in 9.866s — the check is equal (2 database(s), 48 table(s), 67 row(s)), PG_VERSION 18 +2026/09/25 11:08:17 update.go:1365: [INFO] [stacks] update docmost: phase starting +2026/09/25 11:08:28 update.go:1365: [INFO] [stacks] update docmost: phase verifying +2026/09/25 11:09:58 update.go:1066: [ERROR] [stacks] update docmost FAILED after the new version was started: not healthy: not healthy within 1m30s (last: health check failing) +2026/09/25 11:09:59 update.go:1058: [INFO] [stacks] update docmost: kept 13077 bytes of the app's own log at /opt/docker/stacks/docmost/hold-logs/20260925T110958Z/compose-logs.txt before stopping it (R-621) +2026/09/25 11:09:59 update.go:1365: [INFO] [stacks] update docmost: phase undoing +2026/09/25 11:09:59 undo.go:515: [WARN] [stacks] update docmost: UNDO — putting back the previous version and its 3 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: health check failing)) +2026/09/25 11:10:30 unattended.go:490: [INFO] [stacks] update docmost: step docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine recorded as undone — the automatic leg will not press it again until the catalog's ladder changes (ladder 1cdde4765ba33137) +2026/09/25 11:10:30 undo.go:640: [INFO] [stacks] update docmost: event app_update_undone +2026/09/25 11:10:30 undo.go:568: [INFO] [stacks] update docmost: UNDONE in 31s — the previous version is running on the data from before the update (the app's health check passed) + +undo copies now: +done diff --git a/documentation/audits/night-2026-09-26/D/Dc-restart-mid-convert.txt b/documentation/audits/night-2026-09-26/D/Dc-restart-mid-convert.txt new file mode 100644 index 00000000..ddc7bf35 --- /dev/null +++ b/documentation/audits/night-2026-09-26/D/Dc-restart-mid-convert.txt @@ -0,0 +1,83 @@ +# Dc-restart-mid-convert — 2026-09-25T13:11:04+0200; controller gitea.dooplex.hu/admin/felhom-controller:0.273.0-rc1 +before: PG_VERSION/image=['16', 'postgres:16-alpine@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} catalog={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:18-alpine', 'docmost-redis': 'redis:7-alpine'} +before: seed reads back through the front door = True + [kill] 2026/09/25 11:11:23 pgconvert.go:430: [INFO] [stacks] update docmost: database volume docmost_docmost_postgres_data emptied (its copy docmost_docmost_postgres_data.pre-update-20260925T111119Z holds the old datadir, marker checked twice) | KILLED controller pid 554023 at 11:11:24.114456007 + [kill] waiting for the controller to come back and the journal to be resumed + phase +0.0s safety-dump | Adatbázis pillanatkép… | err=None + phase +2.1s pulling | Új verzió letöltése… | err=None + phase +2.6s copying | Az adatok másolása a frissítés előtt… | err=None + phase +5.8s converting | Adatbázis átalakítása | err=None + phase +8.4s None | None | err=None +result: {"accepted": true, "http": "202", "duration_s": 8.5, "final_phase": null, "update_error": null, "hold_reason": null, "state": null, "after_restart": {"phase": "undone", "error": null, "hold": null}} +after: PG_VERSION/image=['16', 'postgres:16-alpine@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} conversion_copy=None +after: seed reads back through the front door = True +controller lines (docmost / conversion): +2026/09/25 11:11:15 update.go:536: [INFO] [stacks] update docmost: accepted — guarded update started +2026/09/25 11:11:15 update.go:1365: [INFO] [stacks] update docmost: phase checking +2026/09/25 11:11:15 update.go:786: [INFO] [stacks] update docmost: ladder — the last step (2 of 2) — the catalog's current definition +2026/09/25 11:11:15 update.go:800: [INFO] [stacks] update docmost: this step CONVERTS docmost-postgres PostgreSQL 16 → 18 (the ladder entry's mark); database volume docmost_docmost_postgres_data +2026/09/25 11:11:15 update.go:814: [INFO] [stacks] update docmost: precondition met — Tier 1 (own recovery unit) copy from 2026-09-25T11:08:02Z (3m0s old, limit 24h0m0s) +2026/09/25 11:11:15 update.go:1365: [INFO] [stacks] update docmost: phase safety-dump +2026/09/25 11:11:16 update.go:846: [INFO] [stacks] update docmost: safety dump done (1 file(s)) [/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260925T111115Z-docmost-postgres.sql] +2026/09/25 11:11:17 undo.go:287: [INFO] [stacks] update docmost: the undo copy will hold 3 named volume(s), 66.4 MiB +2026/09/25 11:11:17 pgconvert.go:286: [INFO] [stacks] update docmost: conversion space — database volume docmost_docmost_postgres_data is 65 MB; the dump needs up to 82 MB + the 2 GB floor; 27.7 GB free beside the stack dir +2026/09/25 11:11:17 update.go:1365: [INFO] [stacks] update docmost: phase pinning +2026/09/25 11:11:17 pin.go:373: [INFO] [stacks] update docmost: pin advanced to /opt/docker/felhom-controller/data/catalog-cache/templates/docmost/docker-compose.yml (docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine) +2026/09/25 11:11:17 update.go:1365: [INFO] [stacks] update docmost: phase pulling +2026/09/25 11:11:18 update.go:1365: [INFO] [stacks] update docmost: phase copying +2026/09/25 11:11:19 update.go:1365: [INFO] [stacks] update docmost: phase copying +2026/09/25 11:11:20 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_postgres_data → docmost_docmost_postgres_data.pre-update-20260925T111119Z in 794ms +2026/09/25 11:11:20 update.go:1365: [INFO] [stacks] update docmost: phase copying +2026/09/25 11:11:20 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_redis_data → docmost_docmost_redis_data.pre-update-20260925T111119Z in 415ms +2026/09/25 11:11:20 update.go:1365: [INFO] [stacks] update docmost: phase copying +2026/09/25 11:11:21 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_storage → docmost_docmost_storage.pre-update-20260925T111119Z in 423ms +2026/09/25 11:11:21 update.go:1365: [INFO] [stacks] update docmost: phase converting +2026/09/25 11:11:21 pgconvert.go:369: [INFO] [stacks] update docmost: CONVERTING docmost-postgres PostgreSQL 16 → 18 (volume docmost_docmost_postgres_data, kept copy docmost_docmost_postgres_data.pre-update-20260925T111119Z) +2026/09/25 11:11:22 pgconvert.go:414: [INFO] [stacks] update docmost: pg_dumpall from 16 done in 539ms — 130 kB, completion line present; before: 2 database(s), 48 table(s), 69 row(s) +2026/09/25 11:11:23 update.go:1365: [INFO] [stacks] update docmost: phase converting +2026/09/25 11:11:23 pgconvert.go:430: [INFO] [stacks] update docmost: database volume docmost_docmost_postgres_data emptied (its copy docmost_docmost_postgres_data.pre-update-20260925T111119Z holds the old datadir, marker checked twice) +2026/09/25 11:11:24 update.go:1432: [WARN] [stacks] update recovery: docmost was interrupted while UNDOING (started 2026-09-25T11:11:15Z) — marking it Updating and RESUMING the undo +2026/09/25 11:11:24 main.go:496: [WARN] [update] 1 interrupted update(s) will resume after the backup side is wired: [docmost] +2026/09/25 11:11:24 main.go:596: [WARN] [update] resumed 1 interrupted update(s) +2026/09/25 11:11:24 update.go:1489: [INFO] [stacks] update docmost: resuming the UNDO after a controller restart (was converting) +2026/09/25 11:11:24 update.go:1066: [ERROR] [stacks] update docmost FAILED after the new version was started: the controller restarted during the database conversion — undoing it +2026/09/25 11:11:25 update.go:1058: [INFO] [stacks] update docmost: kept 5297 bytes of the app's own log at /opt/docker/stacks/docmost/hold-logs/20260925T111124Z/compose-logs.txt before stopping it (R-621) +2026/09/25 11:11:25 update.go:1365: [INFO] [stacks] update docmost: phase undoing +2026/09/25 11:11:25 undo.go:515: [WARN] [stacks] update docmost: UNDO — putting back the previous version and its 3 volume copy(ies) (reason: the controller restarted during the database conversion — undoing it) +2026/09/25 11:12:05 unattended.go:490: [INFO] [stacks] update docmost: step docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine recorded as undone — the automatic leg will not press it again until the catalog's ladder changes (ladder 1cdde4765ba33137) +2026/09/25 11:12:05 undo.go:640: [INFO] [stacks] update docmost: event app_update_undone +2026/09/25 11:12:05 undo.go:568: [INFO] [stacks] update docmost: UNDONE in 40s — the previous version is running on the data from before the update (the app's health check passed) + +undo copies now: +done +# the whole restart, from the controller's own log +2026/09/25 11:11:20 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_postgres_data → docmost_docmost_postgres_data.pre-update-20260925T111119Z in 794ms +2026/09/25 11:11:20 update.go:1365: [INFO] [stacks] update docmost: phase copying +2026/09/25 11:11:20 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_redis_data → docmost_docmost_redis_data.pre-update-20260925T111119Z in 415ms +2026/09/25 11:11:20 update.go:1365: [INFO] [stacks] update docmost: phase copying +2026/09/25 11:11:21 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_storage → docmost_docmost_storage.pre-update-20260925T111119Z in 423ms +2026/09/25 11:11:21 update.go:1365: [INFO] [stacks] update docmost: phase converting +2026/09/25 11:11:21 pgconvert.go:369: [INFO] [stacks] update docmost: CONVERTING docmost-postgres PostgreSQL 16 → 18 (volume docmost_docmost_postgres_data, kept copy docmost_docmost_postgres_data.pre-update-20260925T111119Z) +2026/09/25 11:11:22 pgconvert.go:414: [INFO] [stacks] update docmost: pg_dumpall from 16 done in 539ms — 130 kB, completion line present; before: 2 database(s), 48 table(s), 69 row(s) +2026/09/25 11:11:23 update.go:1365: [INFO] [stacks] update docmost: phase converting +2026/09/25 11:11:23 pgconvert.go:430: [INFO] [stacks] update docmost: database volume docmost_docmost_postgres_data emptied (its copy docmost_docmost_postgres_data.pre-update-20260925T111119Z holds the old datadir, marker checked twice) +2026/09/25 11:11:24 main.go:336: [INFO] felhom-controller 0.273.0-rc1 starting (customer: demo-hp, domain: enkisfelhom.hu) +2026/09/25 11:11:24 update.go:1432: [WARN] [stacks] update recovery: docmost was interrupted while UNDOING (started 2026-09-25T11:11:15Z) — marking it Updating and RESUMING the undo +2026/09/25 11:11:24 main.go:496: [WARN] [update] 1 interrupted update(s) will resume after the backup side is wired: [docmost] +2026/09/25 11:11:24 main.go:596: [WARN] [update] resumed 1 interrupted update(s) +2026/09/25 11:11:24 update.go:1489: [INFO] [stacks] update docmost: resuming the UNDO after a controller restart (was converting) +2026/09/25 11:11:24 update.go:1066: [ERROR] [stacks] update docmost FAILED after the new version was started: the controller restarted during the database conversion — undoing it +2026/09/25 11:11:25 update.go:1058: [INFO] [stacks] update docmost: kept 5297 bytes of the app's own log at /opt/docker/stacks/docmost/hold-logs/20260925T111124Z/compose-logs.txt before stopping it (R-621) +2026/09/25 11:11:25 update.go:1365: [INFO] [stacks] update docmost: phase undoing +2026/09/25 11:11:25 undo.go:515: [WARN] [stacks] update docmost: UNDO — putting back the previous version and its 3 volume copy(ies) (reason: the controller restarted during the database conversion — undoing it) +2026/09/25 11:11:25 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/docmost/docker-compose.yml +2026/09/25 11:11:25 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/docmost/docker-compose.yml +2026/09/25 11:11:25 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/docmost/docker-compose.yml +2026/09/25 11:11:25 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/docmost/docker-compose.yml +2026/09/25 11:11:25 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/docmost/docker-compose.yml +2026/09/25 11:11:34 healthprobe.go:190: [WARN] Health probe docmost: HTTP GET :3000/ → Get "http://docmost-postgres:3000/": dial tcp: lookup docmost-postgres on 127.0.0.11:53: no such host +2026/09/25 11:11:38 [INFO] [stacks] SaveAppConfig: saved config for docmost +2026/09/25 11:11:38 pin.go:93: [INFO] [stacks] pin docmost: docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:16-alpine, docmost-redis=redis:7-alpine +2026/09/25 11:11:44 healthprobe.go:190: [WARN] Health probe docmost: HTTP GET :3000/ → Get "http://docmost-postgres:3000/": dial tcp: lookup docmost-postgres on 127.0.0.11:53: no such host +2026/09/25 11:12:05 [INFO] [stacks] SaveAppConfig: saved config for docmost +2026/09/25 11:12:05 unattended.go:490: [INFO] [stacks] update docmost: step docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine recorded as undone — the automatic leg will not press it again until the catalog's ladder changes (ladder 1cdde4765ba33137) diff --git a/documentation/audits/night-2026-09-26/D/box-verdict-docmost.json b/documentation/audits/night-2026-09-26/D/box-verdict-docmost.json new file mode 100644 index 00000000..5205bf55 --- /dev/null +++ b/documentation/audits/night-2026-09-26/D/box-verdict-docmost.json @@ -0,0 +1,29 @@ +{ + "app": "docmost", + "venue": "box 9202 (drill catalog, controller 0.273.0-rc1), the product's guarded Update", + "from": { + "docmost": "docmost/docmost:0.96.0", + "docmost-postgres": "postgres:16-alpine", + "docmost-redis": "redis:7-alpine" + }, + "to": { + "docmost": "docmost/docmost:0.96.0", + "docmost-postgres": "postgres:18-alpine", + "docmost-redis": "redis:7-alpine" + }, + "verdict": "proven", + "seed_read_before": true, + "seed_read_after": true, + "healthy_after": true, + "engine_conversion": { + "service": "docmost-postgres", + "engine": "postgres", + "from": 16, + "to": 18, + "result": "converted", + "controller_line": "CONVERTED 16 → 18 in 9.942s — the check is equal (2 database(s), 48 table(s), 71 row(s)), PG_VERSION 18" + }, + "duration_s": 42.1, + "measured_at": "2026-09-25T11:13:34Z", + "evidence": "felhom.eu/documentation/audits/night-2026-09-26/D/D1-happy-path.txt" +} \ No newline at end of file diff --git a/documentation/audits/night-2026-09-26/E/E1-01-window-to-1244.txt b/documentation/audits/night-2026-09-26/E/E1-01-window-to-1244.txt new file mode 100644 index 00000000..3ba6da0d --- /dev/null +++ b/documentation/audits/night-2026-09-26/E/E1-01-window-to-1244.txt @@ -0,0 +1,12 @@ +12:40:28 before: +2026/09/25 09:31:03 scheduler.go:132: [INFO] [scheduler] Daily job offsite-proof scheduled for 2026-09-26 05:30 CEST +2026/09/25 09:31:03 scheduler.go:132: [INFO] [scheduler] Daily job metrics-prune scheduled for 2026-09-26 04:00 CEST +2026/09/25 09:31:03 scheduler.go:132: [INFO] [scheduler] Daily job fill-watch scheduled for 2026-09-26 03:30 CEST +12:40:28 POST /backups/window window_start=12:44 -> 303 +12:40:34 after: +2026/09/25 10:40:28 auth.go:142: [DEBUG] [web] auth: valid session for POST /backups/window +2026/09/25 10:40:28 server.go:545: [DEBUG] [web] ServeHTTP: POST /backups/window from 172.18.0.6:60864 +2026/09/25 10:40:28 scheduler.go:172: [INFO] [scheduler] Daily job db-dump rescheduled 02:30 → 12:44 (next run 2026-09-25 12:44 CEST) +2026/09/25 10:40:28 scheduler.go:172: [INFO] [scheduler] Daily job tier2-backup rescheduled 03:30 → 13:44 (next run 2026-09-25 13:44 CEST) +2026/09/25 10:40:28 scheduler.go:172: [INFO] [scheduler] Daily job offbox-backup rescheduled 04:15 → 14:29 (next run 2026-09-25 14:29 CEST) +2026/09/25 10:40:28 backup_handlers.go:65: [INFO] [web] backup window set to 12:44 (legs 12:44/13:44/14:29) diff --git a/documentation/audits/night-2026-09-26/E/E1-02-dbdump-leg.txt b/documentation/audits/night-2026-09-26/E/E1-02-dbdump-leg.txt new file mode 100644 index 00000000..5496a650 --- /dev/null +++ b/documentation/audits/night-2026-09-26/E/E1-02-dbdump-leg.txt @@ -0,0 +1,68 @@ +Fri Sep 25 10:40:56 UTC 2026 +2026/09/25 10:38:19 router.go:946: [ERROR] [api] Remove failed for nextcloud: A(z) /mnt/felhom-drives/scratch_hdd/userdata/nextcloud tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó viss +2026/09/25 10:38:19 router.go:907: [INFO] [api] Remove requested for stack: nextcloud +2026/09/25 10:38:19 delete.go:597: [INFO] Removing deployed stack: nextcloud (removeHDDData=false, hddDeclared=true, backupPaths=2) +2026/09/25 10:38:19 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml +2026/09/25 10:38:47 delete.go:650: [INFO] [stacks] RemoveStack nextcloud: removed volume(s) [nextcloud_nextcloud_db_data nextcloud_nextcloud_html nextcloud_nextcloud_redis_data] +2026/09/25 10:38:47 delete.go:749: [INFO] Stack nextcloud removed successfully (took 27.4s) +2026/09/25 10:39:01 router.go:431: [INFO] [api] Deploy requested for stack: nextcloud +2026/09/25 10:39:01 [INFO] [stacks] SaveAppConfig: saved config for nextcloud +2026/09/25 10:39:01 deploy.go:391: [INFO] [stacks] Deploying stack nextcloud with 7 env vars: [DOMAIN, SUBDOMAIN, DB_PASSWORD, MYSQL_ROOT_PASSWORD, NEXTCLOUD_ADMIN_USER, NEXTCLOUD_ADMIN_PASSWORD, HDD_PATH] +2026/09/25 10:39:01 manager.go:1592: [INFO] [stacks] Deploying stack nextcloud — checking 3 images... +2026/09/25 10:39:13 deploy.go:481: [INFO] [stacks] Stack nextcloud deployed successfully (took 11.2s) +2026/09/25 10:39:13 [INFO] [stacks] SaveAppConfig: saved config for nextcloud +2026/09/25 10:39:13 [INFO] [stacks] SaveAppConfig: saved config for nextcloud +2026/09/25 10:39:13 installed.go:420: [INFO] [stacks] installed-images nextcloud: recorded 3 service(s) (nextcloud=nextcloud:34.0.4-apache (sha256:a5ace30c695a…), nextcloud-db=mariadb:12.3 (sha256:805c8e104bd5…), nextcloud-red +2026/09/25 10:39:13 [INFO] [stacks] SaveAppConfig: saved config for nextcloud +2026/09/25 10:39:13 pin.go:93: [INFO] [stacks] pin nextcloud: nextcloud=nextcloud:34.0.4-apache, nextcloud-db=mariadb:12.3, nextcloud-redis=redis:7-alpine +2026/09/25 10:39:16 manager.go:1553: [INFO] [stacks] Stack nextcloud post-start status: +2026/09/25 10:39:16 manager.go:1556: [INFO] [stacks] nextcloud nextcloud:34.0.4-apache@sha256:a5ace30c695afe48c2c406e940ee7886a81e13fa382e57cd68b2416d1a66914c running Up 3 seconds (health: starting) +2026/09/25 10:39:16 manager.go:1556: [INFO] [stacks] nextcloud-db mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1 running Up 14 seconds (healthy) +2026/09/25 10:39:16 manager.go:1556: [INFO] [stacks] nextcloud-redis redis:7-alpine@sha256:858f009f9709ce576febc734aa78b8f6d624b82571f9ddb6bda4377c833b3499 running Up 14 seconds (healthy) +2026/09/25 10:40:28 scheduler.go:172: [INFO] [scheduler] Daily job db-dump rescheduled 02:30 → 12:44 (next run 2026-09-25 12:44 CEST) +66548367016e nextcloud-db nextcloud mariadb:12.3 +5435a635557b nextcloud-redis nextcloud redis:7-alpine +2026/09/25 10:40:28 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml +2026/09/25 10:40:28 recovery_unit.go:235: [INFO] [backup] Recovery unit captured for nextcloud → /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud (images=3, secrets-referenced=3, data_keys=0, portable-carried=2/2, withhe +/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud: +total 16 +drwxr-xr-x 3 root root 4096 10:40:28 . +drwxr-xr-x 3 root root 4096 10:40:28 .. +drwxr-xr-x 2 root root 4096 10:40:28 compose +-rw-r--r-- 1 root root 1439 10:40:28 manifest.json +## after the db-dump leg +2026/09/25 10:44:23 backup.go:824: [INFO] [backup] Volume dump: nextcloud/nextcloud_nextcloud_db_data → 164.5 MB +2026/09/25 10:44:28 backup.go:824: [INFO] [backup] Volume dump: nextcloud/nextcloud_nextcloud_html → 761.6 MB +2026/09/25 10:44:28 backup.go:824: [INFO] [backup] Volume dump: nextcloud/nextcloud_nextcloud_redis_data → 187.5 KB +2026/09/25 10:44:28 backup.go:933: [INFO] [backup] Restarting nextcloud after volume dump +2026/09/25 10:44:28 manager.go:1207: [INFO] [stacks] Starting stack: nextcloud +2026/09/25 10:44:35 manager.go:1222: [INFO] [stacks] Stack nextcloud started successfully (took 6.2s) +2026/09/25 10:44:38 manager.go:1553: [INFO] [stacks] Stack nextcloud post-start status: +2026/09/25 10:44:38 manager.go:1556: [INFO] [stacks] nextcloud nextcloud:34.0.4-apache@sha256:a5ace30c695afe48c2c406e940ee7886a81e13fa382e57cd68b2416d1a66914c running Up 3 seconds (healt +2026/09/25 10:44:38 manager.go:1556: [INFO] [stacks] nextcloud-db mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1 running Up 9 seconds (healt +2026/09/25 10:44:38 manager.go:1556: [INFO] [stacks] nextcloud-redis redis:7-alpine@sha256:858f009f9709ce576febc734aa78b8f6d624b82571f9ddb6bda4377c833b3499 running Up 9 seconds (healt +2026/09/25 10:44:56 recovery_unit.go:235: [INFO] [backup] Recovery unit captured for nextcloud → /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud (images=3, secrets-referenced=3, data_keys=0 +2026/09/25 10:44:56 scheduler.go:363: [INFO] [scheduler] Job db-dump completed (took 56.675s) +-rw-r--r-- 1 root root 1587 10:44:56 /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud/manifest.json + +/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud/compose: +total 32 +drwxr-xr-x 2 root root 4096 10:44:56 . +drwxr-xr-x 5 root root 4096 10:44:56 .. +-rw-r--r-- 1 root root 9180 10:44:56 .felhom.yml +-rw------- 1 root root 611 10:44:56 app.yaml +-rw-r--r-- 1 root root 4809 10:44:56 docker-compose.yml + +/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud/db-dumps: +total 864 +drwxr-xr-x 2 root root 4096 10:44:00 . +drwxr-xr-x 5 root root 4096 10:44:56 .. +-rw-r--r-- 1 root root 873120 10:44:00 nextcloud-mariadb.sql + +/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud/volume-dumps: +total 948440 +drwxr-xr-x 2 root root 4096 10:44:28 . +drwxr-xr-x 5 root root 4096 10:44:56 .. +-rw-r--r-- 1 root root 172453376 10:44:22 nextcloud_nextcloud_db_data.tar +-rw-r--r-- 1 root root 798543360 10:44:26 nextcloud_nextcloud_html.tar +-rw-r--r-- 1 root root 192000 10:44:28 nextcloud_nextcloud_redis_data.tar diff --git a/documentation/audits/night-2026-09-26/E/E1-03-window-back.txt b/documentation/audits/night-2026-09-26/E/E1-03-window-back.txt new file mode 100644 index 00000000..52fe52db --- /dev/null +++ b/documentation/audits/night-2026-09-26/E/E1-03-window-back.txt @@ -0,0 +1,12 @@ +12:45:23 before: +2026/09/25 10:40:28 scheduler.go:172: [INFO] [scheduler] Daily job db-dump rescheduled 02:30 → 12:44 (next run 2026-09-25 12:44 CEST) +2026/09/25 10:40:28 scheduler.go:172: [INFO] [scheduler] Daily job tier2-backup rescheduled 03:30 → 13:44 (next run 2026-09-25 13:44 CEST) +2026/09/25 10:40:28 scheduler.go:172: [INFO] [scheduler] Daily job offbox-backup rescheduled 04:15 → 14:29 (next run 2026-09-25 14:29 CEST) +12:45:23 POST /backups/window window_start=02:30 -> 303 +12:45:29 after: +2026/09/25 10:45:23 auth.go:142: [DEBUG] [web] auth: valid session for POST /backups/window +2026/09/25 10:45:23 server.go:545: [DEBUG] [web] ServeHTTP: POST /backups/window from 172.18.0.6:39110 +2026/09/25 10:45:23 scheduler.go:172: [INFO] [scheduler] Daily job db-dump rescheduled 12:44 → 02:30 (next run 2026-09-26 02:30 CEST) +2026/09/25 10:45:23 scheduler.go:172: [INFO] [scheduler] Daily job tier2-backup rescheduled 13:44 → 03:30 (next run 2026-09-26 03:30 CEST) +2026/09/25 10:45:23 scheduler.go:172: [INFO] [scheduler] Daily job offbox-backup rescheduled 14:29 → 04:15 (next run 2026-09-26 04:15 CEST) +2026/09/25 10:45:23 backup_handlers.go:65: [INFO] [web] backup window set to 02:30 (legs 02:30/03:30/04:15) diff --git a/documentation/audits/night-2026-09-26/E/E1-04-remove-keep.txt b/documentation/audits/night-2026-09-26/E/E1-04-remove-keep.txt new file mode 100644 index 00000000..142a96e7 --- /dev/null +++ b/documentation/audits/night-2026-09-26/E/E1-04-remove-keep.txt @@ -0,0 +1,37 @@ +('200', {'ok': True, 'message': 'Stack nextcloud stop completed'}) +remove keep data + keep backups -> 200 {'ok': True, 'data': {'removed': 'nextcloud', 'volumes_removed': ['nextcloud_nextcloud_db_data', 'nextcloud_nextcloud_html', 'nextcloud_nextcloud_redis_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': ['/mnt/felhom-drives/scratch_hdd/appdata/nextcloud (126M)'], 'verified': True}, 'message': 'Stack nextcloud removed'} +total 28 +drwxrwx--- 5 www-data www-data 4096 Sep 25 10:39 . +drwxr-xr-x 4 root root 4096 Sep 25 10:39 .. +-rw-rw-r-- 1 www-data www-data 542 Sep 25 10:39 .htaccess +-rw-rw-r-- 1 www-data www-data 52 Sep 25 10:39 .ncdata +drwxr-xr-x 3 www-data www-data 4096 Sep 25 10:39 admin +drwxr-xr-x 4 www-data www-data 4096 Sep 25 10:39 appdata_ocrpy9goog25 +drwxr-xr-x 4 www-data www-data 4096 Sep 25 10:40 drilld112bd +-rw-rw-r-- 1 www-data www-data 0 Sep 25 10:39 index.html +-rw-rw-r-- 1 www-data www-data 0 Sep 25 10:39 nextcloud.log +/mnt/felhom-drives/scratch_hdd/appdata/nextcloud/admin/files: +Documents +Nextcloud Manual.pdf +Nextcloud intro.mp4 +Nextcloud.png +Photos +Readme.md +Reasons to use Nextcloud.pdf +Templates +Templates credits.md + +/mnt/felhom-drives/scratch_hdd/appdata/nextcloud/drilld112bd/files: +Documents +Nextcloud Manual.pdf +Nextcloud intro.mp4 +Nextcloud.png +Photos +Readme.md +Reasons to use Nextcloud.pdf +Templates +Templates credits.md +after-backup.txt +before-backup.txt +nextcloud + diff --git a/documentation/audits/night-2026-09-26/E/E1-05-restore-removed.txt b/documentation/audits/night-2026-09-26/E/E1-05-restore-removed.txt new file mode 100644 index 00000000..e6a131be --- /dev/null +++ b/documentation/audits/night-2026-09-26/E/E1-05-restore-removed.txt @@ -0,0 +1,62 @@ +# E1 Q1 — the removed-app row on /backups/restore (R-487) +row present: True +csolat kikapcsolva — a központi monitoring nem aktív Rendszermonitor → Biztonsági mentés — Visszaállítás enkisfelhom.hu Visszaállítás Alkalmazás: — Válasszon — Docmost Nextcloud Paperless-ngx