docs: 09 decisions 35-38 + part 10 as built; R-689 filed, R-469 narrowed; night-2026-09-26 evidence (A, B, C, D, F)
gates / gates (push) Successful in 24s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-25 13:27:26 +02:00
parent ccd915ff34
commit 9141817867
51 changed files with 2980 additions and 2 deletions
@@ -402,6 +402,46 @@ R-636's louder repeated alarm.
behaviour shipped in v0.271.0 is now the rule. **Why:** one night's delay for an app is cheap, and a resume
would be one more mechanism acting with nobody watching.
### 2026-09-25 (evening) — two operator rulings
35. **Part 10 is built now, one app first** (operator ruling 2026-09-25 evening, D1 option A). The box
converts a PostgreSQL major as a step of the guarded update. **docmost is the first app.** Each other app
needs its own proof on both venues (the test bench and a scratch box through the real guarded Update)
before the catalog may move it. The engine-major gate stays for every app without that proof. **Why:**
decision 16 said "each of the eleven apps is proven before the catalog may move it"; one app first makes
the first proof small enough to read. As built: §6.4 part 10.
36. **A reinstall over kept data offers a choice** (operator ruling 2026-09-25 evening, D2 option A): "use
my kept data" (the database from a backup, with the kept files) or "start fresh" (the kept files move to
a dated folder; nothing is deleted). **And, from the operator's follow-up question: kept data must never
be a dead end.** The household can see it (read only), load it into a later install, and delete it.
Support is for the rare case, not the normal one. **Not ruled (D3, in `STATUS.md`):** whether the box
ever deletes kept data by itself. Until it is ruled, **nothing deletes kept data automatically** — only a
household's explicit Delete. Where it lives and who deletes it: `07-backup-architecture.md` §"Kept data".
### 2026-09-25 (evening) — two decisions taken by CC unattended, building part 10
37. **docmost converts PostgreSQL 16 → 18, not 16 → 17** — *decided by CC unattended 2026-09-25 — operator may
reverse.* *One sentence:* which major does the first conversion target? **Options:** (a) 17 — no mount
change; (b) 18 — the data volume's mount moves to `/var/lib/postgresql` in the same step. **Costs:** (a) every
box converts again later, and docmost's own upstream compose already ships `postgres:18` at
`/var/lib/postgresql`; (b) the step changes the mount point — measured: `postgres:18` REFUSES (exit 1) even an
EMPTY volume at `/var/lib/postgresql/data` — so the step's definition carries the mount, and the undo must put
the old mount back with the old bytes, which it does (it restores the old definition and every volume).
**Why (b):** one conversion is better than two when the app's upstream runs the newer major. Reversible: the
catalog could pin 17 instead before any customer box moves. `audits/night-2026-09-26/A/README.md` A1.
38. **The box loads from a `pg_dumpall` of the OLD engine, with no error tolerated** — *decided by CC unattended
2026-09-25 — operator may reverse.* *One sentence:* what does the conversion load into the new engine?
**Options:** (a) the existing safety dump (`pg_dump` per database, `--no-owner --no-privileges`); (b) a
`pg_dumpall` taken from the old engine at conversion time, loaded with `ON_ERROR_STOP`. **Costs:** (a) is free,
and silently drops any role, grant or database setting beyond the bootstrap ones (for docmost both routes gave
identical databases — it has one role and one database; the other ten are not measured); (b) one more dump
(0.6 s for docmost) and a load that must step around exactly two objects the new engine's entrypoint makes (the
bootstrap role and database — measured: loaded raw, exactly two `already exists` errors). **Why (b):** decision
16 says "save everything"; the two collisions are removed precisely (the entrypoint's databases dropped only
when they hold no table; `CREATE ROLE <x>;` skipped only for a role that exists — its `ALTER ROLE … PASSWORD`
still runs), so ANY other error stops the load and the undo runs. Proven live: an extension 18 lacks
(`adminpack`) stopped the load and the box undid it. `audits/night-2026-09-26/A/README.md` A2.
---
## 3b. ANSWERED 2026-09-23 — the seven questions Slices 6 and 7 needed
@@ -1175,7 +1215,7 @@ what the part can do to a household's data if it is wrong, not how likely that i
| **7** | **SHIPPED — controller v0.271.0 (`9cf13a3`), proven on 9202 over four simulated nights 2026-09-24/25 and watched through the demo boxes' first real automatic night (`audits/DRILL-night-2026-09-25.md`).** Decisions 31–33 taken unattended. Gaps a scratch box cannot close: R-687; no resume after a restart: R-686. *(Before:)* **SPIKED 2026-09-24 (night, Part C) — the measured build brief is §6.4.2 below.** **The update leg in the chain + the automatic caller + the switch.** A leg that starts when the off-site leg has ENDED (on every exit path, skips included), one app at a time, one step per press, `app_update.unattended` default ON (decision 12), `stacks.update_window` removed, reads `UpdateRefusal.Reason`, remembers a failed step so it never re-presses it, and the full-system backup's gate waits for it until W+5h (decision 20). | 11, 12, 14, 15, 20 | **3** | 1, 2, 5 | medium — the only part that acts with nobody watching; everything above is what makes it safe |
| **8** | **SHIPPED — controller v0.265.0 + hub v0.121.0, proven live on 9202 2026-09-23** (`audits/cleanup-2026-09-23/`; a kernel `oom_kill` counter, not the sticky flag — `08` §6.2). **R-636** — the same OOM key re-firing escalates instead of staying one `warning` for six hours. | R-636 | **1** | — | none |
| **9** | **SHIPPED — controller v0.265.0, proven live on 9202 2026-09-23** (badge „Megállítva — visszaállítás szükséges" / "Stopped — restore needed", no Update button, 409 `held` unchanged). **R-625** — a held app stops offering an Update it will refuse. With the undo, holds become rarer; the lie on the page does not go away by itself. | R-625 | **0.5** | 1 | none |
| **10** | **PostgreSQL majors converted by the box.** A guarded-update step: `pg_dumpall` from the old engine, a NEW datadir (the old one kept aside, never deleted, until the check passes), load, check; then each of the eleven apps proven on the bench before its catalog move. | 16, R-463 | **2 + 3** | 1 (the same load discipline), 4 | **HIGH** — it rebuilds the datadir; bounded by keeping the old datadir aside |
| **10** | **SHIPPED FOR DOCMOST — controller v0.273.0 + catalog `6a4a5f0` (harness v4, the gate's proof clause), proven live on 9202 2026-09-25** (`audits/night-2026-09-26/D/`): one press converted docmost 16 → 18 in **9.9 s** of engine work (**42 s** end to end), the check equal (2 databases, 48 tables, 71 rows), `PG_VERSION` 18, the seed read back; the load failing (an extension 18 lacks) → **undone in 43.6 s**; the app unhealthy on 18 → **undone in 148 s** (90 s drill health timeout); the controller SIGKILLed one second after the volume was emptied → the restart **undid it in 40 s**; each ending on 16 with the seed read back. The old datadir's copy is kept until a backup is proven after the conversion. Decisions 35, 37, 38. Every other PostgreSQL app stays refused by the gate until its own two-venue proof. *(Before:)* **PostgreSQL majors converted by the box.** A guarded-update step: `pg_dumpall` from the old engine, a NEW datadir (the old one kept aside, never deleted, until the check passes), load, check; then each of the eleven apps proven on the bench before its catalog move. | 16, R-463 | **2 + 3** | 1 (the same load discipline), 4 | **HIGH** — it rebuilds the datadir; bounded by keeping the old datadir aside |
| **11** | **Fleet view** — per compose service: installed ref, catalog ref, badge state in the report; the hub lists boxes behind. | 18, R-451 | 2 | — | none — **deferred by the ruling** until the fleet grows |
**Recommended order: 1 → 2 + 3 → 4 → 5 → 6 → 7 → 8 → 9 → 10**, part 11 when the fleet grows. **Total
@@ -0,0 +1,38 @@
# A1 — what the postgres images declare (9202, 2026-09-25T10:20:36Z)
## postgres:16-alpine
digest=postgres@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea PGDATA/ENV=PG_MAJOR=16 PG_VERSION=16.15 PG_SHA256=c1575341fa7bd40f5274ea465b34390f4dc64cdd0770af327005caaeb9f6b7ed PGDATA=/var/lib/postgresql/data VOLUME={"/var/lib/postgresql/data":{}}
## postgres:17-alpine
digest=postgres@sha256:b0f9560a2de083e2cc7382e75f808c7381a32852a7ec49117deedb300e552b24 PGDATA/ENV=PG_MAJOR=17 PG_VERSION=17.11 PG_SHA256=dd27f2b3c59e73ed14aa3324901242bf69a032a6347805f274e6260322d42979 PGDATA=/var/lib/postgresql/data VOLUME={"/var/lib/postgresql/data":{}}
## postgres:18-alpine
digest=postgres@sha256:77f585114c32fbca283dc835b0596f4e52b51b4c6662d7810b2f4084f60a1873 PGDATA/ENV=PG_MAJOR=18 PG_VERSION=18.6 PG_SHA256=555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f PGDATA=/var/lib/postgresql/18/docker VOLUME={"/var/lib/postgresql":{}}
# A1b — postgres:18-alpine with an EMPTY volume at /var/lib/postgresql/data (the templates' mount)
state: exited exit=1
Error: in 18+, these Docker images are configured to store database data in a
format which is compatible with "pg_ctlcluster" (specifically, using
major-version-specific directory names). This better reflects how
PostgreSQL itself works, and how upgrades are to be performed.
See also https://github.com/docker-library/postgres/pull/1259
Counter to that, there appears to be PostgreSQL data in:
/var/lib/postgresql/data (unused mount/volume)
This is usually the result of upgrading the Docker image without
upgrading the underlying database using "pg_upgrade" (which requires both
versions).
The suggested container configuration for 18+ is to place a single mount
at /var/lib/postgresql which will then place PostgreSQL data in a
subdirectory, allowing usage of "pg_upgrade --link" without mount point
boundary issues.
See https://github.com/docker-library/postgres/issues/37 for a (long)
discussion around this process, and suggestions for how to do so.
# A1c — control: postgres:18-alpine with an empty volume at /var/lib/postgresql (its own layout)
state: running
2026-09-25 10:21:17.475 UTC [1] LOG: listening on Unix socket "/var/run/postgresql/.s.PGSQL.5432"
2026-09-25 10:21:17.488 UTC [66] LOG: database system was shut down at 2026-09-25 10:21:17 UTC
2026-09-25 10:21:17.496 UTC [1] LOG: database system is ready to accept connections
18
@@ -0,0 +1,37 @@
# A2/A3/A5 — docmost seeded on 9202, measured on a COPY of its datadir (2026-09-25T10:25:58Z)
datadir size: 65.5M
A3 counts (old engine): 0.49s, 48 tables, 62 rows total
db:docmost owner=docmost enc=UTF8 coll=en_US.utf8
db:postgres owner=docmost enc=UTF8 coll=en_US.utf8
role:docmost super=true login=true pw=true
ext:docmost:pg_trgm
ext:docmost:plpgsql
ext:docmost:unaccent
A2 pg_dumpall: 0.59s, 132184 B, tail: -- PostgreSQL database dump complete
A2 pg_dump --no-owner --no-privileges docmost: 0.27s, 123545 B
what pg_dumpall carries that the per-db dump does not:
1 CREATE ROLE docmost;
1 CREATE DATABASE docmost WITH TEMPLATE = template0 ENCODING = 'UTF8' LOCALE_PROVIDER = libc LOCALE = 'en_US.utf8';
1 ALTER TABLE public.workspaces OWNER TO docmost;
1 ALTER TABLE public.workspace_invitations OWNER TO docmost;
1 ALTER TABLE public.watchers OWNER TO docmost;
1 ALTER TABLE public.users OWNER TO docmost;
1 ALTER TABLE public.user_tokens OWNER TO docmost;
1 ALTER TABLE public.user_sessions OWNER TO docmost;
1 ALTER TABLE public.user_mfa OWNER TO docmost;
1 ALTER TABLE public.templates OWNER TO docmost;
1 ALTER TABLE public.spaces OWNER TO docmost;
1 ALTER TABLE public.space_members OWNER TO docmost;
## route dumpall: new engine init 2.64s, load 1.76s, counts 0.48s, PG_VERSION=18, load errors=2
1 ERROR: database "docmost" already exists
1 ERROR: role "docmost" already exists
dbs/owners/extensions/row counts: IDENTICAL
roles before: role:docmost super=true login=true pw=true
roles after: role:docmost super=true login=true pw=true
datadir after: 49.9M
## route pgdump: new engine init 2.64s, load 1.44s, counts 0.49s, PG_VERSION=18, load errors=0
dbs/owners/extensions/row counts: IDENTICAL
roles before: role:docmost super=true login=true pw=true
roles after: role:docmost super=true login=true pw=true
datadir after: 49.8M
A5 space: copy of the old datadir = 68707765 B; dump = 132184 B; new datadir = 52271397 B
@@ -0,0 +1,13 @@
# A4 — the undo after the DB volume was EMPTIED, by hand with the product's own helper commands (undo.go Copy/Restore), 9202 2026-09-25T10:27:34Z
app stopped: 0 running
copy with marker: ok
volume emptied: 0 entries left
Restore (the product's command): ok
PG_VERSION after restore: 16
2026-09-25 12:27:39.221 CEST [29] LOG: database system was shut down at 2026-09-25 12:27:35 CEST
2026-09-25 12:27:39.230 CEST [1] LOG: database system is ready to accept connections
copy removed
12:34:29 docmost: login as the seeded user http=200 ok=True
12:34:29 A4-after-undo-retry: read docmost=True
12:34:30 docmost: login as the seeded user http=200 ok=True
12:34:30 A4-after-undo-productstart: read docmost=True
@@ -0,0 +1,88 @@
# Part A — the PostgreSQL conversion spike on docmost (9202, 2026-09-25 midday)
Written BEFORE any build. docmost 0.96.0 / `postgres:16-alpine` installed on 9202 through the product
(`POST /api/stacks/docmost/deploy`), seeded through its own API (a workspace + an account), read back
(`data-readback.txt`: `A-read0`). Every measurement below ran on a COPY of its datadir.
## A1 — the target major: 18 (CC-unattended decision, `09` §3 decision 37)
**Measured (`A1-images.txt`):**
- `postgres:16-alpine` and `17-alpine`: `PGDATA=/var/lib/postgresql/data`, `VOLUME /var/lib/postgresql/data`.
- `postgres:18-alpine` (18.6): `PGDATA=/var/lib/postgresql/18/docker`, `VOLUME /var/lib/postgresql`.
- **18 with an EMPTY volume at `/var/lib/postgresql/data` (where all eleven templates mount it) REFUSES:
exit 1** — *"in 18+, these Docker images are configured to store database data in a format which is
compatible with pg_ctlcluster … there appears to be PostgreSQL data in: /var/lib/postgresql/data (unused
mount/volume)"*. It refuses even when that mount is empty. Control: the same image with the volume at
`/var/lib/postgresql` initialises and serves, `PG_VERSION` 18.
- **docmost's upstream compose ships `image: postgres:18` with `db_data:/var/lib/postgresql`**
(github docmost/docmost `docker-compose.yml`, read 2026-09-25). Its docs name no other version.
**So the brief's claim is TRUE**, and stronger than stated: 18 refuses at the old mount, even empty. A move to
18 therefore needs the mount moved to `/var/lib/postgresql` in the same step, which the step's own definition
carries (the template's compose / `steps/<key>.yml`).
**Decision 37 — docmost converts 16 → 18, not 16 → 17.**
*One sentence:* which major does the first conversion target? **Options:** (a) 17 — no mount change;
(b) 18 — the mount moves to `/var/lib/postgresql` in the same step. **Costs:** (a) is a second conversion
later, and upstream already ships 18, so every box converts twice; (b) the step changes the volume's mount
point, so an undo must put the OLD mount back with the old data — which the undo does anyway (it restores the
old definition and the volume's bytes). **Why (b):** one conversion is better than two when the app's own
upstream runs the newer major; the mount change rides the step's definition and the undo's existing
definition restore. Reversible (the catalog can pin 17 instead before any box moves).
## A2 — what to load from: `pg_dumpall` from the old engine (decision 38)
`A2-A3-A5-measure.txt`. On the seeded datadir (48 tables, 62 rows, 49 MB):
| | size | time | what it carries |
|---|---|---|---|
| `pg_dumpall` (old engine) | 132 184 B | 0.59 s | roles WITH their password hashes, `CREATE DATABASE … LOCALE`, owners, grants |
| `pg_dump --no-owner --no-privileges docmost` (today's safety dump) | 123 545 B | 0.27 s | one database's schema + data; no roles, no owners, no database settings |
Both loaded into a fresh 18 gave **identical** databases, owners, encodings, collations, extensions
(`pg_trgm`, `plpgsql`, `unaccent`) and row counts for docmost — because docmost has ONE role (the bootstrap
superuser the new engine's entrypoint recreates from the same env) and ONE database.
**So the brief's claim is TRUE** (the safety dump is per-database `pg_dump` with `--no-owner --no-privileges`
— `appbackup/dbdump.go`), **and for docmost it would have been enough.** It would not be for an app with a
second role or a second database, which the other ten have not been measured for.
**Decision 38 — the box loads from a `pg_dumpall` taken from the OLD engine at conversion time, and the load
tolerates NO error.** *One sentence:* what does the box load into the new engine? **Options:** (a) the
existing safety dump; (b) a `pg_dumpall` of the old engine, loaded with `ON_ERROR_STOP`. **Costs:** (a) is
free, and silently drops any role, grant or database setting beyond the bootstrap ones; (b) one more dump
(0.6 s here) and a load that must not trip on the two objects the new engine's entrypoint already made (the
bootstrap role and database: measured, `pg_dumpall` loaded raw gives exactly two `already exists` ERRORs).
**Why (b):** decision 16 says "save everything". The two known collisions are removed precisely — the
entrypoint's empty databases are dropped before the load (only when they hold no table), and the dump's
`CREATE ROLE <name>;` line is skipped for a role that already exists (its `ALTER ROLE … PASSWORD` still
runs) — so any OTHER error stops the load and the undo runs. Proven by a test and live.
## A3 — the check
Old engine, before it stops; new engine, after the load — the same query set (0.5 s each here):
per database: owner, encoding, collation; every role (superuser, login, has a password); every extension by
name; every table's exact row count (`count(*)` per table). After: plus `PG_VERSION` of the new datadir
(`$PGDATA/PG_VERSION`, which follows 17's and 18's different layouts). All IDENTICAL for docmost on both
load routes.
## A4 — the undo after the volume was EMPTIED: proven by hand
`A4-undo-after-empty.txt`: copy with the product's own helper command (+ marker) → volume emptied (0
entries) → the product's `Restore` command → `PG_VERSION` 16 → the old engine "ready to accept
connections" → **the seeded account logged in through docmost's front door** (`A4-after-undo-productstart`).
**The brief's claim is TRUE.**
*Recorded as it happened:* my first start after the restore was a bare `docker compose up -d` in the stack
dir, which has no `.env` (the controller passes the app's env itself). docmost started without its secrets,
crash-looped, and **the box stopped it after 7 restarts in 10 min (decision 28) — the product working**, not
a fault. Started again with the product's own Start (`POST /api/stacks/docmost/start`); the seed read back.
## A5 — space
The conversion adds, on top of today's undo copy (the old datadir, 68.7 MB here): the dump (132 KB — 0.2 % of
the datadir) and nothing for the new datadir (it is built IN the emptied volume: 52.3 MB). The box's check:
**the dump's bound is the DB volume's own size** (a logical dump of live rows is smaller than the datadir
holding them, measured 0.2 %), with a 25 % margin, on the filesystem that holds the stack directory (where the
dump is written), plus the fixed 2 GB floor the update already keeps. Refused before anything moves.
@@ -0,0 +1,55 @@
### RED-PROOF B1 no mark → refused — mutation in pgconvert.go
--- FAIL: TestConvert_MajorWithoutMarkIsRefused (0.00s)
pgconvert_test.go:215: the preflight let it through: <nil>
--- FAIL: TestConvert_JobRefusesAMarkThatDoesNotMatch (0.01s)
pgconvert_test.go:235: phase "done" key "", want failed with the no-test sentence
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.020s
### RED-PROOF B2 marker validated again before emptying — mutation in pgconvert.go
--- FAIL: TestConvert_NeverEmptiedWithoutMarker (0.01s)
pgconvert_test.go:338: Empty was called although the copy had no marker
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.015s
FAIL
### RED-PROOF B3 counts compared — mutation in pgconvert.go
--- FAIL: TestConvert_CountsDifferAreUndone (0.01s)
pgconvert_test.go:271: ended "done" (), want undone
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.021s
FAIL
### RED-PROOF B3 PG_VERSION checked — mutation in pgconvert.go
--- FAIL: TestConvert_WrongVersionIsUndone (0.01s)
pgconvert_test.go:304: ended "done" (), want undone
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.019s
FAIL
### RED-PROOF B3 cut-off dump refused — mutation in pgconvert.go
--- FAIL: TestConvert_CutOffDumpNeverEmptiesTheVolume (0.01s)
pgconvert_test.go:315: ended "done" (), want undone
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.018s
FAIL
### RED-PROOF B4 restart during converting → undo — mutation in update.go
--- FAIL: TestConvert_RestartDuringConvertingIsUndone (0.06s)
pgconvert_test.go:399: RecoverUpdates resumed []
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.062s
FAIL
### RED-PROOF B5 old datadir copy kept — mutation in update.go
--- FAIL: TestConvert_HappyPath (0.01s)
pgconvert_test.go:196: the kept copy is not recorded: <nil>
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.014s
FAIL
### RED-PROOF B5 release wired — mutation in main.go
--- FAIL: TestConvert_ReleaseIsWiredAtStartup (0.01s)
conversion_wiring_test.go:11: ReleaseConversionCopies is never called — a converted app's old datadir copy would stay on disk for ever
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.015s
FAIL
### RED-PROOF A5 space checked before anything moves — mutation in update.go
--- FAIL: TestConvert_NoSpaceIsRefusedWithNothingMoved (0.01s)
pgconvert_test.go:353: phase "done" key ""
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.014s
FAIL
@@ -0,0 +1,18 @@
--- PASS: TestConvert_HappyPath (0.01s)
--- PASS: TestConvert_MajorWithoutMarkIsRefused (0.01s)
--- PASS: TestConvert_JobRefusesAMarkThatDoesNotMatch (0.00s)
--- PASS: TestConvert_CountsDifferAreUndone (0.01s)
--- PASS: TestConvert_LoadFailureIsUndone (0.01s)
--- PASS: TestConvert_HealthFailureIsUndone (0.01s)
--- PASS: TestConvert_WrongVersionIsUndone (0.01s)
--- PASS: TestConvert_CutOffDumpNeverEmptiesTheVolume (0.01s)
--- PASS: TestConvert_NeverEmptiedWithoutMarker (0.01s)
--- PASS: TestConvert_NoSpaceIsRefusedWithNothingMoved (0.00s)
--- PASS: TestConvert_RestartDuringConvertingIsUndone (0.06s)
--- PASS: TestConvert_ReleaseAfterAProvenBackup (0.01s)
--- PASS: TestConvert_PostgresMajorFromTags (0.00s)
--- PASS: TestConvert_PostgresFamilyMatchesTheBackupSide (0.00s)
--- PASS: TestConvert_FilterSkipsOnlyExactLines (0.00s)
--- PASS: TestConvert_ValidateDumpAll (0.00s)
--- PASS: TestConvert_MarkDoesNotChangeOldLadderPrints (0.00s)
ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.189s
@@ -0,0 +1,13 @@
local dir active 40453376 22728440 15637820 56.18%
nvme-scratch dir active 983379700 75082704 858270384 7.64%
template debian-13-standard_13.6-1_arm64.tar.zst
sync_wait: 34 An error occurred in another process (expected sequence number 7)
__lxc_start: 2288 Failed to spawn container "9401"
startup for container '9401' failed
purging CT 9401 from related configurations..
destroyed wrong-arch 9401
debian-13-standard_13.6-1_amd64.tar.zst
192.168.0.121/24
Docker version 26.1.5+dfsg1, build a72d7cd
Docker Compose version 2.26.1-4
net-ok
@@ -0,0 +1,11 @@
ok 18 moves the data mount to /var/lib/postgresql
ok 18 leaves the redis mount alone
ok 16 keeps /var/lib/postgresql/data
ok back from 18 to 16 restores /data
ok pg_conversion names the one moving service
ok within a major is no conversion
ok pgvector's pg16 tag reads 16
ok an entry with a good mark is well-formed
ok a mark naming a service outside `to` is refused
ok a mark going DOWN is refused
pg conversion tests: 0 failed
@@ -0,0 +1,8 @@
ok FACT: docmost-postgres 16-alpine -> 17-alpine bundled with the app bump rc=1 (expected 1)
ok FACT: docmost-postgres postgres:16-alpine -> 17-alpine rc=1 (expected 1)
ok GENUINE: docmost-postgres 16 -> 18 ALONE with its proven two-venue MARKED entry rc=0 (expected 0)
ok DECOY: the entry is proven on both venues but carries NO conversion mark rc=1 (expected 1)
ok DECOY: the marked entry cites ONE venue only (no box_evidence) rc=1 (expected 1)
ok DECOY: the mark sits in a COMMENT, the entry has none rc=1 (expected 1)
ok DECOY: the marked entry, but the move is BUNDLED with the app's own bump rc=1 (expected 1)
ok FACT: adventurelog's postgis 16 -> 17 (the postgis family was never judged before) rc=1 (expected 1)
@@ -0,0 +1,7 @@
# RED-PROOF of check-engine-major.py's PostgreSQL clause (2026-09-25, catalog tree before commit)
# mutation: `if why is None and not others:` -> `if not others:` (the ladder proof is not consulted)
# python3 scripts/test_gate_decoys.py, lines quoted verbatim:
FAIL: FACT: docmost-postgres 16-alpine -> 17-alpine bundled with the app bump: rc=0 expected 1; missing ['ENGINE-MAJOR GATE FAILED']
FAIL: FACT: docmost-postgres postgres:16-alpine -> 17-alpine: rc=0 expected 1; missing ['R-463']
FAIL: DECOY: the entry is proven on both venues but carries NO conversion mark: rc=0 expected 1; missing ['NOT PROVEN FOR THIS APP', 'engine_conversion None']
# tree restored from the saved copy; the suite green again: C2-decoys-green.txt
@@ -0,0 +1,5 @@
# 9202 before: gitea.dooplex.hu/admin/felhom-controller:0.272.0
gitea.dooplex.hu/admin/felhom-controller:0.273.0-rc1
# 9202 after: gitea.dooplex.hu/admin/felhom-controller:0.273.0-rc1 healthy
2026/09/25 11:04:49 main.go:336: [INFO] felhom-controller 0.273.0-rc1 starting (customer: demo-hp, domain: enkisfelhom.hu)
2026/09/25 11:04:49 undo.go:675: [INFO] [stacks] applied-meta backfill (R-646): recorded 0 []; skipped 0 [] — not current with the catalog, their pinned version's .felhom.yml is no longer on the box
@@ -0,0 +1,10 @@
git:
branch: main
repo_url: https://gitea.dooplex.hu/admin/app-catalog-drill.git
sync_interval: 15m
token: <redacted>
username: "admin"
hub:
update:
health_timeout: 90s
@@ -0,0 +1,48 @@
# D1-happy-path — 2026-09-25T13:12:42+0200; controller gitea.dooplex.hu/admin/felhom-controller:0.273.0-rc1
before: PG_VERSION/image=['16', 'postgres:16-alpine@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} catalog={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:18-alpine', 'docmost-redis': 'redis:7-alpine'}
before: seed reads back through the front door = True
phase +0.0s safety-dump | Adatbázis pillanatkép… | err=None
phase +2.1s pulling | Új verzió letöltése… | err=None
phase +3.2s copying | Az adatok másolása a frissítés előtt… | err=None
phase +5.8s converting | Adatbázis átalakítása | err=None
phase +15.8s starting | Indítás az új verzióval… | err=None
phase +26.8s verifying | Működés ellenőrzése… | err=None
phase +42.1s done | Frissítve | err=None
result: {"accepted": true, "http": "202", "duration_s": 42.1, "final_phase": "done", "update_error": null, "hold_reason": null, "state": "running"}
after: PG_VERSION/image=['18', 'postgres:18-alpine@sha256:77f585114c32fbca283dc835b0596f4e52b51b4c6662d7810b2f4084f60a1873'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:18-alpine', 'docmost-redis': 'redis:7-alpine'} conversion_copy={'volume': 'docmost_docmost_postgres_data', 'copy': 'docmost_docmost_postgres_data.pre-update-20260925T111256Z', 'at': '2026-09-25T11:13:34Z', 'from': 16, 'to': 18}
after: seed reads back through the front door = True
controller lines (docmost / conversion):
2026/09/25 11:12:52 update.go:536: [INFO] [stacks] update docmost: accepted — guarded update started
2026/09/25 11:12:52 update.go:1365: [INFO] [stacks] update docmost: phase checking
2026/09/25 11:12:52 update.go:786: [INFO] [stacks] update docmost: ladder — the last step (2 of 2) — the catalog's current definition
2026/09/25 11:12:52 update.go:800: [INFO] [stacks] update docmost: this step CONVERTS docmost-postgres PostgreSQL 16 → 18 (the ladder entry's mark); database volume docmost_docmost_postgres_data
2026/09/25 11:12:52 update.go:814: [INFO] [stacks] update docmost: precondition met — Tier 1 (own recovery unit) copy from 2026-09-25T11:11:16Z (2m0s old, limit 24h0m0s)
2026/09/25 11:12:52 update.go:1365: [INFO] [stacks] update docmost: phase safety-dump
2026/09/25 11:12:53 update.go:846: [INFO] [stacks] update docmost: safety dump done (1 file(s)) [/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260925T111253Z-docmost-postgres.sql]
2026/09/25 11:12:54 undo.go:287: [INFO] [stacks] update docmost: the undo copy will hold 3 named volume(s), 66.4 MiB
2026/09/25 11:12:54 pgconvert.go:286: [INFO] [stacks] update docmost: conversion space — database volume docmost_docmost_postgres_data is 65 MB; the dump needs up to 82 MB + the 2 GB floor; 27.7 GB free beside the stack dir
2026/09/25 11:12:54 update.go:1365: [INFO] [stacks] update docmost: phase pinning
2026/09/25 11:12:54 pin.go:373: [INFO] [stacks] update docmost: pin advanced to /opt/docker/felhom-controller/data/catalog-cache/templates/docmost/docker-compose.yml (docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine)
2026/09/25 11:12:54 update.go:1365: [INFO] [stacks] update docmost: phase pulling
2026/09/25 11:12:55 update.go:1365: [INFO] [stacks] update docmost: phase copying
2026/09/25 11:12:56 update.go:1365: [INFO] [stacks] update docmost: phase copying
2026/09/25 11:12:57 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_postgres_data → docmost_docmost_postgres_data.pre-update-20260925T111256Z in 829ms
2026/09/25 11:12:57 update.go:1365: [INFO] [stacks] update docmost: phase copying
2026/09/25 11:12:58 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_redis_data → docmost_docmost_redis_data.pre-update-20260925T111256Z in 440ms
2026/09/25 11:12:58 update.go:1365: [INFO] [stacks] update docmost: phase copying
2026/09/25 11:12:58 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_storage → docmost_docmost_storage.pre-update-20260925T111256Z in 418ms
2026/09/25 11:12:58 update.go:1365: [INFO] [stacks] update docmost: phase converting
2026/09/25 11:12:58 pgconvert.go:369: [INFO] [stacks] update docmost: CONVERTING docmost-postgres PostgreSQL 16 → 18 (volume docmost_docmost_postgres_data, kept copy docmost_docmost_postgres_data.pre-update-20260925T111256Z)
2026/09/25 11:13:02 pgconvert.go:414: [INFO] [stacks] update docmost: pg_dumpall from 16 done in 566ms — 131 kB, completion line present; before: 2 database(s), 48 table(s), 71 row(s)
2026/09/25 11:13:02 update.go:1365: [INFO] [stacks] update docmost: phase converting
2026/09/25 11:13:03 pgconvert.go:430: [INFO] [stacks] update docmost: database volume docmost_docmost_postgres_data emptied (its copy docmost_docmost_postgres_data.pre-update-20260925T111256Z holds the old datadir, marker checked twice)
2026/09/25 11:13:07 pgconvert.go:459: [INFO] [stacks] update docmost: loaded into 18 in 1.753s (dropped the entrypoint's empty database(s) [docmost]; skipped CREATE ROLE for [docmost])
2026/09/25 11:13:08 pgconvert.go:475: [INFO] [stacks] update docmost: CONVERTED 16 → 18 in 9.942s — the check is equal (2 database(s), 48 table(s), 71 row(s)), PG_VERSION 18
2026/09/25 11:13:08 update.go:1365: [INFO] [stacks] update docmost: phase starting
2026/09/25 11:13:19 update.go:1365: [INFO] [stacks] update docmost: phase verifying
2026/09/25 11:13:34 update.go:997: [INFO] [stacks] update docmost: healthy after 15s (the app's health check passed)
2026/09/25 11:13:34 update.go:1015: [INFO] [stacks] update docmost: KEEPING the pre-conversion datadir copy docmost_docmost_postgres_data.pre-update-20260925T111256Z (PostgreSQL 16) until a backup of the converted app is proven
2026/09/25 11:13:34 update.go:1033: [INFO] [stacks] update docmost: DONE in 42s
undo copies now: docmost_docmost_postgres_data.pre-update-20260925T111256Z
done
@@ -0,0 +1,9 @@
# 'Mentés most' on docmost after the conversion (the release of the kept copy waits for a PROVEN backup) 2026-09-25T13:13:59+0200
13:13:59 [4] backup press SKIPPED for docmost (R-648: whole-box only; the update's backing-up phase backs up docmost alone)
None
-rw-r--r-- 1 root root 154767 2026-09-25T11:06:21 pre-restore-20260925T110621Z-docmost-postgres.sql
-rw-r--r-- 1 root root 154903 2026-09-25T11:08:02 pre-restore-20260925T110802Z-docmost-postgres.sql
-rw-r--r-- 1 root root 155357 2026-09-25T11:11:16 pre-restore-20260925T111115Z-docmost-postgres.sql
-rw-r--r-- 1 root root 155810 2026-09-25T11:12:53 pre-restore-20260925T111253Z-docmost-postgres.sql
docmost_docmost_postgres_data.pre-update-20260925T111256Z
@@ -0,0 +1,47 @@
# Da-load-fails — 2026-09-25T13:06:11+0200; controller gitea.dooplex.hu/admin/felhom-controller:0.273.0-rc1
before: PG_VERSION/image=['16', 'postgres:16-alpine@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} catalog={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:18-alpine', 'docmost-redis': 'redis:7-alpine'}
before: seed reads back through the front door = True
phase +0.0s safety-dump | Adatbázis pillanatkép… | err=None
phase +2.1s pulling | Új verzió letöltése… | err=None
phase +2.7s copying | Az adatok másolása a frissítés előtt… | err=None
phase +5.8s converting | Adatbázis átalakítása | err=None
phase +13.7s undoing | Visszaállítás az előző változatra… | err=None
phase +43.6s undone | Visszaállítva az előző változatra | err=None
result: {"accepted": true, "http": "202", "duration_s": 43.6, "final_phase": "undone", "update_error": null, "hold_reason": null, "state": "running"}
after: PG_VERSION/image=['16', 'postgres:16-alpine@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} conversion_copy=None
after: seed reads back through the front door = True
controller lines (docmost / conversion):
2026/09/25 11:06:21 update.go:536: [INFO] [stacks] update docmost: accepted — guarded update started
2026/09/25 11:06:21 update.go:1365: [INFO] [stacks] update docmost: phase checking
2026/09/25 11:06:21 update.go:786: [INFO] [stacks] update docmost: ladder — the last step (2 of 2) — the catalog's current definition
2026/09/25 11:06:21 update.go:800: [INFO] [stacks] update docmost: this step CONVERTS docmost-postgres PostgreSQL 16 → 18 (the ladder entry's mark); database volume docmost_docmost_postgres_data
2026/09/25 11:06:21 update.go:814: [INFO] [stacks] update docmost: precondition met — Tier 1 (own recovery unit) copy from 2026-09-25T11:04:49Z (2m0s old, limit 24h0m0s)
2026/09/25 11:06:21 update.go:1365: [INFO] [stacks] update docmost: phase safety-dump
2026/09/25 11:06:21 update.go:846: [INFO] [stacks] update docmost: safety dump done (1 file(s)) [/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260925T110621Z-docmost-postgres.sql]
2026/09/25 11:06:22 undo.go:287: [INFO] [stacks] update docmost: the undo copy will hold 3 named volume(s), 66.3 MiB
2026/09/25 11:06:23 pgconvert.go:286: [INFO] [stacks] update docmost: conversion space — database volume docmost_docmost_postgres_data is 65 MB; the dump needs up to 82 MB + the 2 GB floor; 27.7 GB free beside the stack dir
2026/09/25 11:06:23 update.go:1365: [INFO] [stacks] update docmost: phase pinning
2026/09/25 11:06:23 pin.go:373: [INFO] [stacks] update docmost: pin advanced to /opt/docker/felhom-controller/data/catalog-cache/templates/docmost/docker-compose.yml (docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine)
2026/09/25 11:06:23 update.go:1365: [INFO] [stacks] update docmost: phase pulling
2026/09/25 11:06:23 update.go:1365: [INFO] [stacks] update docmost: phase copying
2026/09/25 11:06:24 update.go:1365: [INFO] [stacks] update docmost: phase copying
2026/09/25 11:06:25 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_postgres_data → docmost_docmost_postgres_data.pre-update-20260925T110624Z in 869ms
2026/09/25 11:06:25 update.go:1365: [INFO] [stacks] update docmost: phase copying
2026/09/25 11:06:26 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_redis_data → docmost_docmost_redis_data.pre-update-20260925T110624Z in 431ms
2026/09/25 11:06:26 update.go:1365: [INFO] [stacks] update docmost: phase copying
2026/09/25 11:06:26 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_storage → docmost_docmost_storage.pre-update-20260925T110624Z in 422ms
2026/09/25 11:06:26 update.go:1365: [INFO] [stacks] update docmost: phase converting
2026/09/25 11:06:26 pgconvert.go:369: [INFO] [stacks] update docmost: CONVERTING docmost-postgres PostgreSQL 16 → 18 (volume docmost_docmost_postgres_data, kept copy docmost_docmost_postgres_data.pre-update-20260925T110624Z)
2026/09/25 11:06:30 pgconvert.go:414: [INFO] [stacks] update docmost: pg_dumpall from 16 done in 572ms — 130 kB, completion line present; before: 2 database(s), 48 table(s), 65 row(s)
2026/09/25 11:06:30 update.go:1365: [INFO] [stacks] update docmost: phase converting
2026/09/25 11:06:31 pgconvert.go:430: [INFO] [stacks] update docmost: database volume docmost_docmost_postgres_data emptied (its copy docmost_docmost_postgres_data.pre-update-20260925T110624Z holds the old datadir, marker checked twice)
2026/09/25 11:06:34 update.go:1066: [ERROR] [stacks] update docmost FAILED after the new version was started: conversion failed: loading the dump into 18: exit status 3 (psql: ERROR: extension "adminpack" is not available)
2026/09/25 11:06:34 update.go:1058: [INFO] [stacks] update docmost: kept 11240 bytes of the app's own log at /opt/docker/stacks/docmost/hold-logs/20260925T110634Z/compose-logs.txt before stopping it (R-621)
2026/09/25 11:06:34 update.go:1365: [INFO] [stacks] update docmost: phase undoing
2026/09/25 11:06:34 undo.go:515: [WARN] [stacks] update docmost: UNDO — putting back the previous version and its 3 volume copy(ies) (reason: conversion failed: loading the dump into 18: exit status 3 (psql: ERROR: extension "adminpack" is not available))
2026/09/25 11:07:04 unattended.go:490: [INFO] [stacks] update docmost: step docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine recorded as undone — the automatic leg will not press it again until the catalog's ladder changes (ladder 1cdde4765ba33137)
2026/09/25 11:07:04 undo.go:640: [INFO] [stacks] update docmost: event app_update_undone
2026/09/25 11:07:04 undo.go:568: [INFO] [stacks] update docmost: UNDONE in 30s — the previous version is running on the data from before the update (the app's health check passed)
undo copies now:
done
@@ -0,0 +1,6 @@
# case (a) setup: an object 16 has and 18 does not — the adminpack extension (removed in PostgreSQL 17)
CREATE EXTENSION
adminpack pg_trgm plpgsql unaccent
control: 18 has no adminpack: 0
# after case (a): adminpack removed again
DROP EXTENSION
@@ -0,0 +1,53 @@
# Db-health-fails — 2026-09-25T13:07:52+0200; controller gitea.dooplex.hu/admin/felhom-controller:0.273.0-rc1
before: PG_VERSION/image=['16', 'postgres:16-alpine@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} catalog={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:18-alpine', 'docmost-redis': 'redis:7-alpine'}
before: seed reads back through the front door = True
phase +0.0s safety-dump | Adatbázis pillanatkép… | err=None
phase +2.1s pulling | Új verzió letöltése… | err=None
phase +2.6s copying | Az adatok másolása a frissítés előtt… | err=None
phase +5.8s converting | Adatbázis átalakítása | err=None
phase +15.2s starting | Indítás az új verzióval… | err=None
phase +26.3s verifying | Működés ellenőrzése… | err=None
phase +117.1s undoing | Visszaállítás az előző változatra… | err=None
phase +148.1s undone | Visszaállítva az előző változatra | err=None
result: {"accepted": true, "http": "202", "duration_s": 148.1, "final_phase": "undone", "update_error": null, "hold_reason": null, "state": "running"}
after: PG_VERSION/image=['16', 'postgres:16-alpine@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} conversion_copy=None
after: seed reads back through the front door = True
controller lines (docmost / conversion):
2026/09/25 11:08:02 update.go:536: [INFO] [stacks] update docmost: accepted — guarded update started
2026/09/25 11:08:02 update.go:1365: [INFO] [stacks] update docmost: phase checking
2026/09/25 11:08:02 update.go:786: [INFO] [stacks] update docmost: ladder — the last step (2 of 2) — the catalog's current definition
2026/09/25 11:08:02 update.go:800: [INFO] [stacks] update docmost: this step CONVERTS docmost-postgres PostgreSQL 16 → 18 (the ladder entry's mark); database volume docmost_docmost_postgres_data
2026/09/25 11:08:02 update.go:814: [INFO] [stacks] update docmost: precondition met — Tier 1 (own recovery unit) copy from 2026-09-25T11:06:21Z (2m0s old, limit 24h0m0s)
2026/09/25 11:08:02 update.go:1365: [INFO] [stacks] update docmost: phase safety-dump
2026/09/25 11:08:02 update.go:846: [INFO] [stacks] update docmost: safety dump done (1 file(s)) [/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260925T110802Z-docmost-postgres.sql]
2026/09/25 11:08:03 undo.go:287: [INFO] [stacks] update docmost: the undo copy will hold 3 named volume(s), 66.3 MiB
2026/09/25 11:08:04 pgconvert.go:286: [INFO] [stacks] update docmost: conversion space — database volume docmost_docmost_postgres_data is 65 MB; the dump needs up to 82 MB + the 2 GB floor; 27.7 GB free beside the stack dir
2026/09/25 11:08:04 update.go:1365: [INFO] [stacks] update docmost: phase pinning
2026/09/25 11:08:04 pin.go:373: [INFO] [stacks] update docmost: pin advanced to /opt/docker/felhom-controller/data/catalog-cache/templates/docmost/docker-compose.yml (docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine)
2026/09/25 11:08:04 update.go:1365: [INFO] [stacks] update docmost: phase pulling
2026/09/25 11:08:04 update.go:1365: [INFO] [stacks] update docmost: phase copying
2026/09/25 11:08:05 update.go:1365: [INFO] [stacks] update docmost: phase copying
2026/09/25 11:08:06 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_postgres_data → docmost_docmost_postgres_data.pre-update-20260925T110805Z in 839ms
2026/09/25 11:08:06 update.go:1365: [INFO] [stacks] update docmost: phase copying
2026/09/25 11:08:07 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_redis_data → docmost_docmost_redis_data.pre-update-20260925T110805Z in 435ms
2026/09/25 11:08:07 update.go:1365: [INFO] [stacks] update docmost: phase copying
2026/09/25 11:08:07 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_storage → docmost_docmost_storage.pre-update-20260925T110805Z in 441ms
2026/09/25 11:08:07 update.go:1365: [INFO] [stacks] update docmost: phase converting
2026/09/25 11:08:07 pgconvert.go:369: [INFO] [stacks] update docmost: CONVERTING docmost-postgres PostgreSQL 16 → 18 (volume docmost_docmost_postgres_data, kept copy docmost_docmost_postgres_data.pre-update-20260925T110805Z)
2026/09/25 11:08:10 pgconvert.go:414: [INFO] [stacks] update docmost: pg_dumpall from 16 done in 566ms — 130 kB, completion line present; before: 2 database(s), 48 table(s), 67 row(s)
2026/09/25 11:08:11 update.go:1365: [INFO] [stacks] update docmost: phase converting
2026/09/25 11:08:12 pgconvert.go:430: [INFO] [stacks] update docmost: database volume docmost_docmost_postgres_data emptied (its copy docmost_docmost_postgres_data.pre-update-20260925T110805Z holds the old datadir, marker checked twice)
2026/09/25 11:08:16 pgconvert.go:459: [INFO] [stacks] update docmost: loaded into 18 in 1.747s (dropped the entrypoint's empty database(s) [docmost]; skipped CREATE ROLE for [docmost])
2026/09/25 11:08:17 pgconvert.go:475: [INFO] [stacks] update docmost: CONVERTED 16 → 18 in 9.866s — the check is equal (2 database(s), 48 table(s), 67 row(s)), PG_VERSION 18
2026/09/25 11:08:17 update.go:1365: [INFO] [stacks] update docmost: phase starting
2026/09/25 11:08:28 update.go:1365: [INFO] [stacks] update docmost: phase verifying
2026/09/25 11:09:58 update.go:1066: [ERROR] [stacks] update docmost FAILED after the new version was started: not healthy: not healthy within 1m30s (last: health check failing)
2026/09/25 11:09:59 update.go:1058: [INFO] [stacks] update docmost: kept 13077 bytes of the app's own log at /opt/docker/stacks/docmost/hold-logs/20260925T110958Z/compose-logs.txt before stopping it (R-621)
2026/09/25 11:09:59 update.go:1365: [INFO] [stacks] update docmost: phase undoing
2026/09/25 11:09:59 undo.go:515: [WARN] [stacks] update docmost: UNDO — putting back the previous version and its 3 volume copy(ies) (reason: not healthy: not healthy within 1m30s (last: health check failing))
2026/09/25 11:10:30 unattended.go:490: [INFO] [stacks] update docmost: step docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine recorded as undone — the automatic leg will not press it again until the catalog's ladder changes (ladder 1cdde4765ba33137)
2026/09/25 11:10:30 undo.go:640: [INFO] [stacks] update docmost: event app_update_undone
2026/09/25 11:10:30 undo.go:568: [INFO] [stacks] update docmost: UNDONE in 31s — the previous version is running on the data from before the update (the app's health check passed)
undo copies now:
done
@@ -0,0 +1,83 @@
# Dc-restart-mid-convert — 2026-09-25T13:11:04+0200; controller gitea.dooplex.hu/admin/felhom-controller:0.273.0-rc1
before: PG_VERSION/image=['16', 'postgres:16-alpine@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} catalog={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:18-alpine', 'docmost-redis': 'redis:7-alpine'}
before: seed reads back through the front door = True
[kill] 2026/09/25 11:11:23 pgconvert.go:430: [INFO] [stacks] update docmost: database volume docmost_docmost_postgres_data emptied (its copy docmost_docmost_postgres_data.pre-update-20260925T111119Z holds the old datadir, marker checked twice) | KILLED controller pid 554023 at 11:11:24.114456007
[kill] waiting for the controller to come back and the journal to be resumed
phase +0.0s safety-dump | Adatbázis pillanatkép… | err=None
phase +2.1s pulling | Új verzió letöltése… | err=None
phase +2.6s copying | Az adatok másolása a frissítés előtt… | err=None
phase +5.8s converting | Adatbázis átalakítása | err=None
phase +8.4s None | None | err=None
result: {"accepted": true, "http": "202", "duration_s": 8.5, "final_phase": null, "update_error": null, "hold_reason": null, "state": null, "after_restart": {"phase": "undone", "error": null, "hold": null}}
after: PG_VERSION/image=['16', 'postgres:16-alpine@sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea'] pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} conversion_copy=None
after: seed reads back through the front door = True
controller lines (docmost / conversion):
2026/09/25 11:11:15 update.go:536: [INFO] [stacks] update docmost: accepted — guarded update started
2026/09/25 11:11:15 update.go:1365: [INFO] [stacks] update docmost: phase checking
2026/09/25 11:11:15 update.go:786: [INFO] [stacks] update docmost: ladder — the last step (2 of 2) — the catalog's current definition
2026/09/25 11:11:15 update.go:800: [INFO] [stacks] update docmost: this step CONVERTS docmost-postgres PostgreSQL 16 → 18 (the ladder entry's mark); database volume docmost_docmost_postgres_data
2026/09/25 11:11:15 update.go:814: [INFO] [stacks] update docmost: precondition met — Tier 1 (own recovery unit) copy from 2026-09-25T11:08:02Z (3m0s old, limit 24h0m0s)
2026/09/25 11:11:15 update.go:1365: [INFO] [stacks] update docmost: phase safety-dump
2026/09/25 11:11:16 update.go:846: [INFO] [stacks] update docmost: safety dump done (1 file(s)) [/mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260925T111115Z-docmost-postgres.sql]
2026/09/25 11:11:17 undo.go:287: [INFO] [stacks] update docmost: the undo copy will hold 3 named volume(s), 66.4 MiB
2026/09/25 11:11:17 pgconvert.go:286: [INFO] [stacks] update docmost: conversion space — database volume docmost_docmost_postgres_data is 65 MB; the dump needs up to 82 MB + the 2 GB floor; 27.7 GB free beside the stack dir
2026/09/25 11:11:17 update.go:1365: [INFO] [stacks] update docmost: phase pinning
2026/09/25 11:11:17 pin.go:373: [INFO] [stacks] update docmost: pin advanced to /opt/docker/felhom-controller/data/catalog-cache/templates/docmost/docker-compose.yml (docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine)
2026/09/25 11:11:17 update.go:1365: [INFO] [stacks] update docmost: phase pulling
2026/09/25 11:11:18 update.go:1365: [INFO] [stacks] update docmost: phase copying
2026/09/25 11:11:19 update.go:1365: [INFO] [stacks] update docmost: phase copying
2026/09/25 11:11:20 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_postgres_data → docmost_docmost_postgres_data.pre-update-20260925T111119Z in 794ms
2026/09/25 11:11:20 update.go:1365: [INFO] [stacks] update docmost: phase copying
2026/09/25 11:11:20 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_redis_data → docmost_docmost_redis_data.pre-update-20260925T111119Z in 415ms
2026/09/25 11:11:20 update.go:1365: [INFO] [stacks] update docmost: phase copying
2026/09/25 11:11:21 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_storage → docmost_docmost_storage.pre-update-20260925T111119Z in 423ms
2026/09/25 11:11:21 update.go:1365: [INFO] [stacks] update docmost: phase converting
2026/09/25 11:11:21 pgconvert.go:369: [INFO] [stacks] update docmost: CONVERTING docmost-postgres PostgreSQL 16 → 18 (volume docmost_docmost_postgres_data, kept copy docmost_docmost_postgres_data.pre-update-20260925T111119Z)
2026/09/25 11:11:22 pgconvert.go:414: [INFO] [stacks] update docmost: pg_dumpall from 16 done in 539ms — 130 kB, completion line present; before: 2 database(s), 48 table(s), 69 row(s)
2026/09/25 11:11:23 update.go:1365: [INFO] [stacks] update docmost: phase converting
2026/09/25 11:11:23 pgconvert.go:430: [INFO] [stacks] update docmost: database volume docmost_docmost_postgres_data emptied (its copy docmost_docmost_postgres_data.pre-update-20260925T111119Z holds the old datadir, marker checked twice)
2026/09/25 11:11:24 update.go:1432: [WARN] [stacks] update recovery: docmost was interrupted while UNDOING (started 2026-09-25T11:11:15Z) — marking it Updating and RESUMING the undo
2026/09/25 11:11:24 main.go:496: [WARN] [update] 1 interrupted update(s) will resume after the backup side is wired: [docmost]
2026/09/25 11:11:24 main.go:596: [WARN] [update] resumed 1 interrupted update(s)
2026/09/25 11:11:24 update.go:1489: [INFO] [stacks] update docmost: resuming the UNDO after a controller restart (was converting)
2026/09/25 11:11:24 update.go:1066: [ERROR] [stacks] update docmost FAILED after the new version was started: the controller restarted during the database conversion — undoing it
2026/09/25 11:11:25 update.go:1058: [INFO] [stacks] update docmost: kept 5297 bytes of the app's own log at /opt/docker/stacks/docmost/hold-logs/20260925T111124Z/compose-logs.txt before stopping it (R-621)
2026/09/25 11:11:25 update.go:1365: [INFO] [stacks] update docmost: phase undoing
2026/09/25 11:11:25 undo.go:515: [WARN] [stacks] update docmost: UNDO — putting back the previous version and its 3 volume copy(ies) (reason: the controller restarted during the database conversion — undoing it)
2026/09/25 11:12:05 unattended.go:490: [INFO] [stacks] update docmost: step docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine recorded as undone — the automatic leg will not press it again until the catalog's ladder changes (ladder 1cdde4765ba33137)
2026/09/25 11:12:05 undo.go:640: [INFO] [stacks] update docmost: event app_update_undone
2026/09/25 11:12:05 undo.go:568: [INFO] [stacks] update docmost: UNDONE in 40s — the previous version is running on the data from before the update (the app's health check passed)
undo copies now:
done
# the whole restart, from the controller's own log
2026/09/25 11:11:20 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_postgres_data → docmost_docmost_postgres_data.pre-update-20260925T111119Z in 794ms
2026/09/25 11:11:20 update.go:1365: [INFO] [stacks] update docmost: phase copying
2026/09/25 11:11:20 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_redis_data → docmost_docmost_redis_data.pre-update-20260925T111119Z in 415ms
2026/09/25 11:11:20 update.go:1365: [INFO] [stacks] update docmost: phase copying
2026/09/25 11:11:21 undo.go:437: [INFO] [stacks] update docmost: copied docmost_docmost_storage → docmost_docmost_storage.pre-update-20260925T111119Z in 423ms
2026/09/25 11:11:21 update.go:1365: [INFO] [stacks] update docmost: phase converting
2026/09/25 11:11:21 pgconvert.go:369: [INFO] [stacks] update docmost: CONVERTING docmost-postgres PostgreSQL 16 → 18 (volume docmost_docmost_postgres_data, kept copy docmost_docmost_postgres_data.pre-update-20260925T111119Z)
2026/09/25 11:11:22 pgconvert.go:414: [INFO] [stacks] update docmost: pg_dumpall from 16 done in 539ms — 130 kB, completion line present; before: 2 database(s), 48 table(s), 69 row(s)
2026/09/25 11:11:23 update.go:1365: [INFO] [stacks] update docmost: phase converting
2026/09/25 11:11:23 pgconvert.go:430: [INFO] [stacks] update docmost: database volume docmost_docmost_postgres_data emptied (its copy docmost_docmost_postgres_data.pre-update-20260925T111119Z holds the old datadir, marker checked twice)
2026/09/25 11:11:24 main.go:336: [INFO] felhom-controller 0.273.0-rc1 starting (customer: demo-hp, domain: enkisfelhom.hu)
2026/09/25 11:11:24 update.go:1432: [WARN] [stacks] update recovery: docmost was interrupted while UNDOING (started 2026-09-25T11:11:15Z) — marking it Updating and RESUMING the undo
2026/09/25 11:11:24 main.go:496: [WARN] [update] 1 interrupted update(s) will resume after the backup side is wired: [docmost]
2026/09/25 11:11:24 main.go:596: [WARN] [update] resumed 1 interrupted update(s)
2026/09/25 11:11:24 update.go:1489: [INFO] [stacks] update docmost: resuming the UNDO after a controller restart (was converting)
2026/09/25 11:11:24 update.go:1066: [ERROR] [stacks] update docmost FAILED after the new version was started: the controller restarted during the database conversion — undoing it
2026/09/25 11:11:25 update.go:1058: [INFO] [stacks] update docmost: kept 5297 bytes of the app's own log at /opt/docker/stacks/docmost/hold-logs/20260925T111124Z/compose-logs.txt before stopping it (R-621)
2026/09/25 11:11:25 update.go:1365: [INFO] [stacks] update docmost: phase undoing
2026/09/25 11:11:25 undo.go:515: [WARN] [stacks] update docmost: UNDO — putting back the previous version and its 3 volume copy(ies) (reason: the controller restarted during the database conversion — undoing it)
2026/09/25 11:11:25 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/docmost/docker-compose.yml
2026/09/25 11:11:25 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/docmost/docker-compose.yml
2026/09/25 11:11:25 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/docmost/docker-compose.yml
2026/09/25 11:11:25 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/docmost/docker-compose.yml
2026/09/25 11:11:25 [INFO] [stacks] ParseComposeHDDMounts: found 0 HDD mounts for /opt/docker/stacks/docmost/docker-compose.yml
2026/09/25 11:11:34 healthprobe.go:190: [WARN] Health probe docmost: HTTP GET :3000/ → Get "http://docmost-postgres:3000/": dial tcp: lookup docmost-postgres on 127.0.0.11:53: no such host
2026/09/25 11:11:38 [INFO] [stacks] SaveAppConfig: saved config for docmost
2026/09/25 11:11:38 pin.go:93: [INFO] [stacks] pin docmost: docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:16-alpine, docmost-redis=redis:7-alpine
2026/09/25 11:11:44 healthprobe.go:190: [WARN] Health probe docmost: HTTP GET :3000/ → Get "http://docmost-postgres:3000/": dial tcp: lookup docmost-postgres on 127.0.0.11:53: no such host
2026/09/25 11:12:05 [INFO] [stacks] SaveAppConfig: saved config for docmost
2026/09/25 11:12:05 unattended.go:490: [INFO] [stacks] update docmost: step docmost=docmost/docmost:0.96.0, docmost-postgres=postgres:18-alpine, docmost-redis=redis:7-alpine recorded as undone — the automatic leg will not press it again until the catalog's ladder changes (ladder 1cdde4765ba33137)
@@ -0,0 +1,29 @@
{
"app": "docmost",
"venue": "box 9202 (drill catalog, controller 0.273.0-rc1), the product's guarded Update",
"from": {
"docmost": "docmost/docmost:0.96.0",
"docmost-postgres": "postgres:16-alpine",
"docmost-redis": "redis:7-alpine"
},
"to": {
"docmost": "docmost/docmost:0.96.0",
"docmost-postgres": "postgres:18-alpine",
"docmost-redis": "redis:7-alpine"
},
"verdict": "proven",
"seed_read_before": true,
"seed_read_after": true,
"healthy_after": true,
"engine_conversion": {
"service": "docmost-postgres",
"engine": "postgres",
"from": 16,
"to": 18,
"result": "converted",
"controller_line": "CONVERTED 16 → 18 in 9.942s — the check is equal (2 database(s), 48 table(s), 71 row(s)), PG_VERSION 18"
},
"duration_s": 42.1,
"measured_at": "2026-09-25T11:13:34Z",
"evidence": "felhom.eu/documentation/audits/night-2026-09-26/D/D1-happy-path.txt"
}
@@ -0,0 +1,12 @@
12:40:28 before:
2026/09/25 09:31:03 scheduler.go:132: [INFO] [scheduler] Daily job offsite-proof scheduled for 2026-09-26 05:30 CEST
2026/09/25 09:31:03 scheduler.go:132: [INFO] [scheduler] Daily job metrics-prune scheduled for 2026-09-26 04:00 CEST
2026/09/25 09:31:03 scheduler.go:132: [INFO] [scheduler] Daily job fill-watch scheduled for 2026-09-26 03:30 CEST
12:40:28 POST /backups/window window_start=12:44 -> 303
12:40:34 after:
2026/09/25 10:40:28 auth.go:142: [DEBUG] [web] auth: valid session for POST /backups/window
2026/09/25 10:40:28 server.go:545: [DEBUG] [web] ServeHTTP: POST /backups/window from 172.18.0.6:60864
2026/09/25 10:40:28 scheduler.go:172: [INFO] [scheduler] Daily job db-dump rescheduled 02:30 → 12:44 (next run 2026-09-25 12:44 CEST)
2026/09/25 10:40:28 scheduler.go:172: [INFO] [scheduler] Daily job tier2-backup rescheduled 03:30 → 13:44 (next run 2026-09-25 13:44 CEST)
2026/09/25 10:40:28 scheduler.go:172: [INFO] [scheduler] Daily job offbox-backup rescheduled 04:15 → 14:29 (next run 2026-09-25 14:29 CEST)
2026/09/25 10:40:28 backup_handlers.go:65: [INFO] [web] backup window set to 12:44 (legs 12:44/13:44/14:29)
@@ -0,0 +1,68 @@
Fri Sep 25 10:40:56 UTC 2026
2026/09/25 10:38:19 router.go:946: [ERROR] [api] Remove failed for nextcloud: A(z) /mnt/felhom-drives/scratch_hdd/userdata/nextcloud tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó viss
2026/09/25 10:38:19 router.go:907: [INFO] [api] Remove requested for stack: nextcloud
2026/09/25 10:38:19 delete.go:597: [INFO] Removing deployed stack: nextcloud (removeHDDData=false, hddDeclared=true, backupPaths=2)
2026/09/25 10:38:19 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/25 10:38:47 delete.go:650: [INFO] [stacks] RemoveStack nextcloud: removed volume(s) [nextcloud_nextcloud_db_data nextcloud_nextcloud_html nextcloud_nextcloud_redis_data]
2026/09/25 10:38:47 delete.go:749: [INFO] Stack nextcloud removed successfully (took 27.4s)
2026/09/25 10:39:01 router.go:431: [INFO] [api] Deploy requested for stack: nextcloud
2026/09/25 10:39:01 [INFO] [stacks] SaveAppConfig: saved config for nextcloud
2026/09/25 10:39:01 deploy.go:391: [INFO] [stacks] Deploying stack nextcloud with 7 env vars: [DOMAIN, SUBDOMAIN, DB_PASSWORD, MYSQL_ROOT_PASSWORD, NEXTCLOUD_ADMIN_USER, NEXTCLOUD_ADMIN_PASSWORD, HDD_PATH]
2026/09/25 10:39:01 manager.go:1592: [INFO] [stacks] Deploying stack nextcloud — checking 3 images...
2026/09/25 10:39:13 deploy.go:481: [INFO] [stacks] Stack nextcloud deployed successfully (took 11.2s)
2026/09/25 10:39:13 [INFO] [stacks] SaveAppConfig: saved config for nextcloud
2026/09/25 10:39:13 [INFO] [stacks] SaveAppConfig: saved config for nextcloud
2026/09/25 10:39:13 installed.go:420: [INFO] [stacks] installed-images nextcloud: recorded 3 service(s) (nextcloud=nextcloud:34.0.4-apache (sha256:a5ace30c695a…), nextcloud-db=mariadb:12.3 (sha256:805c8e104bd5…), nextcloud-red
2026/09/25 10:39:13 [INFO] [stacks] SaveAppConfig: saved config for nextcloud
2026/09/25 10:39:13 pin.go:93: [INFO] [stacks] pin nextcloud: nextcloud=nextcloud:34.0.4-apache, nextcloud-db=mariadb:12.3, nextcloud-redis=redis:7-alpine
2026/09/25 10:39:16 manager.go:1553: [INFO] [stacks] Stack nextcloud post-start status:
2026/09/25 10:39:16 manager.go:1556: [INFO] [stacks] nextcloud nextcloud:34.0.4-apache@sha256:a5ace30c695afe48c2c406e940ee7886a81e13fa382e57cd68b2416d1a66914c running Up 3 seconds (health: starting)
2026/09/25 10:39:16 manager.go:1556: [INFO] [stacks] nextcloud-db mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1 running Up 14 seconds (healthy)
2026/09/25 10:39:16 manager.go:1556: [INFO] [stacks] nextcloud-redis redis:7-alpine@sha256:858f009f9709ce576febc734aa78b8f6d624b82571f9ddb6bda4377c833b3499 running Up 14 seconds (healthy)
2026/09/25 10:40:28 scheduler.go:172: [INFO] [scheduler] Daily job db-dump rescheduled 02:30 → 12:44 (next run 2026-09-25 12:44 CEST)
66548367016e nextcloud-db nextcloud mariadb:12.3
5435a635557b nextcloud-redis nextcloud redis:7-alpine
2026/09/25 10:40:28 [INFO] [stacks] ParseComposeHDDMounts: found 1 HDD mounts for /opt/docker/stacks/nextcloud/docker-compose.yml
2026/09/25 10:40:28 recovery_unit.go:235: [INFO] [backup] Recovery unit captured for nextcloud → /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud (images=3, secrets-referenced=3, data_keys=0, portable-carried=2/2, withhe
/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud:
total 16
drwxr-xr-x 3 root root 4096 10:40:28 .
drwxr-xr-x 3 root root 4096 10:40:28 ..
drwxr-xr-x 2 root root 4096 10:40:28 compose
-rw-r--r-- 1 root root 1439 10:40:28 manifest.json
## after the db-dump leg
2026/09/25 10:44:23 backup.go:824: [INFO] [backup] Volume dump: nextcloud/nextcloud_nextcloud_db_data → 164.5 MB
2026/09/25 10:44:28 backup.go:824: [INFO] [backup] Volume dump: nextcloud/nextcloud_nextcloud_html → 761.6 MB
2026/09/25 10:44:28 backup.go:824: [INFO] [backup] Volume dump: nextcloud/nextcloud_nextcloud_redis_data → 187.5 KB
2026/09/25 10:44:28 backup.go:933: [INFO] [backup] Restarting nextcloud after volume dump
2026/09/25 10:44:28 manager.go:1207: [INFO] [stacks] Starting stack: nextcloud
2026/09/25 10:44:35 manager.go:1222: [INFO] [stacks] Stack nextcloud started successfully (took 6.2s)
2026/09/25 10:44:38 manager.go:1553: [INFO] [stacks] Stack nextcloud post-start status:
2026/09/25 10:44:38 manager.go:1556: [INFO] [stacks] nextcloud nextcloud:34.0.4-apache@sha256:a5ace30c695afe48c2c406e940ee7886a81e13fa382e57cd68b2416d1a66914c running Up 3 seconds (healt
2026/09/25 10:44:38 manager.go:1556: [INFO] [stacks] nextcloud-db mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1 running Up 9 seconds (healt
2026/09/25 10:44:38 manager.go:1556: [INFO] [stacks] nextcloud-redis redis:7-alpine@sha256:858f009f9709ce576febc734aa78b8f6d624b82571f9ddb6bda4377c833b3499 running Up 9 seconds (healt
2026/09/25 10:44:56 recovery_unit.go:235: [INFO] [backup] Recovery unit captured for nextcloud → /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud (images=3, secrets-referenced=3, data_keys=0
2026/09/25 10:44:56 scheduler.go:363: [INFO] [scheduler] Job db-dump completed (took 56.675s)
-rw-r--r-- 1 root root 1587 10:44:56 /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud/manifest.json
/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud/compose:
total 32
drwxr-xr-x 2 root root 4096 10:44:56 .
drwxr-xr-x 5 root root 4096 10:44:56 ..
-rw-r--r-- 1 root root 9180 10:44:56 .felhom.yml
-rw------- 1 root root 611 10:44:56 app.yaml
-rw-r--r-- 1 root root 4809 10:44:56 docker-compose.yml
/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud/db-dumps:
total 864
drwxr-xr-x 2 root root 4096 10:44:00 .
drwxr-xr-x 5 root root 4096 10:44:56 ..
-rw-r--r-- 1 root root 873120 10:44:00 nextcloud-mariadb.sql
/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud/volume-dumps:
total 948440
drwxr-xr-x 2 root root 4096 10:44:28 .
drwxr-xr-x 5 root root 4096 10:44:56 ..
-rw-r--r-- 1 root root 172453376 10:44:22 nextcloud_nextcloud_db_data.tar
-rw-r--r-- 1 root root 798543360 10:44:26 nextcloud_nextcloud_html.tar
-rw-r--r-- 1 root root 192000 10:44:28 nextcloud_nextcloud_redis_data.tar
@@ -0,0 +1,12 @@
12:45:23 before:
2026/09/25 10:40:28 scheduler.go:172: [INFO] [scheduler] Daily job db-dump rescheduled 02:30 → 12:44 (next run 2026-09-25 12:44 CEST)
2026/09/25 10:40:28 scheduler.go:172: [INFO] [scheduler] Daily job tier2-backup rescheduled 03:30 → 13:44 (next run 2026-09-25 13:44 CEST)
2026/09/25 10:40:28 scheduler.go:172: [INFO] [scheduler] Daily job offbox-backup rescheduled 04:15 → 14:29 (next run 2026-09-25 14:29 CEST)
12:45:23 POST /backups/window window_start=02:30 -> 303
12:45:29 after:
2026/09/25 10:45:23 auth.go:142: [DEBUG] [web] auth: valid session for POST /backups/window
2026/09/25 10:45:23 server.go:545: [DEBUG] [web] ServeHTTP: POST /backups/window from 172.18.0.6:39110
2026/09/25 10:45:23 scheduler.go:172: [INFO] [scheduler] Daily job db-dump rescheduled 12:44 → 02:30 (next run 2026-09-26 02:30 CEST)
2026/09/25 10:45:23 scheduler.go:172: [INFO] [scheduler] Daily job tier2-backup rescheduled 13:44 → 03:30 (next run 2026-09-26 03:30 CEST)
2026/09/25 10:45:23 scheduler.go:172: [INFO] [scheduler] Daily job offbox-backup rescheduled 14:29 → 04:15 (next run 2026-09-26 04:15 CEST)
2026/09/25 10:45:23 backup_handlers.go:65: [INFO] [web] backup window set to 02:30 (legs 02:30/03:30/04:15)
@@ -0,0 +1,37 @@
('200', {'ok': True, 'message': 'Stack nextcloud stop completed'})
remove keep data + keep backups -> 200 {'ok': True, 'data': {'removed': 'nextcloud', 'volumes_removed': ['nextcloud_nextcloud_db_data', 'nextcloud_nextcloud_html', 'nextcloud_nextcloud_redis_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': ['/mnt/felhom-drives/scratch_hdd/appdata/nextcloud (126M)'], 'verified': True}, 'message': 'Stack nextcloud removed'}
total 28
drwxrwx--- 5 www-data www-data 4096 Sep 25 10:39 .
drwxr-xr-x 4 root root 4096 Sep 25 10:39 ..
-rw-rw-r-- 1 www-data www-data 542 Sep 25 10:39 .htaccess
-rw-rw-r-- 1 www-data www-data 52 Sep 25 10:39 .ncdata
drwxr-xr-x 3 www-data www-data 4096 Sep 25 10:39 admin
drwxr-xr-x 4 www-data www-data 4096 Sep 25 10:39 appdata_ocrpy9goog25
drwxr-xr-x 4 www-data www-data 4096 Sep 25 10:40 drilld112bd
-rw-rw-r-- 1 www-data www-data 0 Sep 25 10:39 index.html
-rw-rw-r-- 1 www-data www-data 0 Sep 25 10:39 nextcloud.log
/mnt/felhom-drives/scratch_hdd/appdata/nextcloud/admin/files:
Documents
Nextcloud Manual.pdf
Nextcloud intro.mp4
Nextcloud.png
Photos
Readme.md
Reasons to use Nextcloud.pdf
Templates
Templates credits.md
/mnt/felhom-drives/scratch_hdd/appdata/nextcloud/drilld112bd/files:
Documents
Nextcloud Manual.pdf
Nextcloud intro.mp4
Nextcloud.png
Photos
Readme.md
Reasons to use Nextcloud.pdf
Templates
Templates credits.md
after-backup.txt
before-backup.txt
nextcloud
@@ -0,0 +1,62 @@
# E1 Q1 — the removed-app row on /backups/restore (R-487)
row present: True
csolat kikapcsolva — a központi monitoring nem aktív Rendszermonitor → Biztonsági mentés — Visszaállítás enkisfelhom.hu Visszaállítás Alkalmazás: — Válasszon — Docmost Nextcloud Paperless-ngx <option value="pr
snapshot_id in the row's form: []
12:46:25 [R] restoring nextcloud from snapshot 'helyi' (of 0 offered)
12:46:25 [R] POST /backup/restore -> HTTP/2 302 ['location: /backups/restore?flash=flash.restore.started']
12:46:25 + 0.0s restore (True, None, None)
12:48:02 + 97.2s restore (False, None, None)
12:48:02 [R] after restore: state=degraded hold=None phase=None
{'http': 'HTTP/2 302', 'location': ['location: /backups/restore?flash=flash.restore.started'], 'seconds': 97.2, 'state_after': 'degraded', 'hold_after': None}
12:58:16 app never answered on c-nc/status.php (last rc=0 code=404)
12:58:16 nextcloud: the app never served /status.php
12:58:16 E1-Q1-after-restore: read nextcloud=False
12:58:16 PROPFIND files of drilld112bd -> 404; before-backup.txt visible=False; listing=[]
12:58:16 PROPFIND files of drilld112bd -> 404; after-backup.txt visible=False; listing=[]
state degraded
Error response from daemon: container 96bf44773afc146b3fb41cf9f04b0faeb94a9b2a6fa02411af34a94da4ffa4f4 is not running
/var/lib/docker/volumes/nextcloud_nextcloud_html/_data->/var/www/html /appdata/nextcloud->/var/www/html/data
## E1 Q1 evidence, 2026-09-25T10:58:39Z
### env KEYS in the restored app.yaml (values never printed)
### env KEYS in the unit's app.yaml
env DB_PASSWORD DOMAIN HDD_PATH MYSQL_ROOT_PASSWORD NEXTCLOUD_ADMIN_USER SUBDOMAIN
### nextcloud container mounts
/var/lib/docker/volumes/nextcloud_nextcloud_html/_data -> /var/www/html
/appdata/nextcloud -> /var/www/html/data
total 12
drwxr-xr-x 3 root root 4096 Sep 15 09:17 .
drwxr-xr-x 20 root root 4096 Sep 24 20:52 ..
### nextcloud-db log tail
2026-09-25 12:56:23+02:00 [Note] [Entrypoint]: Entrypoint script for MariaDB Server 1:12.3.3+maria~ubu2404 started.
2026-09-25 12:56:23+02:00 [ERROR] [Entrypoint]: Database is uninitialized and password option is not specified
You need to specify one of MARIADB_ROOT_PASSWORD, MARIADB_ROOT_PASSWORD_HASH, MARIADB_ALLOW_EMPTY_ROOT_PASSWORD and MARIADB_RANDOM_ROOT_PASSWORD
2026-09-25 12:57:23+02:00 [Note] [Entrypoint]: Entrypoint script for MariaDB Server 1:12.3.3+maria~ubu2404 started.
2026-09-25 12:57:24+02:00 [Warn] [Entrypoint]: /sys/fs/cgroup///memory.pressure not writable, functionality unavailable to MariaDB
2026-09-25 12:57:24+02:00 [Note] [Entrypoint]: Switching to dedicated user 'mysql'
2026-09-25 12:57:24+02:00 [Note] [Entrypoint]: Entrypoint script for MariaDB Server 1:12.3.3+maria~ubu2404 started.
2026-09-25 12:57:24+02:00 [ERROR] [Entrypoint]: Database is uninitialized and password option is not specified
You need to specify one of MARIADB_ROOT_PASSWORD, MARIADB_ROOT_PASSWORD_HASH, MARIADB_ALLOW_EMPTY_ROOT_PASSWORD and MARIADB_RANDOM_ROOT_PASSWORD
2026-09-25 12:58:24+02:00 [Note] [Entrypoint]: Entrypoint script for MariaDB Server 1:12.3.3+maria~ubu2404 started.
2026-09-25 12:58:24+02:00 [Warn] [Entrypoint]: /sys/fs/cgroup///memory.pressure not writable, functionality unavailable to MariaDB
2026-09-25 12:58:24+02:00 [Note] [Entrypoint]: Switching to dedicated user 'mysql'
2026-09-25 12:58:24+02:00 [Note] [Entrypoint]: Entrypoint script for MariaDB Server 1:12.3.3+maria~ubu2404 started.
2026-09-25 12:58:25+02:00 [ERROR] [Entrypoint]: Database is uninitialized and password option is not specified
You need to specify one of MARIADB_ROOT_PASSWORD, MARIADB_ROOT_PASSWORD_HASH, MARIADB_ALLOW_EMPTY_ROOT_PASSWORD and MARIADB_RANDOM_ROOT_PASSWORD
### controller restore lines
2026/09/25 10:39:01 deploy.go:391: [INFO] [stacks] Deploying stack nextcloud with 7 env vars: [DOMAIN, SUBDOMAIN, DB_PASSWORD, MYSQL_ROOT_PASSWORD, NEXTCLOUD_ADMIN_USER, NEXTCLOUD_ADMIN_PASSWORD, HDD_PATH]
2026/09/25 10:40:28 recovery_unit.go:235: [INFO] [backup] Recovery unit captured for nextcloud → /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud (images=3, secrets-referenced=3, data_keys=0, portable-carried=2
2026/09/25 10:44:56 recovery_unit.go:235: [INFO] [backup] Recovery unit captured for docmost → /mnt/sys_drive/felhom-data/backups/primary/docmost (images=3, secrets-referenced=2, data_keys=0, portable-carried=2/2, with
2026/09/25 10:44:56 recovery_unit.go:235: [INFO] [backup] Recovery unit captured for nextcloud → /mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud (images=3, secrets-referenced=3, data_keys=0, portable-carried=2
2026/09/25 10:46:25 handlers.go:1707: [WARN] [web] Restore requested (async): stack=nextcloud, snapshot=helyi from 172.18.0.6:39110
2026/09/25 10:46:25 restore_unit.go:306: [WARN] [backup] No readable recovery unit for nextcloud at /mnt/sys_drive/felhom-data/backups/primary/nextcloud — falling back to volume-only restore
2026/09/25 10:46:25 restore.go:49: [INFO] [backup] Starting app-data restore for nextcloud (drive=/mnt/sys_drive)
2026/09/25 10:46:27 manager.go:1488: [ERROR] [stacks] stderr: time="2026-09-25T10:46:25Z" level=warning msg="The \"HDD_PATH\" variable is not set. Defaulting to a blank string."
stderr: time="2026-09-25T10:46:25Z" level=warning msg="The \"HDD_PATH\" variable is not set. Defaulting to a blank string."
2026/09/25 10:46:27 restore.go:77: [WARN] RESTORE could not restart nextcloud after restore: starting stack nextcloud: exit code 1
stderr: time="2026-09-25T10:46:25Z" level=warning msg="The \"HDD_PATH\" variable is not set. Defaulting to a blank string."
2026/09/25 10:48:01 restore.go:91: [WARN] [backup] Restore completed but app health check failed: stack nextcloud did not reach running state within 1m30s after restore
2026/09/25 10:48:01 restore.go:97: [INFO] RESTORE completed: stack=nextcloud
2026/09/25 10:48:01 handlers.go:1719: [INFO] [web] Restore completed (async): stack=nextcloud in 1m35.743249387s (volumes 0/0, dbs 0/0)
@@ -0,0 +1,12 @@
before: deployed False state degraded
('200', {'ok': True, 'message': 'Stack nextcloud stop completed'})
remove (keep data, keep backups) -> 200 {'ok': True, 'data': {'removed': 'nextcloud', 'volumes_removed': ['nextcloud_nextcloud_db_data', 'nextcloud_nextcloud_html', 'nextcloud_nextcloud_redis_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'verified': True}, 'message': 'Stack nextcloud removed'}
Documents Nextcloud Manual.pdf Nextcloud intro.mp4 Nextcloud.png Photos Readme.md Reasons to use Nextcloud.pdf Templates Templates credits.md after-backup.txt before-backup.txt compose
db-dumps
manifest.json
volume-dumps
total 12
drwxr-xr-x 3 root root 4096 Sep 15 09:17 .
drwxr-xr-x 20 root root 4096 Sep 24 20:52 ..
drwxr-xr-x 4 root root 4096 Sep 15 09:17 paperless
@@ -0,0 +1,28 @@
# E1 Q2 — BY HAND: the unit's own volume tars poured into fresh volumes + the kept drive folder, started with the unit's HDD_PATH. 2026-09-25T11:00:44Z
volume db_data restored from the unit's tar
volume html restored from the unit's tar
volume redis_data restored from the unit's tar
Container nextcloud-db Healthy
Container nextcloud Starting
Container nextcloud Started
- installed: true
- version: 34.0.4.1
- versionstring: 34.0.4
- edition:
13:01:14 nextcloud: readback of the seeded user found=True
13:01:14 E1-Q2-hand: read nextcloud=True
13:01:15 PROPFIND files of drilld112bd -> 207; before-backup.txt visible=True; listing=['Nextcloud%20Manual.pdf', 'Nextcloud%20intro.mp4', 'Nextcloud.png', 'Readme.md', 'Reasons%20to%20use%20Nextcloud.pdf', 'Templates%20credits.md', 'before-backup.txt']
13:01:15 PROPFIND files of drilld112bd -> 207; after-backup.txt visible=False; listing=['Nextcloud%20Manual.pdf', 'Nextcloud%20intro.mp4', 'Nextcloud.png', 'Readme.md', 'Reasons%20to%20use%20Nextcloud.pdf', 'Templates%20credits.md', 'before-backup.txt']
## occ files:scan --all
Starting scan for user 2 out of 2 (drilld112bd)
+---------+-------+-----+---------+---------+--------+--------------+
| Folders | Files | New | Updated | Removed | Errors | Elapsed time |
+---------+-------+-----+---------+---------+--------+--------------+
| 11 | 130 | 0 | 4 | 0 | 0 | 00:00:00 |
+---------+-------+-----+---------+---------+--------+--------------+
took 0.6 s
13:01:26 PROPFIND files of drilld112bd -> 207; after-backup.txt visible=True; listing=['Nextcloud%20Manual.pdf', 'Nextcloud%20intro.mp4', 'Nextcloud.png', 'Readme.md', 'Reasons%20to%20use%20Nextcloud.pdf', 'Templates%20credits.md', 'after-backup.txt', 'before-backup.txt']
## teardown of the hand-made stack
Network nextcloud_nextcloud-internal Removing
Network nextcloud_nextcloud-internal Removed
0
@@ -0,0 +1,21 @@
# baseline on 0.272.0: install nextcloud over the kept folder (R-657)
13:01:53 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/nextcloud']
13:01:53 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['NEXTCLOUD_ADMIN_PASSWORD', 'HDD_PATH']
13:01:53 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
13:02:13 [1] deployed, controller state=running, pinned={'nextcloud': 'nextcloud:34.0.4-apache', 'nextcloud-db': 'mariadb:12.3', 'nextcloud-redis': 'redis:7-alpine'}
True
[Fri Sep 25 13:01:59.795525 2026] [mpm_prefork:notice] [pid 1:tid 1] AH00163: Apache/2.4.68 (Debian) PHP/8.5.10 configured -- resuming normal operations
[Fri Sep 25 13:01:59.795564 2026] [core:notice] [pid 1:tid 1] AH00094: Command line: 'apache2 -D FOREGROUND'
127.0.0.1 - - [25/Sep/2026:13:02:04 +0200] "GET /status.php HTTP/1.1" 200 1064 "-" "curl/8.14.1"
127.0.0.1 - - [25/Sep/2026:13:02:34 +0200] "GET /status.php HTTP/1.1" 200 1064 "-" "curl/8.14.1"
127.0.0.1 - - [25/Sep/2026:13:03:05 +0200] "GET /status.php HTTP/1.1" 200 1064 "-" "curl/8.14.1"
- installed: true
- version: 34.0.4.1
- versionstring: 34.0.4
# Q3: remove with backups ALSO deleted, data kept
backup paths the remove dialog offers: ['/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud']
remove -> 200 {'ok': True, 'data': {'removed': 'nextcloud', 'volumes_removed': ['nextcloud_nextcloud_db_data', 'nextcloud_nextcloud_html', 'nextcloud_nextcloud_redis_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': ['/mnt/felhom-drives/scratch_hdd/appdata/nextcloud (126M)'], 'backup_paths_removed': ['/mnt/felhom-drives/scratch_hdd/backups/primary/nextcloud (928M)'], 'verified': True}, 'message': 'Stack nextcloud removed'}
126M /mnt/felhom-drives/scratch_hdd/appdata/nextcloud
Documents Nextcloud Manual.pdf Nextcloud intro.mp4 Nextcloud.png Photos Readme.md Reasons to use Nextcloud.pdf Templates Templates credits.md after-backup.txt before-backup.txt
removed-app row for nextcloud on /backups/restore: False
@@ -0,0 +1,51 @@
# E1 — kept data: the spike on 9202 (controller 0.272.0, live catalog), 2026-09-25 12:36–13:04 CEST
Written before any build. nextcloud installed through the product with `HDD_PATH=/mnt/felhom-drives/scratch_hdd`
(the drive root, as households have it — demo-hp's apps read `HDD_PATH: /mnt/felhom-drives/hdd_1`), so its files
live at `<drive>/appdata/nextcloud`. Seeded through `occ user:add`; `before-backup.txt` written through WebDAV as that
user; a unit taken by the product's own nightly db-dump leg (window moved to 12:44 by `POST /backups/window`, then back
to 02:30 — `E1-01`, `E1-02`, `E1-03`); `after-backup.txt` written through WebDAV AFTER the unit.
## Q1 — remove „keep my data" (+ backups kept) → the removed-app row → restore: does nextcloud come back whole?
**FALSE on 0.272.0 — it comes back BROKEN.** (`E1-04`, `E1-05`)
- The row is there (R-487: `Nextcloud` in the `/backups/restore` picker), but the picker's snapshot list for it is
EMPTY (`snapshots()` offered 0) — so a household cannot even choose a copy. The form post with `snapshot_id=helyi`
(what the R-487 row sends) was accepted.
- The restore then logged `No readable recovery unit for nextcloud at /mnt/sys_drive/felhom-data/backups/primary/
nextcloud — falling back to volume-only restore` — **the unit sat on the DATA drive**, `.../scratch_hdd/backups/
primary/nextcloud`, fully readable. `compose up` ran WITHOUT the app's env: `The "HDD_PATH" variable is not set`,
the app container mounted `/appdata/nextcloud` on the guest's ROOT disk, `nextcloud-db` crash-looped
(`Database is uninitialized and password option is not specified`), `volumes 0/0, dbs 0/0`, no app.yaml.
- **Root cause (read from source):** `backup.primaryUnitDirFor` and `backup.ListRestorePoints` decide "is this app
removed?" with `stackProvider.GetStackComposePath(name)`, which in production answers **true for every catalog
app** (it asks whether the STACK exists — every template is a stack). The R-487 test's fake answers it for
deployed apps only, so `TestR487_PrimaryUnitDirForNamesTheRemovedUnitWhereItSits` passes while the box fails —
a seam that lies. Neither R-487 branch has ever run on a box for a catalog app on a data drive.
**Filed R-690, fixed in this build** (`isStackDeployed`, the list's own predicate); `/appdata/paperless` on
9202's root (2026-09-15) looks like the same accident, earlier.
## Q2 — a file written after the last backup: visible, or does nextcloud need `occ files:scan`?
**TRUE — it needs the scan.** (`E1-07`, BY HAND — the product route was broken, so the unit's OWN volume tars were
poured into fresh volumes and the stack started with the unit's `HDD_PATH`; the database is what the unit holds.)
After the start: the seeded account reads back; `before-backup.txt` visible; **`after-backup.txt` NOT visible**
though it is on the drive. `occ files:scan --all` (0.6 s, 130 files, 4 updated) → `after-backup.txt` visible.
So nextcloud's template gets `after_load:` = `php occ files:scan --all` as `www-data` in `nextcloud`.
## Q3 — remove with the backups ALSO deleted → files only
**TRUE.** (`E1-08`) `remove_backups` deleted the unit (928 MB); `appdata/nextcloud` (126 MB, both files) stayed; the
removed-app row is gone. So "use my kept data" must be OFF here (no database copy) — the brief's `use-off` sentence.
## And R-657's loop — measured again as a baseline
A fresh install over the kept folder on 0.272.0 did **not** loop this time: `occ status` → `installed: true` after
60 s. The old account (`drilld112bd`) and its files are invisible to the new install — its database knows nothing of
them. So the harm is data that silently stops being reachable, and the loop is one way it can show. Either way a
silent install over kept files is what E2 replaces with the choice.
Teardown: nextcloud removed through the product (keep data, backups deleted); the hand-made stack `down -v`;
`appdata/nextcloud` (126 MB) kept on purpose — it is E5's kept-data fixture. docmost untouched.
Controller on 9202 was 0.272.0 throughout — not interrupted.
@@ -0,0 +1,6 @@
R-690 red-proof, 2026-09-25
(1) primaryUnitDirFor with GetStackComposePath as the removed? test:
FAIL: the restore opens <sys>/felhom-data/backups/primary/nextcloud, want the kept unit <drive>/backups/primary/nextcloud
(2) ListRestorePoints with GetStackComposePath:
FAIL: the picker offers [] (found=true), want the one kept copy on HDD
Restored: ok
@@ -0,0 +1,8 @@
### RED-PROOF R-687 (1): copy the steps with append(nil, …)
--- FAIL: TestR687_EmptyLegReportsEmptySteps (0.00s)
unattended_test.go:473: an empty leg must report steps as [], got {"trigger":"after-offsite","started_at":"0001-01-01T00:00:00Z","ended_at":"0001-01-01T00:00:00Z","deadline":"0001-01-01T00:00:00Z","enabled":false,"done":0,"undone":0,"held":0,"failed":0,"skipped":0,"steps":null}
FAIL
### RED-PROOF R-687 (2): drop the taken-step log line
--- FAIL: TestLeg_FilesMayChangeNeedsAWholeCopy (0.01s)
unattended_test.go:250: a taken files_may_change step must log which whole copy allowed it
FAIL
@@ -0,0 +1,4 @@
### RED-PROOF R-688: the preview drops cloudflare_manual
--- FAIL: TestR688_PreviewNamesCloudflareByHand (0.03s)
customer_delete_test.go:607: preview missing "\"cloudflare_manual\":["
body: {"_x":["the Cloudflare tunnel that serves acme.example (this customer's config carried its own tunnel token)","the DNS records of acme.example (the apex and *.acme.example)"],"claim_present":true,"customer_id":"acme","customer_name":"Teszt","dr_recipe_present":true,"has_config":true,"host_count":1,"hosts":[{"host_id":"acme-01","online":false,"status":"pending"}],"offsite_enabled":false,"offsite_identifier":"","offsite_type":"","one_time_secret":true,"online_host_present":false,"pbs_tenancy_configured":false,"pending_journal":null,"residue":{"app_log_tails":0,"app_telemetry":0,"appliance_registrations":0,"log_tail_requests":0,"notification_prefs":0,"reports":0,"selfbind_tokens":0},"residue_total":0,"superseded_blobs":1}
@@ -0,0 +1,9 @@
A-seed: seed docmost=True
A-read0: read docmost=True
A4-after-undo: read docmost=False
A4-after-undo-retry: read docmost=True
A4-after-undo-productstart: read docmost=True
E1-seed: seed nextcloud=True
E1-seed2: seed nextcloud=True
E1-Q1-after-restore: read nextcloud=False
E1-Q2-hand: read nextcloud=True
@@ -0,0 +1,7 @@
# read 2026-09-25T10:16:49Z — debug ring, lines naming the night's legs
5000 /var/lib/docker/volumes/felhom-controller-data/_data/data/debug-ring.log
{"timestamp":"2026-09-25T09:46:12Z","level":"DEBUG","message":"[scheduler] job status-refresh: execution starting","source":"scheduler.go:67"}
{"timestamp":"2026-09-25T09:46:12Z","level":"DEBUG","message":"[scheduler] job ring-spill: execution starting","source":"scheduler.go:67"}
# app.yaml last_auto_update
# settings app_update + backup window
{'app_update': None, 'backup_window': None, 'backup': None}
@@ -0,0 +1,7 @@
Sep 25 04:57:53 demo-hp felhom-agent[995644]: time=2026-09-25T04:57:53.359+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=local archive=local:backup/felhom-golden-0.236.0.tar.zst landed=2026-09-13T20:17:01Z reason="newest settled archive (landed 2026-09-13T20:17:01Z) has not been proven (last proven archive was a different one)"
Sep 25 04:57:53 demo-hp felhom-agent[995644]: time=2026-09-25T04:57:53.375+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=local:backup/felhom-golden-0.236.0.tar.zst err="reconcile: restore-test extract archive config: proxmox: GET /nodes/demo-hp/vzdump/extractconfig?volume=local%3Abackup%2Ffelhom-golden-0.236.0.tar.zst -> HTTP 403: permission denied at /vms/ (missing privilege VM.Backup)"
-rw-r--r-- 1 root root 654115664 Sep 13 22:16 felhom-golden-0.236.0.tar.zst
3
Sep 24 10:36:30 demo-hp felhom-agent[1195783]: time=2026-09-24T10:36:30.331+02:00 level=ERROR msg="backup: scheduled res
Sep 25 04:57:53 demo-hp felhom-agent[995644]: time=2026-09-25T04:57:53.375+02:00 level=ERROR msg="backup: scheduled rest
Sep 25 10:57:53 demo-hp felhom-agent[995644]: time=2026-09-25T10:57:53.372+02:00 level=ERROR msg="backup: scheduled rest
@@ -0,0 +1,64 @@
# read-only, demo-hp 9201, GET /api/stacks/docmost 200
deployed = true
state = "running"
catalog_images = {"docmost": "docmost/docmost:0.96.0", "docmost-postgres": "postgres:16-alpine", "docmost-redis": "redis:7-alpine"}
installed_images = {'docmost': {'ref': 'docmost/docmost:0.96.0', 'digest': 'sha256:b56947fcfd08aab8fae12a377e1792784786adbf8b96e4281f14ef4fc072685a', 'at': '2026-09-22T09:07:17Z'}, 'docmost-postgres': {'ref': 'postgres:16-alpine', 'digest': 'sha256:721873c34ceb9f8d8fc265984940dc982404c105f19ad51be9fdc5970a6080ea', 'at': '2026-09-22T09:07:17Z'}, 'docmost-redis': {'ref': 'redis:7-alpine', 'digest': 'sha256:858f009f9709ce576febc734aa78b8f6d624b82571f9ddb6bda4377c833b3499', 'at': '2026-09-22T09:07:17Z'}}
pinned_images = {'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'}
deployed_at = 2026-08-31T12:24:41Z
['name', 'meta', 'compose_path', 'state', 'deployed', 'protected', 'orphaned', 'containers', 'app_config', 'deploying', 'updating', 'health_probe', 'last_updated', 'restarting_since', 'template_images', 'catalog_images', 'catalog_digests', 'catalog_tested_at']
# GET /api/stacks/docmost/backup-data 200
{
"stack": "docmost",
"backup_paths": [
{
"path": "/mnt/sys_drive/felhom-data/backups/primary/docmost",
"size_bytes": 134773579,
"size_human": "129M",
"exists": true
},
{
"path": "/mnt/felhom-drives/hdd_1/backups/secondary/docmost",
"size_bytes": 132154188,
"size_human": "127M",
"exists": true
}
],
"has_backups": true
}
# units on the box (ls -la --time-style=full-iso)
/mnt/sys_drive/felhom-data/backups/primary/docmost
/mnt/felhom-drives/hdd_1/backups/secondary/docmost
## /mnt/sys_drive/felhom-data/backups/primary/docmost
total 24
drwxr-xr-x 5 root root 4096 2026-09-25T09:31:43 .
drwxr-xr-x 10 root root 4096 2026-09-23T12:33:30 ..
drwxr-xr-x 2 root root 4096 2026-09-25T09:31:43 compose
drwxr-xr-x 2 root root 4096 2026-09-25T02:15:01 db-dumps
-rw-r--r-- 1 root root 1376 2026-09-25T09:31:43 manifest.json
drwxr-xr-x 2 root root 4096 2026-09-25T02:15:37 volume-dumps
{
"schema_version": 2,
"app_name": "docmost",
"display_name": "Docmost",
"controller_version": "0.272.0",
"created_at": "2026-09-25T09:31:43Z",
"drive": "/mnt/sys_drive",
"namespace_root": "/mnt/sys_drive/felhom-data",
"image_pins": [
"docmost/docmost:0.96.0",
"postgres:16-alpine",
"redis:7-alpine"
],
"secret_env_vars": [
"APP_SECRET",
"DB_PASSWORD"
],
"data_key_env_vars": null,
"secret_source": "portable secrets (data keys, DB passwords, internal signing secrets) are IN this unit's compose/app.yaml (0600); internet-reachable admin logins are NOT,
## /mnt/felhom-drives/hdd_1/backups/secondary/docmost
total 16
drwxr-xr-x 3 root root 4096 2026-08-23T01:30:00 .
drwxr-xr-x 9 root root 4096 2026-09-14T01:30:00 ..
-rw-r--r-- 1 root root 1 2026-09-25T01:30:13 .felhom-tier2-layout
drwxr-xr-x 5 root root 4096 2026-09-24T20:59:16 recovery-unit
@@ -0,0 +1,65 @@
# read 2026-09-25T10:16:45Z — debug ring, lines naming the night's legs
5000 /var/lib/docker/volumes/felhom-controller-data/_data/data/debug-ring.log
{"timestamp":"2026-09-25T05:49:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T05:49:14Z","level":"INFO","message":"[stacks] Status refresh: 4 containers across 56 stacks","source":""}
{"timestamp":"2026-09-25T05:49:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T05:54:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T05:54:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T05:59:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T05:59:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T06:04:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T06:04:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T06:09:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T06:09:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T06:14:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T06:14:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T06:19:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T06:19:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T06:24:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T06:24:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T06:29:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T06:29:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T06:34:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T06:34:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T06:39:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T06:39:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T06:44:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T06:44:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T06:49:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T06:49:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T06:54:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T06:54:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T06:59:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T06:59:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T07:04:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T07:04:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T07:09:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T07:09:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T07:14:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T07:14:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T07:19:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T07:19:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T07:24:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T07:24:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T07:29:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T07:29:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T07:34:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T07:34:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T07:39:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T07:39:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T07:44:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T07:44:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T07:49:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T07:49:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T07:54:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T07:54:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
{"timestamp":"2026-09-25T07:59:14Z","level":"INFO","message":"[scheduler] Running job: offsite-credential-retry","source":""}
{"timestamp":"2026-09-25T07:59:14Z","level":"INFO","message":"[scheduler] Job offsite-credential-retry completed (took 0s)","source":""}
# app.yaml last_auto_update
/opt/docker/stacks/opengist/app.yaml
last_auto_update:
at: "2026-09-25T02:15:46Z"
outcome: done
from:
# settings app_update + backup window
{'app_update': None, 'backup_window': None, 'backup': None}
@@ -0,0 +1,5 @@
# read from a copy of the hub DB (+wal), 2026-09-25 ~12:30 CEST; the copy is deleted after
demo-felhom first seen in report received 2026-09-25 02:29:15 update_leg= {"trigger": "after-offsite", "started_at": "2026-09-25T02:15:46.825918487Z", "ended_at": "2026-09-25T02:16:06.961273956Z", "deadline": "2026-09-25T07:30:00+02:00", "enabled": true, "done": 1, "undone": 0, "held": 0, "failed": 0, "skipped": 0, "steps": [{"app": "opengist", "outcome": "done", "from": {"opengist": "ghcr.io/thomiceli/opengist:1.13"}, "to": {"opengist": "ghcr.io/thomiceli/opengist:1.15"}, "seconds": 20.1}]}
demo-felhom latest report 2026-09-25 10:16:42 controller 0.272.0
demo-hp first seen in report received 2026-09-25 02:29:17 update_leg= {"trigger": "after-offsite", "started_at": "2026-09-25T02:18:41.05475777Z", "ended_at": "2026-09-25T02:18:41.126639901Z", "deadline": "2026-09-25T07:30:00+02:00", "enabled": true, "done": 0, "undone": 0, "held": 0, "failed": 0, "skipped": 0, "steps": null}
demo-hp latest report 2026-09-25 10:16:43 controller 0.272.0
@@ -0,0 +1,47 @@
== felhom-pve
Sep 24 21:34:05 demo-felhom felhom-agent[3559458]: time=2026-09-24T21:34:05.465+02:00 level=INFO msg="audit: gate decision" class=agent_update host=demo-felhom-8363b5 guest="" source=one_shot_job disposition=destructive allowed=true reason=
Sep 24 21:34:05 demo-felhom felhom-agent[3559458]: time=2026-09-24T21:34:05.465+02:00 level=INFO msg="gate decision" class=agent_update guest="" source=one_shot_job disposition=destructive allowed=true reason=signed
Sep 24 21:34:09 demo-felhom felhom-agent[3858230]: time=2026-09-24T21:34:09.787+02:00 level=INFO msg="controller-supervisor: started" interval=30s confirm_sweeps=2 crashloop_max=3 crashloop_window=15m0s slow_crashloop_max=5 slow_crashloop_w
Sep 24 21:36:37 demo-felhom felhom-agent[3860787]: time=2026-09-24T21:36:37.184+02:00 level=INFO msg="backup: restore-test scheduler starting (per-archive due-check)" eval_interval=6h0m0s settle=24h0m0s
Sep 24 21:36:38 demo-felhom felhom-agent[3860787]: time=2026-09-24T21:36:38.416+02:00 level=INFO msg="controller-supervisor: started" interval=30s confirm_sweeps=2 crashloop_max=3 crashloop_window=15m0s slow_crashloop_max=5 slow_crashloop_w
Sep 24 22:51:41 demo-felhom felhom-agent[3860787]: time=2026-09-24T22:51:41.639+02:00 level=INFO msg="audit: gate decision" class=agent_update host=demo-felhom-8363b5 guest="" source=one_shot_job disposition=destructive allowed=true reason=
Sep 24 22:51:41 demo-felhom felhom-agent[3860787]: time=2026-09-24T22:51:41.639+02:00 level=INFO msg="gate decision" class=agent_update guest="" source=one_shot_job disposition=destructive allowed=true reason=signed
Sep 24 22:51:44 demo-felhom felhom-agent[3925782]: time=2026-09-24T22:51:44.757+02:00 level=INFO msg="backup: restore-test scheduler starting (per-archive due-check)" eval_interval=6h0m0s settle=24h0m0s
Sep 24 22:51:45 demo-felhom felhom-agent[3925782]: time=2026-09-24T22:51:45.908+02:00 level=INFO msg="controller-supervisor: started" interval=30s confirm_sweeps=2 crashloop_max=3 crashloop_window=15m0s slow_crashloop_max=5 slow_crashloop_w
Sep 25 04:51:45 demo-felhom felhom-agent[3925782]: time=2026-09-25T04:51:45.379+02:00 level=INFO msg="backup: restore-test evaluated — nothing due" verdicts="felhom-pbs: newest settled archive (landed 2026-09-22T04:12:20Z) is already prov
Sep 25 07:30:31 demo-felhom felhom-agent[3925782]: time=2026-09-25T07:30:31.104+02:00 level=INFO msg="backup: completed" vmid=9201 target=felhom-backup archive=felhom-backup:backup/vzdump-lxc-9201-2026_09_25-07_29_19.tar.zst size_bytes=2752
Sep 25 07:30:31 demo-felhom felhom-agent[3925782]: time=2026-09-25T07:30:31.104+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=felhom-backup job=backup-9201-1790314159465759090 archive=felhom-backup:backup/vzdump-lxc
Sep 25 10:51:45 demo-felhom felhom-agent[3925782]: time=2026-09-25T10:51:45.064+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-backup archive=felhom-backup:backup/vzd
Sep 25 10:51:45 demo-felhom felhom-agent[3925782]: time=2026-09-25T10:51:45.909+02:00 level=INFO msg="janitor: stale-lock sweep deferred — a heavy operation is in flight" busy=restore-test
Sep 25 10:52:58 demo-felhom felhom-agent[3925782]: time=2026-09-25T10:52:58.751+02:00 level=INFO msg="audit: gate decision" class=guest_destroy host=demo-felhom-8363b5 guest=990000 source=one_shot_job disposition=benign allowed=true reason=
Sep 25 10:52:58 demo-felhom felhom-agent[3925782]: time=2026-09-25T10:52:58.751+02:00 level=INFO msg="gate decision" class=guest_destroy guest=990000 source=one_shot_job disposition=benign allowed=true reason=benign
Sep 25 10:53:04 demo-felhom felhom-agent[3925782]: time=2026-09-25T10:53:04.805+02:00 level=INFO msg="backup: scheduled restore-test passed" archive=felhom-backup:backup/vzdump-lxc-9201-2026_09_24-07_27_27.tar.zst duration_s=79.735774907
-rw-r--r-- 1 root root 5840583923 2026-07-28_06:49:30 vzdump-lxc-9201-2026_07_28-06_44_22.tar.zst
-rw-r--r-- 1 root root 16 2026-07-28_06:49:30 vzdump-lxc-9201-2026_07_28-06_44_22.tar.zst.notes
-rw-r--r-- 1 root root 1579 2026-07-28_17:58:35 vzdump-lxc-9201-2026_07_28-17_53_26.log
-rw-r--r-- 1 root root 5929975285 2026-07-28_17:58:34 vzdump-lxc-9201-2026_07_28-17_53_26.tar.zst
-rw-r--r-- 1 root root 16 2026-07-28_17:58:34 vzdump-lxc-9201-2026_07_28-17_53_26.tar.zst.notes
== demo-hp
Sep 24 21:48:57 demo-hp felhom-agent[3717779]: time=2026-09-24T21:48:57.876+02:00 level=INFO msg="audit: gate decision" class=agent_update host=demo-hp-bb76ea guest="" source=one_shot_job disposition=destructive allowed=true reason=signed k
Sep 24 21:48:57 demo-hp felhom-agent[3717779]: time=2026-09-24T21:48:57.876+02:00 level=INFO msg="gate decision" class=agent_update guest="" source=one_shot_job disposition=destructive allowed=true reason=signed
Sep 24 21:49:02 demo-hp felhom-agent[728032]: time=2026-09-24T21:49:02.430+02:00 level=INFO msg="controller-supervisor: started" interval=30s confirm_sweeps=2 crashloop_max=3 crashloop_window=15m0s slow_crashloop_max=5 slow_crashloop_window
Sep 24 21:56:35 demo-hp felhom-agent[752235]: time=2026-09-24T21:56:35.100+02:00 level=INFO msg="backup: restore-test scheduler starting (per-archive due-check)" eval_interval=6h0m0s settle=24h0m0s
Sep 24 21:56:35 demo-hp felhom-agent[752235]: time=2026-09-24T21:56:35.952+02:00 level=INFO msg="controller-supervisor: started" interval=30s confirm_sweeps=2 crashloop_max=3 crashloop_window=15m0s slow_crashloop_max=5 slow_crashloop_window
Sep 24 21:57:44 demo-hp felhom-agent[756657]: time=2026-09-24T21:57:44.968+02:00 level=INFO msg="backup: restore-test scheduler starting (per-archive due-check)" eval_interval=6h0m0s settle=24h0m0s
Sep 24 21:57:45 demo-hp felhom-agent[756657]: time=2026-09-24T21:57:45.838+02:00 level=INFO msg="controller-supervisor: started" interval=30s confirm_sweeps=2 crashloop_max=3 crashloop_window=15m0s slow_crashloop_max=5 slow_crashloop_window
Sep 24 22:06:23 demo-hp felhom-agent[756657]: time=2026-09-24T22:06:23.065+02:00 level=INFO msg="backup: completed" vmid=9201 target=local archive=local:backup/vzdump-lxc-9201-2026_09_24-21_59_25.tar.zst size_bytes=8182759056 uncovered_volu
Sep 24 22:06:23 demo-hp felhom-agent[756657]: time=2026-09-24T22:06:23.065+02:00 level=INFO msg="local-api: backup job complete" vmid=9201 target=local job=backup-9201-1790279965190292701 archive=local:backup/vzdump-lxc-9201-2026_09_24-21_5
Sep 24 22:07:45 demo-hp felhom-agent[756657]: time=2026-09-24T22:07:45.839+02:00 level=INFO msg="janitor: stale-lock sweep deferred — a heavy operation is in flight" busy=backup:felhom-pbs
Sep 24 22:57:49 demo-hp felhom-agent[756657]: time=2026-09-24T22:57:49.433+02:00 level=INFO msg="audit: gate decision" class=agent_update host=demo-hp-bb76ea guest="" source=one_shot_job disposition=destructive allowed=true reason=signed ke
Sep 24 22:57:49 demo-hp felhom-agent[756657]: time=2026-09-24T22:57:49.433+02:00 level=INFO msg="gate decision" class=agent_update guest="" source=one_shot_job disposition=destructive allowed=true reason=signed
Sep 24 22:57:53 demo-hp felhom-agent[995644]: time=2026-09-24T22:57:53.035+02:00 level=INFO msg="backup: restore-test scheduler starting (per-archive due-check)" eval_interval=6h0m0s settle=24h0m0s
Sep 24 22:57:53 demo-hp felhom-agent[995644]: time=2026-09-24T22:57:53.898+02:00 level=INFO msg="controller-supervisor: started" interval=30s confirm_sweeps=2 crashloop_max=3 crashloop_window=15m0s slow_crashloop_max=5 slow_crashloop_window
Sep 25 04:57:53 demo-hp felhom-agent[995644]: time=2026-09-25T04:57:53.359+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=local archive=local:backup/felhom-golden-0.236.0.ta
Sep 25 04:57:53 demo-hp felhom-agent[995644]: time=2026-09-25T04:57:53.375+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=local:backup/felhom-golden-0.236.0.tar.zst err="reconcile: restore-test extract archive config:
Sep 25 10:57:53 demo-hp felhom-agent[995644]: time=2026-09-25T10:57:53.357+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=local archive=local:backup/felhom-golden-0.236.0.ta
Sep 25 10:57:53 demo-hp felhom-agent[995644]: time=2026-09-25T10:57:53.372+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=local:backup/felhom-golden-0.236.0.tar.zst err="reconcile: restore-test extract archive config:
-rw-r--r-- 1 root root 187 2026-09-24_08:22:55 vzdump-lxc-9201-2026_09_24-08_22_55.log
-rw-r--r-- 1 root root 1494 2026-09-24_22:06:19 vzdump-lxc-9201-2026_09_24-21_59_25.log
-rw-r--r-- 1 root root 8182759056 2026-09-24_22:06:15 vzdump-lxc-9201-2026_09_24-21_59_25.tar.zst
-rw-r--r-- 1 root root 16 2026-09-24_22:06:15 vzdump-lxc-9201-2026_09_24-21_59_25.tar.zst.notes
-rw-r--r-- 1 root root 1780 2026-09-24_22:45:37 vzdump-lxc-9201-2026_09_24-22_40_21.log
@@ -0,0 +1,57 @@
"""d1.py <label> [--kill-on-empty] — ONE press of docmost's Update on 9202 (drill catalog), with the
facts before and after: PG_VERSION (the engine's own file), the seed through docmost's front door,
the phases with times, the controller's own conversion lines, the undo copies left. Evidence: D/<label>.txt."""
import json, sys, subprocess, time, threading
import walk as w, fixtures
label = sys.argv[1]
KILL = "--kill-on-empty" in sys.argv
OUT = open(f"{w.EV}/D/{label}.txt", "w", buffering=1)
def say(*a):
w.say(*a); OUT.write(" ".join(map(str, a)) + "\n")
TOK = json.load(open(w.SC + "/seed-tokens.json"))
def pgver():
return w.guest('docker exec docmost-postgres sh -c \'cat "$PGDATA/PG_VERSION"\' 2>&1; docker inspect docmost-postgres --format "{{.Config.Image}}" 2>&1').split()
def seed():
return fixtures.FIXTURES["docmost"].verify(w, "c-docs", TOK["docmost"], lambda m: None)
w.login()
w.sync_rescan("docmost", "postgres:18-alpine")
st = w.stack("docmost")
say(f"# {label} — {time.strftime('%Y-%m-%dT%H:%M:%S%z')}; controller {w.guest('cat /etc/felhom-controller-image').strip()}")
say(f"before: PG_VERSION/image={pgver()} pinned={(st.get('app_config') or {}).get('pinned_images')} catalog={st.get('catalog_images')}")
say(f"before: seed reads back through the front door = {seed()}")
since = w.guest("date -u +%Y-%m-%dT%H:%M:%SZ").strip()
killer = None
if KILL:
def kill():
# wait INSIDE the guest for the controller's own "emptied" line, then SIGKILL its main process
out = w.guest(f"""timeout 600 sh -c 'docker logs -f --since {since} felhom-controller 2>&1 | grep -m1 "emptied"'
pid=$(docker inspect felhom-controller --format '{{{{.State.Pid}}}}'); kill -9 $pid; echo "KILLED controller pid $pid at $(date -u +%T.%N)" """, timeout=700)
say(" [kill] " + out.strip().replace("\n", " | "))
killer = threading.Thread(target=kill); killer.start(); time.sleep(2)
r = w.press_update("docmost", poll=0.5, cap_s=1500)
if killer:
killer.join()
say(" [kill] waiting for the controller to come back and the journal to be resumed")
for i in range(120):
time.sleep(5)
try:
w.login(); s2 = w.stack("docmost")
except SystemExit:
continue
if not s2.get("updating") and s2.get("update_phase") in ("undone", "failed", "done"):
r["after_restart"] = {"phase": s2.get("update_phase"), "error": s2.get("update_error"), "hold": s2.get("hold_reason")}
break
for p in r.get("phases", []):
OUT.write(f" phase +{p['t']}s {p['phase']} | {p['label']} | err={p['error']}\n")
say(f"result: {json.dumps({k: v for k, v in r.items() if k != 'phases'}, ensure_ascii=False)[:900]}")
for i in range(40):
if w.app_curl("c-docs", "/")[1] == "200":
break
time.sleep(5)
st = w.stack("docmost")
say(f"after: PG_VERSION/image={pgver()} pinned={(st.get('app_config') or {}).get('pinned_images')} conversion_copy={(st.get('app_config') or {}).get('conversion_copy')}")
say(f"after: seed reads back through the front door = {seed()}")
say("controller lines (docmost / conversion):")
OUT.write(w.guest(f"docker logs --since {since} felhom-controller 2>&1 | grep -E 'update docmost|CONVERT|pg_dumpall|emptied|loaded into|resum|UNDO' | cut -c1-400") + "\n")
OUT.write("undo copies now: " + w.guest("docker volume ls -q --filter label=felhom.undo-copy-of=docmost | tr '\\n' ' '") + "\n")
say("done")
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,19 @@
#!/usr/bin/env python3
"""Part C step 1 on 9202: move the backup window forward (the product's own form, POST /backups/window) so the
night chain runs now: db-dump at W, Tier 2 at W+60, off-site at W+105 (9202 has none), gate at W+2h."""
import sys, time, json
sys.path.insert(0, ".")
import walk as w
W = sys.argv[1]
w.login()
before = w.guest("docker logs --since 20h felhom-controller 2>&1 | grep 'Daily job' | tail -3").strip()
w.say("before:\n" + before)
sess = open(f"{w.SC}/sess{__import__('os').getpid()}.txt").read().strip()
csrf = open(f"{w.SC}/csrf{__import__('os').getpid()}.txt").read().strip()
r = w.sh(["curl", "-sk", "-o", "/dev/null", "-w", "%{http_code}", "-H", w.HOSTHDR, "-H", f"Cookie: {sess}", "-X", "POST",
"--data-urlencode", f"_csrf={csrf}", "--data-urlencode", f"window_start={W}", f"{w.BASE}/backups/window"])
w.say(f"POST /backups/window window_start={W} -> {r.stdout}")
time.sleep(3)
after = w.guest("docker logs --since 2m felhom-controller 2>&1 | grep -E 'Daily job|window' | tail -6").strip()
w.say("after:\n" + after)
open(sys.argv[2], "w").write("\n".join(w.LOG) + "\n")
@@ -0,0 +1,15 @@
"""ncfile.py put|ls <name> — a file through Nextcloud's OWN WebDAV as the seeded user (front door)."""
import json, sys, subprocess, re
import walk as w
t = json.load(open(w.SC + "/seed-tokens.json"))["nextcloud"]
mode, name = sys.argv[1], sys.argv[2]
base = f"https://192.168.0.114/remote.php/dav/files/{t['uid']}/"
a = ["curl", "-sk", "-u", f"{t['uid']}:{t['pw']}", "-H", "Host: c-nc.enkisfelhom.hu", "-w", "\n%{http_code}"]
if mode == "put":
r = subprocess.run(a + ["-X", "PUT", "--data-binary", f"E1 {name}", base + name], capture_output=True, text=True)
w.say(f"PUT {name} -> {r.stdout.strip().splitlines()[-1]}")
else:
r = subprocess.run(a + ["-X", "PROPFIND", "-H", "Depth: 1", base], capture_output=True, text=True)
code = r.stdout.strip().splitlines()[-1]
names = sorted(set(re.findall(r"<d:href>[^<]*/([^/<]+)</d:href>", r.stdout)))
w.say(f"PROPFIND files of {t['uid']} -> {code}; {name} visible={name in names}; listing={names}")
@@ -0,0 +1,60 @@
#!/usr/bin/env python3
"""Point guest 9202 at the drill catalog (and a 90 s health timeout), or restore the saved config.
`09` §6.5: `git.repo_url` alone is INERT (R-615) — the cache dir must go too. The saved copy is
`controller.yaml.pre-night0926` (NOT the older `.pre-28`, which a restore must never pick up).
"""
import re, sys, io
sys.path.insert(0, '.')
import walk as w
VOL = "/var/lib/docker/volumes/felhom-controller-data/_data"
DRILL_REPO = "https://gitea.dooplex.hu/admin/app-catalog-drill.git"
def creds():
for l in io.open("/home/kisfenyo/.git-credentials").read().strip().split("\n"):
m = re.match(r'https://(admin):([^@]+)@gitea\.dooplex\.hu', l)
if m:
return m.group(1), m.group(2)
raise SystemExit("no admin credential")
def to_drill():
u, t = creds()
print(w.guest(f"""
set -e
test -f {VOL}/controller.yaml.pre-night0926 || cp -p {VOL}/controller.yaml {VOL}/controller.yaml.pre-night0926
python3 - <<'PY'
import re
p = "{VOL}/controller.yaml"
s = open(p).read()
s = re.sub(r'(^\\s+repo_url: ).*$', r'\\g<1>{DRILL_REPO}', s, count=1, flags=re.M)
s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+token: ).*$', r'\\g<1>"{t}"', s, count=1, flags=re.M)
s = re.sub(r'(^git:(?:\\n\\s+.*)*?\\n\\s+username: ).*$', r'\\g<1>"{u}"', s, count=1, flags=re.M)
if not re.search(r'^update:', s, re.M):
s += "update:\\n health_timeout: 90s\\n"
open(p, "w").write(s)
PY
rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache
docker restart felhom-controller >/dev/null
sleep 15
grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: <redacted>/'
grep -A2 '^update:' {VOL}/controller.yaml
"""))
def restore():
print(w.guest(f"""
set -e
cp -p {VOL}/controller.yaml.pre-night0926 {VOL}/controller.yaml
rm -rf {VOL}/catalog-cache {VOL}/data/catalog-cache
docker restart felhom-controller >/dev/null
sleep 15
grep -A6 '^git:' {VOL}/controller.yaml | sed 's/token:.*/token: <redacted>/'
grep -c '^update:' {VOL}/controller.yaml || true
"""))
if __name__ == "__main__":
to_drill() if sys.argv[1] == "drill" else restore()
@@ -0,0 +1,24 @@
"""seed | read — the four drill apps' data through their OWN front doors (fixtures.py, R-156). Tokens are kept in the
scratchpad (they can carry generated credentials); the evidence gets only True/False per app."""
import json, sys, os
import walk as w, fixtures
SUBS = {"wishlist": "c-wish", "navidrome": "c-navi", "romm": "c-romm", "vikunja": "c-vik"}
TOK = w.SC + "/seed-tokens.json"
mode, label = sys.argv[1], sys.argv[2]
w.login()
res = {}
if mode == "seed":
toks = {}
for a, sub in SUBS.items():
toks[a] = fixtures.FIXTURES[a].seed(w, sub, w.say)
res[a] = toks[a] is not None
json.dump(toks, open(TOK, "w"), default=str); os.chmod(TOK, 0o600)
else:
toks = json.load(open(TOK))
for a, sub in SUBS.items():
if len(sys.argv) > 3 and a not in sys.argv[3].split(","):
continue
res[a] = bool(toks.get(a)) and fixtures.FIXTURES[a].verify(w, sub, toks[a], w.say)
line = f"{label}: {mode} " + " ".join(f"{a}={v}" for a, v in res.items())
w.say(line)
open(w.EV + "/C/data-readback.txt", "a").write(line + "\n")
@@ -0,0 +1,19 @@
"""sr.py seed|read <label> app=sub[,app=sub] — seed or read back through each app's OWN front door (fixtures.py).
Tokens stay in the scratchpad; the evidence gets True/False only."""
import json, sys, os
import walk as w, fixtures
mode, label = sys.argv[1], sys.argv[2]
SUBS = dict(x.split("=") for x in sys.argv[3].split(","))
TOK = w.SC + "/seed-tokens.json"
w.login()
toks = json.load(open(TOK)) if os.path.exists(TOK) else {}
res = {}
for a, sub in SUBS.items():
if mode == "seed":
toks[a] = fixtures.FIXTURES[a].seed(w, sub, w.say); res[a] = toks[a] is not None
else:
res[a] = bool(toks.get(a)) and fixtures.FIXTURES[a].verify(w, sub, toks[a], w.say)
if mode == "seed":
json.dump(toks, open(TOK, "w"), default=str); os.chmod(TOK, 0o600)
line = f"{label}: {mode} " + " ".join(f"{a}={v}" for a, v in res.items())
w.say(line); open(w.EV + "/data-readback.txt", "a").write(line + "\n")
@@ -0,0 +1,29 @@
#!/usr/bin/env python3
"""Night 2026-09-25 Part G (adapted from 09-24), the MACHINE layer on 9202: tonight's apps removed THROUGH THE PRODUCT, leftovers read by name,
the backup window put back to 02:30 through the product's own form. Apps present before tonight stay:
privatebin, paperless-ngx, filebrowser (gokapi was removed in chaos round 5 — R-644's disposition)."""
import json, sys, time
sys.path.insert(0, ".")
import walk as w
w.login()
apps = ["wishlist", "navidrome", "romm", "vikunja", "n8n", "mealie"]
q = lambda: w.guest("for p in " + " ".join(apps) + "; do echo \"$p containers=$(docker ps -aq --filter label=com.docker.compose.project=$p | wc -l) volumes=$(docker volume ls -q | grep -c ^${p}_) undo_copies=$(docker volume ls -q --filter label=felhom.undo-copy-of=$p | wc -l)\"; done; docker ps --format '{{.Names}}' | sort | tr '\\n' ' '")
before = q(); w.say("before:\n" + before)
res = {}
for a in apps:
st = w.stack(a)
if not st.get("deployed"):
res[a] = "not deployed"; continue
res[a] = w.remove(a)
after = q(); w.say("after the product's removes:\n" + after)
dirs = w.guest("ls -d /mnt/felhom-drives/scratch_hdd/userdata/*/ 2>/dev/null; ls /mnt/sys_drive/felhom-data/backups/secondary/ 2>/dev/null; ls -d /mnt/felhom-drives/scratch_hdd/userdata/*/backups/secondary/* 2>/dev/null")
w.say("drive dirs left (read by name):\n" + dirs)
sess = open(f"{w.SC}/sess{__import__('os').getpid()}.txt").read().strip()
csrf = open(f"{w.SC}/csrf{__import__('os').getpid()}.txt").read().strip()
r = w.sh(["curl", "-sk", "-o", "/dev/null", "-w", "%{http_code}", "-H", w.HOSTHDR, "-H", f"Cookie: {sess}", "-X", "POST",
"--data-urlencode", f"_csrf={csrf}", "--data-urlencode", "window_start=02:30", f"{w.BASE}/backups/window"])
time.sleep(2)
win = w.guest("docker logs --since 1m felhom-controller 2>&1 | grep -E 'rescheduled|window set' | tail -4")
w.say(f"backup window -> 02:30: {r.stdout}\n{win}")
json.dump({"before": before, "remove": res, "after": after, "dirs": dirs, "window": win}, open("../G/G2-machine-removes.json", "w"), indent=2)
open("../G/G2-machine-removes.log", "w").write("\n".join(w.LOG) + "\n")
@@ -0,0 +1,9 @@
import sys, time, walk as w
head = sys.argv[1]; w.login()
for k in range(40):
w.sync_rescan()
got = w.guest("cd /var/lib/docker/volumes/felhom-controller-data/_data/data/catalog-cache && git rev-parse --short HEAD").strip()
if got.startswith(head[:7]):
print(f"box clone at {got} after {k+1} sync(s)"); sys.exit(0)
time.sleep(8)
sys.exit("box never caught up to " + head)
@@ -0,0 +1,507 @@
#!/usr/bin/env python3
"""walk.py — ONE app's full update walk on guest 9202, through the product's own endpoints.
EVIDENCE, NOT PRODUCT. It presses exactly the buttons a person presses:
POST /api/stacks/<n>/deploy · POST /api/sync · POST /api/stacks/rescan
POST /api/stacks/<n>/update · POST /api/stacks/<n>/remove
and reads GET /api/stacks/<n>. No controller code exists for it.
The walk, per `09` §6.4 and the update-night brief §4:
1 deploy from the DRILL catalog at the LIVE pin
2 seed through the app's OWN front door (R-156: never a volume, never SQL)
3 read the seed back <- control C1; a fixture that cannot prove itself proves nothing
4 „Mentés most"
5 commit the real one-step bump to the DRILL repo, sync, rescan, read the badge in BOTH languages
6 press the guarded Update, record every phase with timestamps
7 read the seed back through the front door
8 the four version observables side by side
9 write the verdict record in `09`'s JSON shape
`inconclusive` is a first-class verdict and is NEVER collapsed into `failed`.
"""
import argparse, json, os, re, subprocess, sys, time
from datetime import datetime, timezone
SC = "/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/4e6cd5c6-d1fa-409b-ab25-405f6b8bb892/scratchpad"
EV = "/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/night-2026-09-26"
DRILL = "/mnt/5_hdd/felhom.eu/drill/app-catalog-drill"
# GUEST=9201 selects demo-hp's hub-enabled guest (the mail proof); default 9202, the scratch guest.
GUEST = os.environ.get("GUEST", "9202")
BASE = os.environ.get("BASE") or {"9202": "https://192.168.0.114", "9201": "https://192.168.0.155"}[GUEST]
DOMAIN = os.environ.get("DOMAIN", "enkisfelhom.hu")
HOSTHDR = f"Host: felhom.{DOMAIN}"
HP = "demo-hp"
LOG = []
def say(*a):
line = " ".join(str(x) for x in a)
ts = datetime.now().strftime("%H:%M:%S")
print(f"{ts} {line}", flush=True)
LOG.append(f"{ts} {line}")
def sh(args, timeout=300, inp=None):
try:
return subprocess.run(args, capture_output=True, text=True, timeout=timeout, input=inp)
except (subprocess.TimeoutExpired, OSError) as e:
return subprocess.CompletedProcess(args, 124, "", f"{e}")
def guest(script, timeout=600):
"""Run a bash script inside guest 9202. Piped as a file — never as an argument (quoting)."""
# ONE TEMP FILE PER CALL (night 2026-09-23): the shared /tmp/w<guest>.sh swapped scripts under
# two concurrent walks (memory: guest-helper-shares-one-tmp-file).
import secrets as _s
t = f"/tmp/w{GUEST}-{os.getpid()}-{_s.token_hex(4)}.sh"
r = sh(["ssh", "-o", "ConnectTimeout=20", "-o", "StrictHostKeyChecking=accept-new", HP,
f"export LC_ALL=C; cat > {t}; pct push {GUEST} {t} {t} >/dev/null 2>&1; "
f"pct exec {GUEST} -- bash {t}; pct exec {GUEST} -- rm -f {t}; rm -f {t}"],
timeout=timeout, inp=script)
return r.stdout or ""
def login():
pw = open(f"{SC}/.ctlpw").read().strip()
sh(["curl", "-sk", "-D", f"{SC}/hdr{os.getpid()}.txt", "-o", "/dev/null", "-H", HOSTHDR,
"-X", "POST", "--data-urlencode", f"password={pw}", f"{BASE}/login"])
h = open(f"{SC}/hdr{os.getpid()}.txt").read()
m = re.search(r"felhom_session=[A-Za-z0-9._-]+", h, re.I)
if not m:
sys.exit("login failed: no session cookie")
open(f"{SC}/sess{os.getpid()}.txt", "w").write(m.group(0))
r = sh(["curl", "-sk", "-L", "-H", HOSTHDR, "-H", f"Cookie: {m.group(0)}", f"{BASE}/"])
c = re.search(r'<meta name="csrf-token" content="([^"]+)"', r.stdout or "")
if not c:
sys.exit("login failed: no csrf token")
open(f"{SC}/csrf{os.getpid()}.txt", "w").write(c.group(1))
def ctl(method, path, data=None, raw=False, tries=2):
"""One controller API call. Re-logs in once on a 302/401 — the controller's session store is
in memory, so any controller restart during the night invalidates it silently."""
for attempt in range(tries):
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
csrf = open(f"{SC}/csrf{os.getpid()}.txt").read().strip()
args = ["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", "-w", "\n%{http_code}"]
if method != "GET":
args += ["-H", f"X-CSRF-Token: {csrf}", "-H", "Content-Type: application/json",
"-X", method]
if data is not None:
args += ["--data", json.dumps(data)]
args.append(f"{BASE}{path}")
r = sh(args)
body, _, code = (r.stdout or "").rpartition("\n")
if code.strip() in ("302", "401") and attempt + 1 < tries:
login()
continue
if raw:
return code.strip(), body
try:
return code.strip(), json.loads(body)
except Exception:
return code.strip(), {"_raw": body[:600]}
return code.strip(), {"_raw": body[:600]}
def page(path):
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
r = sh(["curl", "-sk", "-H", HOSTHDR, "-H", f"Cookie: {sess}", f"{BASE}{path}"])
return r.stdout or ""
def app_curl(sub, path, *extra, method=None, data=None, timeout=45):
"""A call to the APP's own front door on 9202 — the household's route, not ours."""
args = ["curl", "-sSk", "--max-time", str(timeout), "-H", f"Host: {sub}.{DOMAIN}",
"-w", "\n%{http_code}"]
if method:
args += ["-X", method]
if data is not None:
args += ["--data-binary", "@-"]
args += list(extra) + [f"{BASE}{path}"]
r = sh(args, timeout=timeout + 30, inp=data)
body, _, code = (r.stdout or "").rpartition("\n")
return r.returncode, code.strip(), body
def stack(name):
_, d = ctl("GET", f"/api/stacks/{name}")
return (d.get("data") or {}) if isinstance(d, dict) else {}
def wait_app(sub, path="/", want=("200", "302", "303", "401", "403"), tries=60, delay=5):
"""Settling says the container runs; this says the APP answers. Not the same thing."""
last = None
for _ in range(tries):
rc, code, _ = app_curl(sub, path, timeout=15)
last = (rc, code)
if rc == 0 and code in want:
return True
time.sleep(delay)
say(f" app never answered on {sub}{path} (last rc={last[0]} code={last[1]})")
return False
# ------------------------------------------------------------------ the walk
DRIVE = "/mnt/felhom-drives/scratch_hdd/userdata"
# What THIS run generated for a deploy, per app. Deploy secrets are ENCRYPTED AT REST in
# `app.yaml` (`ENC:…`), which is right and which means a fixture cannot read an app's admin
# password back off the box — the household sees it once. So the value the harness itself
# generated is kept here for the life of the run, and nowhere else.
GENERATED = {}
def deploy_values(name, sub):
"""Fill EVERY required deploy field the way the wizard would, by asking the box what this app
asks for — `GET /api/stacks/<n>/deploy-fields` — instead of assuming DOMAIN+SUBDOMAIN.
Measured 2026-09-21: three apps in one batch refused at the deploy with a correct 400 because
a required field was absent — `HDD_PATH` (navidrome, audiobookshelf) and an admin password
(grafana). The refusals happen BEFORE anything is created (`deploy.go:324`), which is the only
reason this was safe to discover by running it (live-probes rule).
A `path` field must name a directory that ALREADY EXISTS (`deploy.go:330`), so one is made on
the scratch drive first — the same act the drive browser performs for a household.
"""
code, d = ctl("GET", f"/api/stacks/{name}/deploy-fields")
fields = (((d.get("data") or {}).get("metadata") or {}).get("deploy_fields")) or []
values = {"DOMAIN": DOMAIN, "SUBDOMAIN": sub}
made = []
for f in fields:
ev, ty = f.get("env_var"), f.get("type")
if ev in values:
continue
# `type: password` is MANDATORY whatever `required` says — `deploy.go:305-312` refuses
# when the caller sends none, deliberately ("the user needs to know their password"),
# while `.felhom.yml` declares `required: false` and the API serves that verbatim. A
# caller that trusts the contract gets a 400. Measured tonight on grafana; filed.
if not f.get("required") and ty != "password":
continue # the controller generates the optional secrets itself
if ty == "path":
p = f"{DRIVE}/{name}"
values[ev] = p
made.append(p)
elif ty in ("secret", "password"):
import secrets as _s
values[ev] = "Drill-" + _s.token_hex(12)
GENERATED.setdefault(name, {})[ev] = values[ev]
elif f.get("default"):
values[ev] = f["default"]
else:
values[ev] = f"drill-{name}"
if made:
guest("mkdir -p " + " ".join(made) + "; ls -ld " + " ".join(made))
say(f" [1] made the drive paths this app requires: {made}")
extra = [k for k in values if k not in ("DOMAIN", "SUBDOMAIN")]
if extra:
say(f" [1] required fields filled beyond DOMAIN/SUBDOMAIN: {extra}")
return values
def deploy(name, sub, extra_values=None):
st = stack(name)
if st.get("deployed"):
say(f" [1] {name} already deployed — reusing")
return True
values = deploy_values(name, sub)
if extra_values:
values.update(extra_values)
code, d = ctl("POST", f"/api/stacks/{name}/deploy", {"values": values})
say(f" [1] deploy -> {code} {str(d)[:120]}")
if code != "202":
return False
# WAIT FOR `deployed`, NOT FOR `running`. Measured 2026-09-21 on tandoor: docker reported the
# container `healthy` while the controller's own state read `unhealthy` — a gate on `running`
# alone therefore times out on an app that is up. The state is RECORDED rather than required;
# the real gate is the fixture's own `wait_app`, which asks whether the APP answers.
seen = None
for _ in range(90):
time.sleep(5)
st = stack(name)
seen = st.get("state")
# `deployed` alone is NOT enough and `state` alone is NOT right. Measured 2026-09-21:
# tandoor reads `unhealthy` while serving (R-618), so gating on "running" hangs; and romm
# read `deployed=True, state=degraded, pinned_images=None` twenty seconds in, i.e. the
# deploy had not finished writing app.yaml. The PIN is the deploy's own completion mark
# (`runComposeDeploy` writes it), so that is what to wait for.
pins = (st.get("app_config") or {}).get("pinned_images")
if st.get("deployed") and pins and seen in ("running", "unhealthy", "degraded"):
say(f" [1] deployed, controller state={seen}, "
f"pinned={(st.get('app_config') or {}).get('pinned_images')}")
if seen != "running":
say(f" [1] NOTE: the controller's own state is {seen!r}, not 'running' — recorded, "
f"not treated as a failure; the fixture's front-door wait is the real gate")
return True
say(f" [1] never became deployed (last controller state={seen!r})")
return False
def backup_now(name):
"""R-648 (2026-09-23): NO whole-box „Mentés most" from a drill, ever.
`POST /api/backup/run` is the only backup endpoint and it is WHOLE-BOX: on 9201 it stopped and
restarted 9 of 10 standing apps twice, and on 9202 it broke a deploy in flight (R-634). The product
has NO per-app backup endpoint (router.go: /backup/run, /backup/tier2 only); the per-app backup
exists only inside the guarded update, whose `backing-up` phase calls RunAppBackupNow for the one
app. So this presses nothing: the update takes the throwaway app's own backup, and says so in its
phase list. A seed written "after the backup" is therefore written before the update's own backup
— the undo's last-second copy is still the one that must bring it back."""
say(f" [4] backup press SKIPPED for {name} (R-648: whole-box only; the update's backing-up phase backs up {name} alone)")
return None
def drill_bump(app, frm, to, service_hint=None):
"""Serialised across concurrent walks: one git working tree, one lock."""
import fcntl
with open(f"{SC}/drill.lock", "w") as lk:
fcntl.flock(lk, fcntl.LOCK_EX)
sh(["git", "-C", DRILL, "pull", "-q", "--rebase", "origin", "main"], timeout=120)
return _drill_bump(app, frm, to, service_hint)
def _drill_bump(app, frm, to, service_hint=None):
"""Commit the edge to the DRILL repo. catalog_since set by hand (the drill repo has no gates).
`frm`/`to` may be comma-separated lists of the SAME length: an app whose own version lives in
two images (adventurelog's backend and frontend) moves both in one edge, while its engine
sidecar stays where it is — `09` §3b Q3's rule is per SERVICE, and an app-half edge must move
every service that carries the app's own version and no others.
"""
comp = f"{DRILL}/templates/{app}/docker-compose.yml"
fy = f"{DRILL}/templates/{app}/.felhom.yml"
s = open(comp).read()
froms = [x.strip() for x in frm.split(",") if x.strip()]
tos = [x.strip() for x in to.split(",") if x.strip()]
if len(froms) != len(tos):
say(f" [5] from/to lists differ in length: {froms} vs {tos}")
return None
for f1, t1 in zip(froms, tos):
if f"image: {f1}" not in s:
say(f" [5] FROM ref not found in compose: {f1}")
return None
s = s.replace(f"image: {f1}", f"image: {t1}")
open(comp, "w").write(s)
f = open(fy).read()
today = datetime.now().strftime("%Y-%m-%d")
f = re.sub(r'^catalog_since:.*$', f'catalog_since: "{today}"', f, count=1, flags=re.M)
open(fy, "w").write(f)
sh(["git", "-C", DRILL, "add", "-A"])
sh(["git", "-C", DRILL, "commit", "-q", "-m", f"DRILL {app}: {frm} -> {to}"])
r = sh(["git", "-C", DRILL, "push", "-q", "origin", "main"], timeout=120)
h = sh(["git", "-C", DRILL, "rev-parse", "--short=12", "HEAD"]).stdout.strip()
say(f" [5] drill commit {h}: {app} {frm} -> {to} (push rc={r.returncode})")
return h
def sync_rescan(expect_app=None, expect_ref=None, tries=12, delay=5):
"""Sync, rescan, and — when told what to expect — WAIT FOR THE BADGE TO CATCH UP.
R-607: `POST /api/sync` answers "nincs valtozas" while the catalog HAS moved, and
`catalog_images` stays stale until a separate rescan. Tonight showed the rescan alone is not
enough either: mealie's badge read "Naprakesz" seconds after its bump was pushed, and the
Update that followed moved nothing and still reported "Frissitve". So when the caller knows
which reference should appear, this polls for it and SAYS HOW LONG IT TOOK — which is the
NUMBER R-607 asks for and has never had.
"""
t0 = time.time()
ctl("POST", "/api/sync")
time.sleep(2)
ctl("POST", "/api/stacks/rescan")
time.sleep(2)
if not expect_app or not expect_ref:
return None
for i in range(tries):
cat = stack(expect_app).get("catalog_images") or {}
if expect_ref in cat.values():
waited = round(time.time() - t0, 1)
if i:
say(f" [sync] the badge needed {waited}s and {i+1} sync+rescan rounds to catch up "
f"to {expect_ref} — R-607's window, measured")
return waited
time.sleep(delay)
ctl("POST", "/api/sync")
time.sleep(1)
ctl("POST", "/api/stacks/rescan")
say(f" [sync] the badge NEVER caught up to {expect_ref} in {round(time.time()-t0,1)}s — "
f"catalog_images = {stack(expect_app).get('catalog_images')}")
return None
def badges(name):
out = {}
for lang, suffix in (("hu", ""), ("en", "?lang=en")):
h = page(f"/apps/{name}{suffix}")
m = re.findall(r'<span class="tag tag-[^"]*"[^>]*title="([^"]*)"[^>]*>([^<]*)<', h)
out[lang] = [{"title": a.strip(), "text": b.strip()} for a, b in m][:3]
return out
def press_update(name, poll=1.0, cap_s=1800):
code, d = ctl("POST", f"/api/stacks/{name}/update")
say(f" [6] Update -> {code} {str(d)[:220]}")
if code not in ("202", "200"):
return {"accepted": False, "http": code, "refusal": d, "phases": [], "duration_s": 0}
phases, seen, t0 = [], None, time.time()
while time.time() - t0 < cap_s:
st = stack(name)
ph = st.get("update_phase")
if ph != seen:
seen = ph
rec = {"t": round(time.time() - t0, 1), "phase": ph,
"label": st.get("update_phase_label"), "updating": st.get("updating"),
"error": st.get("update_error"), "hold": st.get("hold_reason")}
phases.append(rec)
say(f" +{rec['t']:>6.1f}s phase={ph} label={rec['label']} "
f"err={rec['error']} hold={rec['hold']}")
if not st.get("updating") and ph in ("done", "failed", "undone", None) and time.time() - t0 > 3:
break
time.sleep(poll)
st = stack(name)
return {"accepted": True, "http": code, "phases": phases,
"duration_s": round(time.time() - t0, 1),
"final_phase": st.get("update_phase"), "update_error": st.get("update_error"),
"hold_reason": st.get("hold_reason"), "state": st.get("state")}
def observables(name):
st = stack(name)
ac = st.get("app_config") or {}
live = guest(f"""
grep -E '^\\s+image:' /opt/docker/stacks/{name}/docker-compose.yml 2>/dev/null | sed 's/^ *//'
echo '---inspect---'
for c in $(docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'); do
echo -n "$c "; docker inspect "$c" --format '{{{{.Config.Image}}}} running={{{{.State.Running}}}} restarts={{{{.RestartCount}}}}'
done
""")
a, _, b = live.partition("---inspect---")
return {
"pinned_images": ac.get("pinned_images"),
"installed_images": {k: (v.get("ref") if isinstance(v, dict) else v)
for k, v in (ac.get("installed_images") or {}).items()},
"catalog_images": st.get("catalog_images"),
"live_compose_image_lines": [x for x in a.strip().splitlines() if x.strip()],
"docker_inspect": [x for x in b.strip().splitlines() if x.strip()],
}
def app_logs(name, lines=400):
"""The app's own container log, DECODED. The endpoint answers a JSON envelope whose `logs` is
one string with escaped newlines — a scan over the envelope sees a single enormous line and
finds nothing, which reads exactly like "the app printed no migration line" and is not. R-96
rule 3 in a new place: an absent line is not evidence when the instrument cannot see lines."""
code, d = ctl("GET", f"/api/stacks/{name}/logs?lines={lines}")
if isinstance(d, dict):
data = d.get("data")
if isinstance(data, dict) and isinstance(data.get("logs"), str):
return data["logs"]
if isinstance(d.get("_raw"), str):
return d["_raw"]
return str(d)
def write_verdict(rec, appdir):
os.makedirs(appdir, exist_ok=True)
p = os.path.join(appdir, "verdict.json")
json.dump(rec, open(p, "w"), indent=2, ensure_ascii=False)
say(f" [9] verdict {rec['verdict']} -> {p}")
def remove(name):
"""Remove through the PRODUCT, never `docker rm` (live-probes rule). The remove endpoint
refuses a running stack — `409 still running` — so the stop is part of the act, not a tidy-up."""
c1, d1 = ctl("POST", f"/api/stacks/{name}/stop")
say(f" [X] stop -> {c1} {str(d1)[:100]}")
for _ in range(24):
time.sleep(5)
if stack(name).get("state") != "running":
break
code, d = ctl("POST", f"/api/stacks/{name}/remove",
{"remove_hdd_data": True, "remove_backups": True})
say(f" [X] remove (with drive data) -> {code} {str(d)[:160]}")
if code == "409":
# R-442's fail-closed guard: when the storage subsystem cannot RESOLVE the app's drive
# path, the removal is REFUSED and the app is kept rather than half-deleted. On guest 9202
# `/api/disks` answers `agent not configured`, so every app deployed with an HDD_PATH hits
# this. The household's other choice — remove the app, KEEP the data — is accepted, and the
# harness takes it, then tidies its own directory by name at teardown.
say(" [X] refused because the drive path cannot be resolved (R-442, fail-closed and right)"
" — removing the app and KEEPING the drive data instead")
code, d = ctl("POST", f"/api/stacks/{name}/remove",
{"remove_hdd_data": False, "remove_backups": True})
say(f" [X] remove (keeping drive data) -> {code} {str(d)[:160]}")
time.sleep(5)
st = stack(name)
left = guest(f"ls -d /opt/docker/stacks/{name} 2>/dev/null; "
f"docker ps -a --filter label=com.docker.compose.project={name} --format '{{{{.Names}}}}'")
say(f" [X] after remove: deployed={st.get('deployed')} leftovers={left.strip()!r}")
return code
def app_env(name, key):
"""Read one deploy value the CUSTOMER was given (e.g. the generated admin password) from the
app's own `app.yaml`. This is not seeding — it is how the household logs in; the controller
shows them the same value. Data still goes in through the app's own front door."""
out = guest(f"grep -E '^\\s*{key}:' /opt/docker/stacks/{name}/app.yaml 2>/dev/null | head -1")
if ":" in out:
return out.split(":", 1)[1].strip().strip('"').strip("'")
return ""
def snapshots(name):
"""The restorable copies the backups page offers for this app."""
code, d = ctl("GET", f"/api/backup/snapshots?stack={name}")
data = d.get("data") if isinstance(d, dict) else None
if isinstance(data, dict):
for k in ("snapshots", "items", "restore_points"):
if isinstance(data.get(k), list):
return data[k]
return data if isinstance(data, list) else []
def restore(name, snapshot_id=None, wait_s=1200):
"""The household's own way out: the „Visszaállítás a mentésből" button on the backups page.
A FORM post, not an API call — `POST /backup/restore` with `_csrf`, `stack_name`,
`snapshot_id` — because that is the button the sentence tells them to press.
"""
snaps = snapshots(name)
if snapshot_id is None:
if not snaps:
say(f" [R] no restorable copy offered for {name}")
return {"ok": False, "why": "no snapshot offered", "snapshots": snaps}
first = snaps[0]
snapshot_id = first.get("id") or first.get("snapshot_id") or first.get("short_id")
say(f" [R] restoring {name} from snapshot {snapshot_id!r} (of {len(snaps)} offered)")
sess = open(f"{SC}/sess{os.getpid()}.txt").read().strip()
csrf = open(f"{SC}/csrf{os.getpid()}.txt").read().strip()
r = sh(["curl", "-sk", "-D", "-", "-o", "/dev/null", "-H", HOSTHDR, "-H", f"Cookie: {sess}",
"-X", "POST",
"--data-urlencode", f"_csrf={csrf}",
"--data-urlencode", f"stack_name={name}",
"--data-urlencode", f"snapshot_id={snapshot_id}",
f"{BASE}/backup/restore"], timeout=180)
head = (r.stdout or "").split("\n")[0].strip()
loc = [l for l in (r.stdout or "").split("\n") if l.lower().startswith("location:")]
say(f" [R] POST /backup/restore -> {head} {loc[:1]}")
t0 = time.time()
last = None
while time.time() - t0 < wait_s:
code, d = ctl("GET", "/api/backup/restore-status")
dd = d.get("data") or {}
cur = (dd.get("running"), dd.get("phase") or dd.get("state"), dd.get("message"))
if cur != last:
say(f" +{round(time.time()-t0,1):>6.1f}s restore {cur}")
last = cur
if not dd.get("running", False) and time.time() - t0 > 5:
break
time.sleep(2)
st = stack(name)
say(f" [R] after restore: state={st.get('state')} hold={st.get('hold_reason')!r} "
f"phase={st.get('update_phase')}")
return {"ok": True, "snapshot_id": snapshot_id, "snapshots": snaps,
"http": head, "location": loc[:1], "seconds": round(time.time() - t0, 1),
"state_after": st.get("state"), "hold_after": st.get("hold_reason"),
"observables_after": observables(name)}
+3 -1
View File
@@ -675,7 +675,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-463** | **[P2-MEDIUM] The day the catalog moves `postgres:16` to `17`, ELEVEN apps are affected and the container image will NOT perform the conversion — and nothing anywhere records that.** MEASURED 2026-09-06: 11 of the 53 templates carry PostgreSQL — **8 on `postgres:16-alpine`**, 1 on `postgres:15-alpine`, plus `postgis/postgis:16-3.5-alpine` and Immich's own `postgres:16-vectorchord…` build. **A grep of the whole register for `pg_upgrade`, "postgres major" or "postgresql major" returns ZERO** (confirmed this session, and confirmed again before filing). **WHY IT IS NOT THE SAME PROBLEM AS R-459, and this is the point of the row: the two engines fail in OPPOSITE directions.** MariaDB starts anyway and skips the conversion quietly, which is why R-459 went unnoticed until a harness looked. **PostgreSQL REFUSES TO START on a datadir from an older major** — the official image performs no `pg_upgrade` and exits with a message naming both versions. So the Postgres case cannot hide; it will present as eight apps down at once, on the sync after the catalog moves. **DELIBERATELY NOT MEASURED HERE, and saying so is the scope discipline:** R-459's task was scoped to MariaDB, and measuring the Postgres analogue is its own piece of work with its own venue. **This row exists so the gap is a record rather than a sentence in an audit nobody greps.** What would settle it: one edge on the existing harness (`postgres:16-alpine` → `17-alpine`) on a scratch host, which would also exercise the `engine_state_after` field's Postgres probe end to end — it is written but has never run against a real Postgres major. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §7 **-- UPDATE NIGHT 2026-09-21:** **Measured 2026-09-21, both halves.** (a) What a household sees TODAY: the guarded Update of `postgres:16-alpine` to `17-alpine` on docmost ended **`failed` in 5.1 s**, the app stopped and held, **the pin naming 17 while `installed_images` still said 16 and nothing ran**, the data intact, and the restore the hold sentence names back in **29.1 s**. The engine's refusal line had to be REPRODUCED independently because `failAndHold` destroyed it (R-621): *FATAL: database files are incompatible with server / DETAIL: The data directory was initialized by PostgreSQL version 16, which is not compatible with this version 17.11.* The datadir was still `16` afterwards, and the same copy started under 16 holding 48 tables as the positive control. (b) The conversion **COSTED** on a real seeded 49 MB / 48-table datadir: `pg_dumpall` **2.6 s / 132 201 B**, fresh 17 plus replay **6.5 s / 48 tables restored**, the app up on 17 saying *Database connection successful*, **the seeded account read back**, total **155.9 s of which ~9 s is engine work**. `pg_upgrade` was NOT run: it needs both majors' binaries in one image and no such image exists here. Full paragraph: `audits/update-night-2026-09-21/24-Q5-postgres-conversion-costed.md`. **-- RULED 2026-09-23 (`09` §3 decision 16):** PostgreSQL majors are converted BY THE BOX as a guarded-update step — save everything from the old engine, start the new one empty, load it back, check. Each of the eleven apps is proven on the test bench before the catalog may move it; the engine-major gate stays until then. | **READY TO BUILD — owner: CC; `09` §6.4; the gate stays until all eleven are proven** |
| **R-464** | **[P3-LOW] MariaDB's entrypoint prints `MariaDB upgrade not required` on an UNSUPPORTED DOWNGRADE, so that line cannot be used as a soundness signal.** MEASURED 2026-09-06. After converting a datadir to `12.3.3-MariaDB` and then starting **11.6** on it, the entrypoint logs, on every start: **`[Note] [Entrypoint]: MariaDB upgrade not required`**. Asked properly, the same engine answers **`FATAL ERROR: Version mismatch (12.3.3-MariaDB -> 11.6.2-MariaDB): Trying to downgrade from a higher to lower version is not supported!`** **The entrypoint compares the datadir's recorded version against its own and concludes there is nothing to DO. That is true, and it is not a statement that the state is sound.** **THIS IS THIS PROJECT'S MOST-REPEATED CLASS, in a new costume** — the same shape as `CLAUDE.md`'s "presence is not success" and as R-443's HTTP 200 over a crash-looping app: a reassuring sentence that answers a narrower question than the one a reader will take it for. **Why it is worth a row rather than a footnote: the obvious cheap instrument for R-459 is to grep container logs for that exact line**, and such an instrument would report "fine" for an unsupported downgrade. **The correct probe is `mariadb-upgrade --check-if-upgrade-is-needed`**, which is what `upgrade-test.py`'s `engine_state_after` now uses. **Also recorded, because it nearly produced a wrong answer here: run without credentials that command returns `ERROR 1045 … FATAL ERROR: Upgrade failed` with exit 1** — an authentication failure wearing the shape of a verdict. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §5.4 | **READY — rank P3-LOW; owner: CC** |
| **R-468** | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** |
| **R-469** | **[P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse.** Since 2026-09-13 `app-catalog-felhom.eu` `CLAUDE.md` rules that *until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version* (four MariaDB, eleven PostgreSQL services), and `scripts/check-engine-major.py` (fourth row of `catalog_gates.py`, run by `.githooks/pre-push` with the push range) refuses one, naming the rule and this expiry. **Why the rule:** every `mariadb:` sidecar now carries `MARIADB_AUTO_UPGRADE=1` (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. **Honest limit, not re-filed:** the gate needs a parent commit and CI fetches at `--depth 1` — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. **When R-448 ships:** delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. **2026-09-13 — UNBLOCKED, NOT LIFTED.** R-448 shipped in controller v0.237.0/v0.238.0 (slice 4): an update now refuses without a restorable, proven Tier-2 copy, backs up first when it is stale, takes a safety dump, and holds an app that does not come up — the precondition this rule was waiting for. **The rule stays in force until someone deliberately removes it**, which is a separate act (and is worth weighing against R-475: an app with no Tier-2 copy cannot be updated at all, so the guard does not yet cover every app a major engine move would touch). **HALF-LIFTED 2026-09-21, catalog `5ff36d098cbc`.** Slice 4 shipped 2026-09-13, so the rule's own expiry condition is met — **for MariaDB**: the four `mariadb:` sidecars have both halves they need, a verified backup in front of the Update (any tier since v0.239.0) and `MARIADB_AUTO_UPGRADE=1` whose conversion the harness WATCHED run on E3/E3b with the seeded data read back after. **PostgreSQL and MySQL stay refused** — postgres performs no `pg_upgrade` and REFUSES to start on an older major's datadir across eleven templates (R-463); a backup is a route BACK, not a conversion. The refusal text now cites R-463 instead of the shipped R-448. **R-450's second half is enforced in its place:** a MariaDB major must be the ONLY image move in its template in that commit (the bookstack `0b73e5e` shape — two migrations behind one edge). The gate now PRINTS what it allowed, by name — a lifted rule that goes quiet is a lifted rule nobody can audit. Two new decoy cases; two red-proofs, each seen to fail; 40 cases green. **What remains of this row:** the PostgreSQL half, which is R-463's to clear — see `09` §3b **Q5**. | **PARTLY CLOSED 2026-09-21 — MariaDB lifted; the PostgreSQL half stands until R-463** |
| **R-469** | **[P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse.** Since 2026-09-13 `app-catalog-felhom.eu` `CLAUDE.md` rules that *until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version* (four MariaDB, eleven PostgreSQL services), and `scripts/check-engine-major.py` (fourth row of `catalog_gates.py`, run by `.githooks/pre-push` with the push range) refuses one, naming the rule and this expiry. **Why the rule:** every `mariadb:` sidecar now carries `MARIADB_AUTO_UPGRADE=1` (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. **Honest limit, not re-filed:** the gate needs a parent commit and CI fetches at `--depth 1` — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. **When R-448 ships:** delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. **2026-09-13 — UNBLOCKED, NOT LIFTED.** R-448 shipped in controller v0.237.0/v0.238.0 (slice 4): an update now refuses without a restorable, proven Tier-2 copy, backs up first when it is stale, takes a safety dump, and holds an app that does not come up — the precondition this rule was waiting for. **The rule stays in force until someone deliberately removes it**, which is a separate act (and is worth weighing against R-475: an app with no Tier-2 copy cannot be updated at all, so the guard does not yet cover every app a major engine move would touch). **HALF-LIFTED 2026-09-21, catalog `5ff36d098cbc`.** Slice 4 shipped 2026-09-13, so the rule's own expiry condition is met — **for MariaDB**: the four `mariadb:` sidecars have both halves they need, a verified backup in front of the Update (any tier since v0.239.0) and `MARIADB_AUTO_UPGRADE=1` whose conversion the harness WATCHED run on E3/E3b with the seeded data read back after. **PostgreSQL and MySQL stay refused** — postgres performs no `pg_upgrade` and REFUSES to start on an older major's datadir across eleven templates (R-463); a backup is a route BACK, not a conversion. The refusal text now cites R-463 instead of the shipped R-448. **R-450's second half is enforced in its place:** a MariaDB major must be the ONLY image move in its template in that commit (the bookstack `0b73e5e` shape — two migrations behind one edge). The gate now PRINTS what it allowed, by name — a lifted rule that goes quiet is a lifted rule nobody can audit. Two new decoy cases; two red-proofs, each seen to fail; 40 cases green. **What remains of this row:** the PostgreSQL half, which is R-463's to clear — see `09` §3b **Q5**. **-- NARROWED 2026-09-25 (evening), catalog `6a4a5f0`, `09` §3 decision 35:** the PostgreSQL half now passes ONE app at a time — only a template whose ladder entry for the step is proven on BOTH venues and carries `engine_conversion` (the box converts it, controller v0.273.0), as the only image move in its commit. Every other PostgreSQL app stays refused; the postgis family is judged now (it was not). CLAUDE.md rule text updated the same commit. Decoys + red-proof: `audits/night-2026-09-26/C/`. | **PARTLY CLOSED 2026-09-21 — MariaDB lifted; the PostgreSQL half stands until R-463** |
| **R-488** | **[P3-LOW] `go test ./internal/backup` takes 5½ minutes: 89 off-site tests wait on real clocks.** MEASURED 2026-09-13 (`-v` timings, run alone: 581 tests, 333 s in total, 89 of them ≥ 1 s — `TestOffbox*`, `TestOffbox3a*`, `TestOffboxRun*`, `TestR4xx*` reconstitute fixtures at 3–8 s each). The controller's per-commit gate is therefore ~6 minutes, most of it sleeping, and two concurrent runs of the package looked like a hang. **Fix shape:** the waits are `waitForHealthy`-style polls and retry back-offs with fixed durations; make them seams the fixtures shorten (the R-457 rule: one clock). Not a correctness defect. | **READY — rank P3-LOW; owner: CC** |
| **R-489** | **[P3-LOW] `POST /api/stacks/{name}/remove` reports `volumes_removed: null` over named volumes it DID remove.** MEASURED 2026-09-13 on demo-hp five times (gokapi, actualbudget, adventurelog ×2, glance): `docker compose down --volumes` removed the app's named volumes (`docker volume ls` count 2 → 0) and the response carried `"volumes_removed":null`. The customer's confirmation dialog therefore cannot say what it deleted. Split out of R-474 (closed in v0.240.0 for the backups half). **Fix shape:** list the volumes before `down --volumes`, diff after, and report the difference (`[]` when none, never `null`). **PARTLY SHIPPED in v0.242.0 (`d698ce3`), measured live on 9202 the same night:** the difference is computed and a fresh compose-created volume IS reported (`["opengist_opengist_data"]`), but the listing filters on the compose project LABEL and a volume recreated by a unit restore (`docker volume create <name>`, `restore.go:154`) carries no labels — compose still removes it and the response says `[]` (`audits/v0242-2026-09-14/19-R489-cause.txt`). **Remaining fix:** list by the `<project>_` name prefix as well (union), or label the recreated volume as compose would. | **READY — rank P3-LOW; owner: CC (residual)** |
| **R-492** | **[P3-LOW] `cfg.Paths.HDDPath` is empty on every box and still has readers; delete it.** R-465 audited its six readers and found every one falling back; R-490 (v0.242.0) gave the last one, `systemInfo`, the same fallback. The global now carries no information on any box and its deletion was deferred twice. **Fix shape:** remove the field, its env binding and the readers' fallback branches; a build proves nothing reads it. Next controller release. | **READY — rank P3-LOW; owner: CC** |
@@ -810,6 +810,8 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-683** | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** |
| **R-687** | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` | **OPEN — P3; owner: CC** |
| **R-688** | **[P3-LOW] The customer delete says it removes the tunnel and zone, but no leg of it calls Cloudflare.** The dialog's acknowledgement reads "the customer will be RESET — offsite repo DESTROYED, PBS revoked, tunnel/zone removed" (`hub/internal/web/customer_delete.go` `deleteCascadeAcks`), while `commitCustomerReset` has legs for Hetzner, PBS, claim, descriptor and DB only. Seen 2026-09-25 retiring `peti-felhom`, whose config carried a Cloudflare tunnel token and API token (`sajatfelhom.hu`): the tokens went with the record; any tunnel or DNS record on Cloudflare's side was neither listed nor removed. **Fix direction:** either a Cloudflare leg (tunnel + DNS by the customer's ids), or the dialog stops promising it and lists what to remove by hand. `audits/RETIRE-peti-2026-09-25.md` | **READY — P3; owner: CC (hub) / operator (Peti's Cloudflare leftovers, if any)** |
| **R-689** | **[P3-LOW] demo-hp's scheduled restore test picks the golden template in `local:backup/` as a backup and fails every 6 h.** Read 2026-09-25 (agent 0.134.0): `restore-test tier is DUE … archive=local:backup/felhom-golden-0.236.0.tar.zst … newest settled archive (landed 2026-09-13T20:17:01Z)` then `scheduled restore-test FAILED … extractconfig … HTTP 403: permission denied at /vms/ (missing privilege VM.Backup)` — 2026-09-24 10:36, 2026-09-25 04:57 and 10:57. The golden is not a backup of any guest; the real archive (`vzdump-lxc-9201-2026_09_24-21_59_25`) was younger than the 24 h settle. So the restore test proves nothing on demo-hp and logs an ERROR each cycle. **Fix direction:** the restore test selects only `vzdump-*` archives of a known guest (or the golden moves out of `local:backup/`). Agent-side; not built (the agent is untouched this session). `audits/night-2026-09-26/part0/P0-demo-hp-restoretest-golden.txt` | **READY — P3; owner: CC (agent)** |
| **R-690** | **[P1-HIGH] The removed-app restore (R-487) never finds a unit kept on a DATA drive: nextcloud came back with no env, no database and its files mounted on the guest's root disk.** MEASURED 2026-09-25 on 9202 (controller 0.272.0): nextcloud removed with „keep my data" and backups kept; the `/backups/restore` picker listed it but offered 0 copies; the restore logged `No readable recovery unit for nextcloud at /mnt/sys_drive/... — falling back to volume-only restore` while the unit sat readable on the data drive; `compose up` without env (`HDD_PATH` blank → bind `/appdata/nextcloud` on the root disk), `nextcloud-db` crash-looping, `volumes 0/0, dbs 0/0`. **Cause:** `backup.primaryUnitDirFor` and `ListRestorePoints` ask `GetStackComposePath` whether the app is removed; production answers true for EVERY catalog app (the stack exists), while the R-487 test's fake answers it for deployed apps only — the test passes, the box fails. **Fix:** ask `isStackDeployed` (the list's own predicate), pinned by a test on a provider that answers like production. `audits/night-2026-09-26/E/E1-README.md` Q1 | **FIX BUILT — controller (Part E build, 2026-09-25); live proof pending the release** — owner CC |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.
Clearing a row means the check was DONE and its result recorded in that R-row —