Update night 2026-09-21: the full record, twelve rows, and the answers to five of the seven questions
gates / gates (push) Successful in 28s

The drill is complete. Teardown done in three layers plus Gitea; the live catalog's every `image:`
line is proven identical to before.

WHAT WAS MEASURED. 21 edges across 19 apps, on scratch guest 9202 through the product's own
guarded Update, against a PRIVATE DRILL CATALOG so the live catalog carried no test reference at
any point: 14 proven, 3 failed, 4 inconclusive. Each app seeded and read back through its OWN
front door, with a negative control on every readback. Ten of the fourteen printed a verbatim
migration line. Up from the three apps this project had ever measured.

THE RESULT THAT MATTERS. R-618, P1: three of the 53 templates name a health probe the app does not
answer, and because the guarded update WAITS on that same probe, a SUCCESSFUL update ends by
STOPPING a working app. tandoor was measured serving HTTP 200 on the new version at four samples
across five minutes, docker's own healthcheck green, and was then stopped and the household sent
to a restore they did not need. zipline and wger are the same defect, both confirmed live. The
gate that catches all three is static and cheap: both health checks already sit in the same file.

WHAT THE NIGHT ANSWERED that was open. The UNATTENDED HOLD (312.9 s, pressed once, never again) —
which needed a purpose-built image store, because the rule that makes automatic updates safe is
the same rule that refuses the obvious way to break one. MariaDB across a major through the real
button, all four observables, first time. PostgreSQL across a major, refusing exactly as predicted,
with the conversion costed at ~9 s of engine work. There is NO single-flight: five updates ran at
once and all ended honest. And the two EARLY power-cut phases nobody had cut in.

TWELVE NEW ROWS (R-615..R-626), register 303 -> 315, and eight existing rows updated with what was
measured — including two CORRECTIONS: R-606 records the pre-flight refusals as reaching an English
household in English and they do not, and R-446/R-458 are both narrower than their rows state.

Two instrument fixes were needed before anything could be trusted: the unattended caller turned
every success into a timeout (R-623), and one of my own reproductions was wrong and is kept
labelled with what it actually measured.

Interventions: zero. No controller, agent or hub code written. The hub was never touched beyond
the floor the operator asked for.

Gates: repo_gates.py --fast, all 15 OK.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-21 22:34:24 +02:00
parent 9c69b3ff07
commit 8d786f7940
39 changed files with 1657 additions and 128 deletions
+22 -4
View File
@@ -3,18 +3,27 @@
**The full record is `documentation/audits/DRILL-update-night-2026-09-21.md`.** This file is the
session report: what ran, what shipped, what is owed.
*(Filled at the end of the run. `<…>` are placeholders.)*
## Not done, or changed from the brief
<!-- NOTDONE -->
**Nothing in the brief was skipped.** Five things were changed, re-run or measured on a different
venue, each named with its reason in the audit's own first section. In short: the PostgreSQL
rehearsal ran on guest 9202 rather than a separate harness LXC; the `pg_upgrade` route was not run
(it needs an image that does not exist here); B5's `safety-dump` cut MISSED first and was recorded
as a miss before being retried and hit; B8 and the rehearsal were re-run after B1's own precondition
swept the app they needed; and the harness RUNS of the new catalog edges are owed although the code
is shipped.
**One thing the brief asked for that this venue cannot produce at all:** every event and every
customer mail. Guest 9202 runs `hub.enabled: false` and the notifier returns before it logs
(**R-620**). Stated on every row of the alarm truth table rather than left blank.
## What ran
- **Phase 0** — the fleet floor to **0.261.0** (both demo boxes in **13 s**, hub `SERVED … from
declared`); a private **drill catalog** with a positive and two negative controls; a throwaway
**image store** on the scratch guest; capacity measured; the upstream drift re-run.
- **Phase 1** — <N> real within-a-major upstream edges walked on guest 9202 through the product's
- **Phase 1** — 21 edges across 19 apps walked on guest 9202 through the product's
own guarded Update, each seeded and read back through the app's own front door.
- **Phase 2** — the two database engines across a major, through the real Update button.
- **Phase 3** — the bad days, B1–B9.
@@ -33,7 +42,16 @@ session report: what ran, what shipped, what is owed.
## What is owed
<!-- OWED -->
- **The harness RUNS of edges U1–U7.** The code is in the catalog repo and the gates are green; the
runs, and with them the per-app ABORT answers, have not been performed.
- **A cut inside `starting` itself.** Both EARLY phases were cut tonight; `starting` lasts well under
a second and still needs an in-process fault injector rather than a faster shell.
- **The mail half of Q4**, and every event: structurally unmeasurable on this venue (R-620).
- **Fixtures for the four inconclusive apps** — and for two of them (vaultwarden, zipline) the honest
maximum is `inconclusive` while the catalog rightly closes their sign-up (R-624).
- **What re-created the removed `navidrome` container** (R-626): observed, not diagnosed, because the
controller had restarted and its log no longer reached that moment.
- **`wger`'s own edge** — it was deployed only to measure its probe and was then removed.
## The live catalog
+31
View File
@@ -1,5 +1,36 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-09-21 (overnight) — I tested the "update my app" button on as many apps as fit in a night, on good days and bad ones. Fourteen updates are proven safe. Three apps are broken in a way that shuts down a working app, and I would fix that first.**
**Decisions I took on my own: none.** Nothing tonight needed a choice you had not already made.
**The fleet version is 0.261.0.** You asked for that. Both demo machines took it **thirteen seconds** after I saved it. The third machine is switched off and will take it when it comes back.
**What I did.** Twenty-one real updates on a scratch machine, each app installed at the version our catalog has today, filled with real data through the app's own front door, backed up, updated to the newer version that really exists upstream, then the data read back. **Fourteen proven, three failed, four I could not judge.** Ten of the fourteen printed their own "I am rewriting the database" line — so the data really was rewritten, and it still came back.
**The one I would fix first — three apps tell the machine they are broken when they are fine.** Tandoor, Zipline and Wger each have one wrong number or address in their settings file, so the machine knocks on the wrong door and hears nothing. That alone would only be a wrong label. **But the update also waits on that same check** — so when one of these apps updates *successfully*, the machine waits five minutes, decides it failed, **shuts the working app down**, and tells the household to restore from backup. I watched Tandoor serve customers for five minutes on its new version and then get switched off. Nothing is lost and the restore works, but the household loses their app and does work they did not need to do. **A cheap check would catch all three: each of those files already contains the right answer a few lines further down.**
**What broke, and whether the household could get out.**
- **Adventurelog's newer version rewrites the database and then never starts.** The machine did everything right: backup one minute before, waited the full five minutes, stopped the app so the data could not be hurt, and said in one sentence where the copy is, when it was made and what is inside it. I pressed that restore: **back in 75 seconds.** That app must not be moved to the newer version.
- **Tandoor, as above.** Restored in 32 seconds.
- **PostgreSQL will not jump a version.** Exactly as expected: the database engine refuses to start, the app stops honestly, the data is untouched, and the restore brings it back in 29 seconds. I also rehearsed the conversion that would let those eleven apps ever move: **about nine seconds of database work, under three minutes end to end.** That is a maintenance window, not a project.
- **MariaDB, by contrast, jumps a version cleanly** — and I pressed that through the real button for the first time. The engine converted the data, said so in its own words, and took its own backup first.
**The machine also passed every bad day I could invent.** A version that cannot be downloaded: refused in one second, app keeps running. Two updates at once, then five: all ran together and all ended honestly. Power cut in the middle of the backup: the machine recovered itself and said so. Disk nearly full: refused before touching anything. An app left stopped by a failed update stayed stopped after a power cut — *"whatever is holding it owns its recovery."*
**Four smaller faults, all written down.** A message that says "Updated" when nothing was updated. The failure message shown in Hungarian on the English page — including the sentence that tells a household where their files are. An app the household deleted that **came back by itself**, empty. And when an update fails, the machine deletes the broken app's log before anyone can read why.
**Rows opened and closed.** Twelve new, eight existing ones updated with what was measured. The list went from 303 to 315.
**What needs you.**
1. **Rotate the Gitea `admin` token.** The machine stores it in plain text inside its copy of the catalog, so an ordinary diagnostic printed it into my log. *If you do nothing:* the token keeps working and anyone with my session transcript has it.
2. **The promotion list** — fourteen updates proven safe enough to move on the real catalog, and three named that must not move. Moving a version is your call, never mine. *If you do nothing:* nothing breaks; those apps drift further from upstream each month.
3. **The seven questions** about automatic updates now have real facts beside them — including the two that had never been measured: what a stopped app looks like when nobody was watching, and what a database conversion costs. They are still yours. *If you do nothing:* the automatic-update work cannot start, because everything hangs off the first one — *may the machine update apps by itself at night?*
**The live catalog was never touched with a test change.** Not once, not for thirteen minutes. Everything ran against a private drill copy on a scratch machine. The only change to the real catalog is test code, and I have proved every app's version line is identical to before.
---
**Updated 2026-09-21 (evening) — I cut the power to a machine in the middle of an app update, three times, after the new version had already changed the data. It survived every time.**
**The fleet version is now 0.260.0.** You approved it. Both demo machines have it. Three machines are
@@ -102,7 +102,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|---|---|---|---|---|
| Deploy an app from the catalog (env config, memory guard, health-aware progress) | controller, catalog (~52 apps, images pinned) | **PROVEN-LIVE** | `CAMPAIGN-2` T-DEPLOY-SET (7 apps, env config, health-aware); `RERUN-p1p3` (×4 PASS) | Memory-guard FIRING is not live-shown (T-RES-MEMGUARD never fired: ample RAM / auth-walled) — implemented + unit-level only |
| App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE — the ACTIONS work. NARROWED 2026-09-13: `CAMPAIGN-3` proved `remove` removes the APP, not the DATA — the "delete my data" half was INERT on every box until controller v0.236.0 (R-442). RE-PROVEN 2026-09-13 on demo-hp: data written by the app itself (63 MB) gone after removal and listed; an unresolvable data location is REFUSED (409) with the app kept; an SSD app gets `[]` and a note.** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove (app only) live in `CAMPAIGN-3`; **remove WITH data: `audits/R442-2026-09-13/`**; **data behaviour: `audits/SPIKE-app-update-2026-09-01.md` (2026-09-01)** | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical |
| **Update is GUARDED: it refuses without a restorable backup, backs up first when the copy is stale, and HOLDS an app that does not come up — on ANY backup tier, and the release itself arrives by the managed floor** | controller **v0.237.0 + v0.238.0 + v0.238.1 + v0.239.0**, hub **v0.112.0** | **PROVEN-LIVE (2026-09-13, and again the same afternoon for any tier + floor delivery)** — **afternoon (`audits/rulings-r472-r475-2026-09-13/`):** an undeclared floor above the golden refused with nothing stored (02); a declared floor 0.239.0 / MinAgent 0.129.0 served `from declared` and both demo boxes self-updated in 14 s and 15 s (03); nothing on any tier → backed up first, Tier 1 chosen, done (04); gokapi updated on its Tier-1 unit alone (05); a never-healthy update held naming „saját meghajtó" (07); restored from „helyi", hold cleared (08). **Morning:** scenarios A (real upgrade, success only after health), B (stale copy → backup first), E (pull failure → pin back, app untouched), F (never healthy → held, hold text on API and page), H (start/restart/update and the boot sweep all refuse the held app) and **the restore walk** (Mentések unit restore → back on the old version, hold cleared), on demo-hp with a throwaway app | **`audits/slice4-2026-09-13/`** (live/, redproofs/, gates/); design `architecture/09-update-architecture.md` §6.1 | ~~**Tier-2-only precondition**~~ — superseded by v0.239.0 (any tier, R-475 CLOSED); a Tier-1 route back restores only what the unit holds (R-479); the card keeps the failure sentence after a successful restore (R-480); no automatic rollback, by measurement; a release does not reach the fleet by floor between golden bakes (R-472) |
| **Update is GUARDED: it refuses without a restorable backup, backs up first when the copy is stale, and HOLDS an app that does not come up — on ANY backup tier, and the release itself arrives by the managed floor** | controller **v0.237.0 + v0.238.0 + v0.238.1 + v0.239.0**, hub **v0.112.0** | **PROVEN-LIVE (2026-09-13, and again the same afternoon for any tier + floor delivery)** — **afternoon (`audits/rulings-r472-r475-2026-09-13/`):** an undeclared floor above the golden refused with nothing stored (02); a declared floor 0.239.0 / MinAgent 0.129.0 served `from declared` and both demo boxes self-updated in 14 s and 15 s (03); nothing on any tier → backed up first, Tier 1 chosen, done (04); gokapi updated on its Tier-1 unit alone (05); a never-healthy update held naming „saját meghajtó" (07); restored from „helyi", hold cleared (08). **Morning:** scenarios A (real upgrade, success only after health), B (stale copy → backup first), E (pull failure → pin back, app untouched), F (never healthy → held, hold text on API and page), H (start/restart/update and the boot sweep all refuse the held app) and **the restore walk** (Mentések unit restore → back on the old version, hold cleared), on demo-hp with a throwaway app | **`audits/slice4-2026-09-13/`** (live/, redproofs/, gates/); design `architecture/09-update-architecture.md` §6.1 | ~~**Tier-2-only precondition**~~ — superseded by v0.239.0 (any tier, R-475 CLOSED); a Tier-1 route back restores only what the unit holds (R-479); the card keeps the failure sentence after a successful restore (R-480); no automatic rollback, by measurement; a release does not reach the fleet by floor between golden bakes (R-472) **WIDENED 2026-09-21 (the update night) from 3 apps to 21 edges across 19 apps, and NARROWED in one place by the same run.** `audits/DRILL-update-night-2026-09-21.md`. On scratch guest 9202 (controller v0.261.0), against a **private drill catalog** so the live catalog carried no test reference at any point, 21 edges across 19 apps real within-a-major upstream edges were walked through the product's own guarded Update, each app seeded and read back **through its own front door** (R-156) with a negative control on every readback: **14 proven, 3 failed, 4 inconclusive.** **What the PROVEN edges prove, precisely:** the app moved, the four version observables agreed, and the data the app itself was given came back through the app's own interface afterwards. Ten of them printed a verbatim migration line. **What the FAILED edges prove, and they are the more valuable half.** `adventurelog` (a real upstream edge that migrates and then never serves), `tandoor` (an update that SUCCEEDED and was stopped by its own wrong health port), and the PostgreSQL engine major, which refused exactly as predicted. `adventurelog v0.12.1 → v0.13.0` applied **nine database migrations successfully** and then never bound its port; the update held after the full health wait, the hold sentence named the tier, the date and what the copy holds, and the restore the sentence names brought the app back. **That is this row's own promise, exercised on a real upstream edge rather than a staged one.** **AND THE NARROWING, which this row must carry because it is the same mechanism:** the `verifying` phase trusts the `.felhom.yml` probe absolutely, and **two of the 53 templates name a probe the app does not answer** — `tandoor` (port 8080; it listens on 80) and `zipline` (`/api/health`; it answers 404 there, while the compose healthcheck in the same file uses `/api/healthcheck` and is green). For those apps a **successful** update is stopped by its own health wait: tandoor was measured **serving HTTP 200 on the new version at four samples across five minutes**, with docker's own healthcheck green, and was then stopped by `failAndHold` and the household sent to a restore they did not need. **R-618, P1.** No data was lost and the restore works — but "the update is guarded" must not be read as "the guard is right about whether the app came up". **Still true and unchanged:** no automatic rollback (by measurement); the route back is the restore; a multi-major jump ends held honestly. **Not measured on this venue, and named rather than assumed:** every event and every customer mail. Guest 9202 runs `hub.enabled: false` and the notifier returns before it logs (**R-620**), so the whole "who was told" half of `08` was structurally unobservable tonight. |
| **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller **v0.235.0** | **CHANGED 2026-09-06 — they NO LONGER upgrade it.** The row below records what shipped; this text records what it replaced, because every box under v0.235.0 still behaves the old way. **Up to v0.234.0: PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. |
| **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. |
| **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. |
@@ -266,6 +266,25 @@ update **proceeds**, and the household is told what the copy holds.
**If nothing is decided:** Slice 6 must be built for the safe subset only, and the file-leg apps stay
manual — which is the third option by default, without anyone choosing it.
**MEASURED 2026-09-21 (update night).** The hold sentence this question turns on was read verbatim
off a REAL failure rather than from source. `adventurelog v0.12.1 -> v0.13.0` applied nine database
migrations successfully, never bound its port, and held:
> „A(z) adventurelog frissitese 2026-09-21 20:53-kor nem sikerult, es az alkalmazas nem indult el az
> uj verzioval. Az alkalmazas biztonsagi okbol leallitva marad, hogy az adatai ne serüljenek.
> Visszaallithato a Mentesek oldalon ebbol a biztonsagi mentesbol: **sajat meghajto, 2026-09-21 20:47
> — ez a masolat a beallitasokat, az adatbazist es az adatkoteteket tartalmazza.**"
(ASCII fragments here; the live page carries its accents.) So the machinery this question's first
option would key on **exists and works**: the sentence names the tier, the date and **what the copy
holds**, unprompted, on a real edge. Whether the AUTOMATIC rule should differ from the button's is
untouched by that and remains the operator's.
**And one thing Q2 did not ask, which tonight makes urgent: after the hold, nobody can find out WHY.**
`failAndHold` removes the containers, so the failing version's own output is gone within seconds
(**R-621**). With a person pressing, they at least watched it happen.
### Q3 — What counts as "within a major" when the tag is not a version number?
*§3 decision 3 says automatic within a major, never across. What about `postgres:16-alpine`,
@@ -289,6 +308,15 @@ one comparator, never a second one.**
**If nothing is decided:** Slice 6 would have to invent a rule under time pressure, which is how a
major gets automated by accident.
**MEASURED 2026-09-21 by RUNNING the comparator rather than reading it.** `CompareImageRefs` orders a
reference carrying a `host:port/` prefix correctly — `splitImageRef` takes the LAST colon and rejects
it only when a `/` follows, so a registry port is never mistaken for a tag. Four positive cases and
one negative control (different repositories are not orderable). **This is what made the unattended
hold measurable at all**: the drill edge `localhost:5000/drill/glance:1.0.0 -> :1.0.1` PASSES the
within-a-major test and still fails, which no real catalog move does. The recommendation is
unchanged; the *same major?* extension it already names is still owed.
### Q4 — A held app: who is told, when, and does the box try again?
*An automatic update that ends HELD happened while everyone was asleep.*
@@ -319,6 +347,33 @@ within-a-major test and still fails its health check — same repository, same m
starts and does not serve — which probably means a purpose-built image rather than a catalog move.
**So this question still rests on the ATTENDED hold measured in slice 4 (v0.238.0, Scenario F).**
**MEASURED 2026-09-21 (update night) — and this is the half that was missing.** The caller pressed
ONCE with nobody watching; the app held after **312.9 s**; passes 2 and 3 pressed nothing at all
(`outcomes={'glance': ('held', 312.9)} never_again=['glance']`).
| the question | the answer, measured |
|---|---|
| does an unattended update ever produce a HOLD? | **yes** — 312.9 s, the full health wait plus the phases |
| does the box try again? | **no** — two further passes pressed nothing |
| is the household told? | **on the screen, yes** — the app page, and a banner on EVERY authenticated page carrying every held app at once |
| told what? | what happened, when, **which copy** and **what that copy holds** — all four scored True |
| by MAIL? | **still unmeasured** — the scratch guest runs `hub.enabled: false` and the notifier returns before it logs (**R-620**) |
| in ENGLISH? | **no** — the sentence is Hungarian on the English page (**R-606**, confirmed on the hold sentence itself) |
**So the mechanism this question's recommended option rests on is already there and already behaves
that way.** What remains in Q4 is the MAIL and the ENGLISH, not the hold.
**Two further facts this measurement produced, neither of which the question anticipated.**
**(1) There is NO single-flight** — five Updates pressed within 0.45 s all ran at once and all ended
honest, so a caller pressing N apps runs N updates simultaneously. **(2) A held app keeps inviting
the household to update it and the button then refuses** (`409 reason='held'`), even after the
catalog publishes a FIXED newer version — the household's only route out is the restore. Correct per
§6.1, and the page says otherwise (**R-625**).
**Also proven across a genuine power cut:** the boot sweep met a held app after an unclean shutdown
and deliberately left it alone — *„whatever is holding it owns its recovery"*.
### Q5 — PostgreSQL: what has to exist before the catalog may move `postgres:16` to `17`?
*Eleven templates, and the image performs no conversion — it refuses to start on an older major's
@@ -335,6 +390,35 @@ rule's MariaDB half is LIFTED (R-469 — MariaDB has both a backup in front of i
`MARIADB_AUTO_UPGRADE=1`); this half is exactly what stays. **If nothing is decided:** nothing breaks
— the gate refuses the move — but the eleven apps drift further from upstream every month.
**MEASURED 2026-09-21, both halves, on a real seeded datadir.**
**(a) What a household would see today — as predicted, and now observed.** The guarded Update of
`postgres:16-alpine -> 17-alpine` ended **`failed` in 5.1 s**; the app was stopped and held; **the pin
named 17 while `installed_images` still said 16 and nothing was running**; the data was intact; and
the restore the hold sentence names brought it back in **29.1 s**. The engine's refusal had to be
REPRODUCED independently, because `failAndHold` destroyed it before any probe could read it
(**R-621**) — *FATAL: database files are incompatible with server / DETAIL: The data directory was
initialized by PostgreSQL version 16, which is not compatible with this version 17.11.* The datadir
was still `16` afterwards; the positive control (the same copy under 16) started and held 48 tables.
**(b) The conversion rehearsal, COSTED.** Logical dump and restore, 49 MB / 48 tables:
`pg_dumpall` **2.6 s / 132 201 B**; fresh 17 datadir plus replay **6.5 s / 48 tables restored**; the
app up on 17 saying *Database connection successful*; **the seeded account read back**; **total
155.9 s, of which ~9 s is engine work.** For eleven apps that is a maintenance window, not a project.
`pg_upgrade` was NOT run — it needs both majors' binaries in one image and no such image exists in
this project; the logical route may make it unnecessary at this size. Full paragraph:
`audits/update-night-2026-09-21/24-Q5-postgres-conversion-costed.md`.
**(c) A fact about the INSTRUMENT, not the engine.** `upgrade-test.py`'s PostgreSQL probe is
`cat /var/lib/postgresql/data/PG_VERSION` **inside the container**. Against the converted datadir it
answered `17`, exit 0 — it works. **But it is blind in exactly the case that matters**: when
PostgreSQL refuses, the container is not running, so `docker exec` cannot ask it anything. Tonight it
recorded `No such container`, which its own honesty rule covers — but it must never be read as *the
engine is content*.
**The recommendation is unchanged.** Tonight gives it a price rather than a new opinion.
### Q6 — Should the catalog record each pin's DIGEST at push time?
*So the box can tell a moved floating tag from an unmoved one without ever reaching a registry.*
@@ -355,6 +439,23 @@ that has demonstrably moved.
**If nothing is decided:** „Naprakész" keeps meaning "the reference matches", which is measurably not
what it sounds like.
**MEASURED ON A BOX 2026-09-21 (update night, leg B8), and it refines the picture in two ways.**
§8.1's numbers came from a registry sweep on DooPlex; this is the same question asked of a
customer-shaped box, where the badge actually renders. On guest 9202, `docmost`'s two floating pins
were read as `installed_images` records them and compared with the upstream digests measured the same
night: `postgres:16-alpine` -> `sha256:721873c34ceb9…` **on both sides**; `redis:7-alpine` ->
`sha256:858f009f9709c…` **on both sides**. **Identical — so „Naprakesz" is TRUE for this box.**
**(1) The defect's size is set by INSTALL AGE, not by the catalog.** A floating pin is wrong only for
a box that pulled BEFORE the tag moved. R-446's six repushed pins measure the tag against the date
the CATALOG set it, which is the right measure for the catalog and not for a box.
**(2) The producer this question needs ALREADY EXISTS on the box.** `installed_images` records a real
`digest` per service — the box knows exactly what it is running. What it cannot do is COMPARE,
because the catalog carries no digest. That is precisely this question's proposal, and only the
catalog half is missing. **The recommendation is unchanged.**
### Q7 — What does the hub's report need to carry for a fleet view?
*Slice 7 lets the operator SEE and MOVE how far behind every box is.*
@@ -767,6 +868,22 @@ headlessly (R-460).
| E | ~~the automatic night~~ **MOSTLY DONE 2026-09-21 (R-611)** — the success night and the no-retry proof both measured. **What remains: the unattended HOLD**, which needs an edge that passes the within-a-major test and still fails health (see Q4) | ~1 CC-hour + a purpose-built image |
| F | the remaining **38 apps**, through the nightly rotation as decision 6 directs | ~1 app/night; fixtures amortised |
**RE-COSTED 2026-09-21 FROM THE NIGHT'S REAL NUMBERS** (`audits/DRILL-update-night-2026-09-21.md`):
| leg | status after the update night |
|---|---|
| A — the 15 database services | **LARGELY DONE.** 21 edges across 19 apps walked box-side in one night, including both engines and 8 database-carrying apps. **Machine time was never the cost and is now known: a proven edge took 11-218 s, median ~45 s.** The cost was fixtures, exactly as costed — and the real surprise is that two apps can NEVER be seeded headlessly while the catalog rightly closes their sign-up (R-624) |
| B — the power cut | **COMPLETE.** The two EARLY phases nobody had cut in were cut tonight: `backing-up` (the box recovered and said so) and `safety-dump` (nothing moved, nothing to say). Only a cut inside `starting` itself remains, and it still needs an in-process fault injector |
| C — the PostgreSQL rehearsal | **DONE and COSTED**: ~9 s of engine work, 155.9 s end to end for 49 MB / 48 tables. `pg_upgrade` still owed and may prove unnecessary |
| D — the downgrade refusal | already done, v0.260.0 |
| E — the automatic night | **COMPLETE.** The unattended HOLD was produced at last (312.9 s), with no retry across two further passes. It needed the image store of §6.5 |
| F — the remaining apps | ~34 still unwalked. The fixtures for 20 exist and amortise |
**What the night ADDED to this table, which none of the legs anticipated:** the `verifying` phase
trusts the `.felhom.yml` probe absolutely, and three of 53 templates name a probe the app does not
answer — so a SUCCESSFUL update of those apps ends by STOPPING a working app (**R-618**, P1). That is
now the first thing Slice 6 has to be safe against, ahead of everything in this table.
**Total for legs A–E: roughly 21–34 CC-hours**, plus ~25–30 GB of images on a scratch host. Legs C
and E are the ones that unblock a decision; leg A is the one that takes the time.
@@ -774,6 +891,64 @@ and E are the ones that unblock a decision; leg A is the one that takes the time
harness on DooPlex for the image-side ones.
## 6.5 The drill catalog and the image store — the standing method for update drills
**Why this section exists.** On 2026-09-21 an afternoon session put a deliberately broken image into
the LIVE catalog for thirteen minutes to produce a failing edge. It was reverted and nothing reached
a customer, but the method was wrong and the brief that asked for it said so. This is the method that
replaces it, proven the same night.
**The rule, and it has no exception:** *nothing broken, dummy, cross-repo or engine-major ever enters
the live catalog — not as a fallback, not for thirteen minutes. If a leg cannot be done without that,
the leg is skipped and named.*
### The two mechanisms
| | what it is | what it makes possible |
|---|---|---|
| **the drill catalog** | `admin/app-catalog-drill` on Gitea — private, a copy of the live catalog's `main` | a scratch box can be pointed at a catalog where a failing edge is *committable*, because it carries none of the live repo's gates |
| **the image store** | a `registry:2` container on the scratch guest at `127.0.0.1:5000` | an edge that **passes the within-a-major test and still fails** — the one shape a real catalog move cannot produce |
**The image store is not a convenience.** `09` §3b Q4 could not be measured for a year of drills
because the only failing edges available were across-a-major, and the within-a-major rule — correctly
— refuses those before the guarded update is ever reached. *The rule that makes automatic updates
safe is the same rule that refuses the obvious way to break one.* Measuring an unattended HOLD needs
`drill/<app>:X.Y.Z` (the real image, retagged) against `drill/<app>:X.Y.(Z+1)` (a built image that
starts, stays up and never serves) — same repository, same major, plain version tags. A third
flavour, a tag simply **absent** from the store, gives the pull-failure leg.
`stacks.CompareImageRefs` orders a `host:port/` reference correctly: `splitImageRef` takes the last
colon and rejects it only when a `/` follows, so a registry port is never read as a tag. **Proven by
running it**, four positive cases and a negative control, 2026-09-21.
### Pointing a box at the drill catalog — the step that is NOT obvious
**`git.repo_url` alone is inert.** `Syncer.gitCloneOrPull` clones only when the cache has no `.git`;
otherwise it fetches from the remote the clone already stores. The cache directory must be removed as
well, or the box goes on following the live catalog and reports success. Filed as **R-615**; until it
is fixed, the drill procedure is:
1. save `controller.yaml` as `controller.yaml.pre-update-night`;
2. set `git.repo_url` (and `username`/`token` — the drill repo is private);
3. **remove `<data>/catalog-cache`**;
4. restart the controller, sync, **rescan** (R-607: a sync can answer „nincs változás" while the
catalog has moved, and the badge answers from the stale value until the rescan);
5. **three controls, all quoted in the report** — the drill bump appears on the scratch box; the
other boxes' caches are unchanged; the live catalog's `main` hash is unchanged.
### What the drill must leave behind
- `controller.yaml` restored from the saved copy, the controller restarted, and `git.repo_url` **read
back and quoted** as the live catalog.
- The registry container and its volume removed; drill images removed **by name**. Never `prune`.
- The drill repo **kept**, private, reset to the live catalog's `main`, so the next drill starts clean.
- A diff of every `image:` line against the live catalog's `main` — expected: identical.
### The fence
Only a scratch guest is ever pointed at the drill catalog. The drill repo's README says so, and no
customer box has credentials for it. The store listens on the guest's loopback only.
## 7. What slices 1 and 2 actually built
### 7.1 The record (slice 1)
@@ -887,8 +1062,20 @@ Version strings stay in the logs, the API and the hub.
PostgreSQL half (R-463) has no equivalent — the image performs no `pg_upgrade` — and the
engine-major rule (§3 precaution 3, R-469) is what keeps both engines inside their major until
Slice 4 gives the Update button a backup.
8. **Only three of 53 apps have ever had an upgrade measured**, and one of them (bookstack) can only
be half-proven headlessly (**R-460**). The widening is **R-462**, costed with real numbers.
8. ~~**Only three of 53 apps have ever had an upgrade measured.**~~ **WIDENED 2026-09-21 to 21
EDGES ACROSS 19 APPS** (`audits/DRILL-update-night-2026-09-21.md`), on scratch guest 9202
through the product's own guarded Update, against a PRIVATE DRILL CATALOG so the live catalog
carried no test reference at any point: **14 proven, 3 failed, 4 inconclusive**, each app
seeded and read back through its OWN front door with a negative control on every readback.
Ten of the fourteen printed a verbatim migration line. **What stays true:** bookstack is still
only half-provable headlessly (**R-460**), and **two apps cannot be seeded AT ALL** while the
catalog rightly closes their sign-up — vaultwarden (`SIGNUPS_ALLOWED=false`, R-512) and
zipline — which is a permanent ceiling on R-462's scope rather than a fixture nobody has
written (**R-624**). **And one thing this widening FOUND that no count would have:** the
`verifying` phase trusts the `.felhom.yml` probe absolutely, and **three of the 53 templates
name a probe the app does not answer**, so a SUCCESSFUL update of those apps ends by STOPPING
a working app (**R-618**, P1 — tandoor measured serving HTTP 200 on the new version at four
samples across five minutes, then stopped).
9. **The hub does not record image tags at all.** Its report's container payload carries name, state,
CPU and memory, and no image field (spike §5). So the fleet view of §6 slice 7 needs a hub-side
change; it is not derivable from what is already reported.
@@ -12,13 +12,43 @@ the per-edge records, `bad-days/<leg>/` the Phase-3 legs.
*(Filled at the end of the run. Every phase and every B-leg is listed here if it was skipped,
shortened or altered, with the reason. Empty only if true — R-611.)*
<!-- NOTDONE -->
**Nothing in the brief was skipped. Four things were CHANGED or RE-RUN, and one was measured on a
different venue than the brief named — each with its reason.**
| what | what happened | why |
|---|---|---|
| **Phase 2.3, the PostgreSQL rehearsal** | run on **guest 9202 itself**, with plain `docker` beside the product, not on a separate harness LXC | the rehearsal needed the SAME app the 5.2 leg had on a real seeded 16 datadir. No harness LXC was created tonight, so none was destroyed — stated again in the teardown |
| **Phase 2.3, the `pg_upgrade` route** | **NOT run.** The logical dump-and-restore route was run end to end and costed | `pg_upgrade` needs both majors' binaries in one image and no such image exists in this project. Building it is the work Q5's first option is really asking for; naming it costs nothing, and the logical route may make it unnecessary at this size |
| **B5's `safety-dump` cut** | **MISSED on the first attempt and recorded as a MISS**, then retried with a real pending edge and HIT | the first attempt's app was level with the catalog, so the update failed in 0.473 s and `safety-dump` was never observed. A miss recorded as a miss, then fixed |
| **B8, and Phase 2.3's first attempt** | **re-run** after B1's own precondition swept the app they needed | B1 removes every other behind-app so the unattended caller has exactly one thing to react to. That is correct and is recorded; it also removed `docmost`. Re-run in `phase2_redo.sh` |
| **The harness runs of the new catalog EDGES (U1–U7)** | **code shipped, runs OWED** | the box-side result for each of those edges exists; the harness adds the ABORT step, and setting up `/opt/upg` was not worth the last hour against the teardown |
**And one thing the brief asked for that this VENUE cannot produce at all, named rather than left
blank: every event and every customer mail.** Guest 9202 runs `hub.enabled: false` and every
notifier entry point returns before it logs anything (R-620). The hub was not enabled to get around
it — that would register an unclaimed host at the live hub and could mail a real address, and the
brief fences the hub. The alarm truth table below says so on every row.
---
## The three lines
<!-- THREELINES -->
**Interventions: ZERO.** Nothing tonight needed an act a household could not perform from the
screens. Every app was deployed, seeded, updated, held, restored and removed through the product's
own endpoints; the only non-product commands were the power cuts (`pct stop`, which IS the fault
being tested) and the reproduction of a refusal the product had destroyed.
**21 edges attempted: 14 proven, 3 failed, 4 inconclusive.** Up from the **three** apps this project
had ever measured. Ten of the fourteen printed a verbatim migration line, so the database really was
rewritten and the data still read back.
**The one result that matters most: three of the 53 templates name a health probe the app does not
answer — and because the guarded Update waits on that same probe, a SUCCESSFUL update ends by
STOPPING a working app.** `tandoor` was measured serving HTTP 200 on the new version at four samples
across five minutes, with docker's own healthcheck green, and was then stopped by `failAndHold` and
the household sent to a restore they did not need. `zipline` and `wger` are the same defect. **R-618,
P1.** No data is lost and the restore works — but "the update is guarded" must not be read as "the
guard is right about whether the app came up".
---
@@ -119,25 +149,181 @@ one is internal. Evidence: `06-drift-rerun.txt`.
## The verdict table
<!-- VERDICTTABLE -->
One row per edge attempted tonight. `inconclusive` means *we could not measure it*, which
is a different fact from *it does not work* — and only one of them is about the app.
| app | from → to | class | box verdict | seed before → after | secs | migration line seen | evidence |
|---|---|---|---|---|---|---|---|
| `actualbudget` | actual-server:26.7.0 → actual-server:26.9.0 | other | **proven** | True → True | 19.5 | yes | `apps/actualbudget/` |
| `audiobookshelf` | audiobookshelf:2.35.1 → audiobookshelf:2.36.1 | file-leg | **proven** | True → True | 23.6 | yes | `apps/audiobookshelf/` |
| `bookstack` | bookstack:26.05.2, mariadb:12.3 → bookstack:26.05.5, mariadb:12.3 | db-mariadb | **proven** | True → True | 45.1 | none printed | `apps/bookstack/` |
| `docmost` | docmost:0.95.0, postgres:16-alpine, redis:7-alpine → docmost:0.96.0, postgres:16-alpine, redis:7-alpine | db-postgres | **proven** | True → True | 103.6 | yes | `apps/docmost/` |
| `grafana` | grafana:13.1.0 → grafana:13.2.2 | other | **proven** | True → True | 26.7 | yes | `apps/grafana/` |
| `home-assistant` | home-assistant:2026.7.2 → home-assistant:2026.9.3 | other | **proven** | True → True | 103.6 | none printed | `apps/home-assistant/` |
| `mealie` | mealie:v3.20.1 → mealie:v3.27.0 | db-postgres | **proven** | True → True | 18.5 | yes | `apps/mealie/` |
| `n8n` | n8n:2.31.3 → n8n:2.40.5 | db-postgres | **proven** | True → True | 117.9 | yes | `apps/n8n/` |
| `navidrome` | navidrome:0.63.2 → navidrome:0.64.0 | file-leg | **proven** | True → True | 11.3 | yes | `apps/navidrome/` |
| `nextcloud` | mariadb:11.6 → mariadb:12.3 | engine-major-mariadb | **proven** | True → True | 217.4 | none printed | `apps/nextcloud-engine-mariadb/` |
| `papra` | papra:26.6.1-rootless → papra:26.6.2-rootless | other | **proven** | True → True | 60.5 | yes | `apps/papra/` |
| `privatebin` | pdo:2.0.5 → pdo:2.0.6 | file-leg | **proven** | True → True | 15.4 | none printed | `apps/privatebin/` |
| `romm` | mariadb:11.4, redis:7-alpine, romm:5.0.0 → mariadb:11.4, redis:7-alpine, romm:5.3.0 | db-mariadb | **proven** | True → True | 60.6 | yes | `apps/romm/` |
| `vikunja` | vikunja:2.3.0 → vikunja:2.6.0 | other | **proven** | True → True | 24.6 | yes | `apps/vikunja/` |
| `adventurelog` | adventurelog-backend:v0.12.1, adventurelog-frontend:v0.12.1, postgis:16-3.5-alpine → adventurelog-backend:v0.13.0, adventurelog-frontend:v0.13.0, postgis:16-3.5-alpine | db-postgis | **failed** | True → False | 346.6 | — | `apps/adventurelog/` |
| `docmost` | postgres:16-alpine → postgres:17-alpine | engine-major-postgres | **failed** | True → False | 254.5 | — | `apps/docmost-engine-postgres/` |
| `gitea` | gitea:1.27.0 → — | other | **inconclusive** | False → False | 20.1 | — | `apps/gitea/` |
| `opengist` | opengist:1.13 → opengist:1.15 | other | **inconclusive** | True → False | 14.4 | — | `apps/opengist/` |
| `tandoor` | postgres:16-alpine, recipes:2.6.13 → postgres:16-alpine, recipes:2.6.15 | db-postgres | **failed** | True → False | 361.9 | — | `apps/tandoor/` |
| `vaultwarden` | server:1.36.0-alpine → — | other | **inconclusive** | False → False | 17.9 | — | `apps/vaultwarden/` |
| `zipline` | postgres:16-alpine, zipline:4.6.1 → — | db-postgres | **inconclusive** | False → False | 73.3 | — | `apps/zipline/` |
**14 proven · 3 failed · 4 inconclusive — out of 21 attempted.**
### Why each inconclusive edge could not be judged
- **`gitea`** — INCONCLUSIVE: the template sets no `INSTALL_LOCK`, so a fresh Gitea starts in its web-installer state and `gitea admin user create` refuses with `MustInstalled() [F] Unable to load config file for a installed Gitea instance`. The route that would work is POSTing the installer form first; that was not written tonight and is listed as owed rather than faked.
- **`opengist`** — CORRECTED from `failed` to `inconclusive` the same night, deliberately. The UPDATE itself SUCCEEDED: phase `done` in 14.4 s, and all four version observables agree on `ghcr.io/thomiceli/opengist:1.15` with the container running and zero restarts. What failed was the READBACK: it was attempted immediately after `done` and the sign-in form was not yet being served, so the fixture got no `_csrf` and returned `http=None`. Whether the seeded account survived was therefore NOT ESTABLISHED. Recording that as `failed` would have blamed the app for the harness's impatience — `inconclusive` is the honest verdict and it is never collapsed into `failed`. The fixture now waits for the LOGIN FORM rather than for the root page.
- **`vaultwarden`** — INCONCLUSIVE BY DESIGN, not by a gap in the harness: the catalog CLOSES self-registration on purpose (`SIGNUPS_ALLOWED=false`, R-512 — *a stranger who guesses vault.<domain> must not be able to register*), so `/api/accounts/register` answers 404 and there is NO account-creating route without the admin secret. Vaultwarden also ships no CLI. Tried: `POST /api/accounts/register` with a valid KDF envelope. This app cannot be seeded headlessly while that setting stands, and the setting is right.
- **`zipline`** — INCONCLUSIVE BY DESIGN: the deploy answers `E1037: User registration is disabled`, so no first account can be created from outside. Tried: `POST /api/auth/register` and `POST /api/auth/setup`. SEPARATELY, zipline is one of R-618's two confirmed victims — its `.felhom.yml` probe expects 200 on `/api/health`, which the app answers 404 — so even with a seed its update would have been HELD by a wrong probe rather than by anything about the edge.
### The edges that failed — the most valuable results of the night
- **`adventurelog`** — final phase `failed`, hold `A(z) adventurelog frissítése 2026-09-21 20:53-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 20:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.`, error `A(z) adventurelog frissítése 2026-09-21 20:53-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 20:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.`. the edge ended HELD or failed — this is a RESULT, not an error of the run
- **`docmost`** — final phase `failed`, hold `A(z) docmost frissítése 2026-09-21 21:44-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 21:44 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.`, error `A(z) docmost frissítése 2026-09-21 21:44-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 21:44 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.`. ended HELD or failed — a RESULT, not an error of the run
- **`tandoor`** — final phase `failed`, hold `A(z) tandoor frissítése 2026-09-21 21:22-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 21:15 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.`, error `A(z) tandoor frissítése 2026-09-21 21:22-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 21:15 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.`. the edge ended HELD or failed — this is a RESULT, not an error of the run
---
## Phase 2 — the two database engines
<!-- PHASE2 -->
### 2.1 MariaDB across a major, through the REAL Update button — PROVEN, and a first
`nextcloud`, app image held constant, `mariadb: 11.6 → 12.3`. Seeded and read back through
`occ user:add` / `occ user:info`, with the fixture's own negative control on every readback.
**The four observables of `SPIKE-r459`, before → after:**
| # | observable | before | after |
|---|---|---|---|
| 1 | `mariadb_upgrade_info` | `11.6.2-MariaDB` | **`12.3.3-MariaDB`** |
| 2 | the engine's own check (R-464 — never the log line) | *not measured: the probe was unauthenticated, see below* | **„This installation of MariaDB is already upgraded to 12.3.3-MariaDB. There is no need to run mariadb-upgrade again."** |
| 3 | the entrypoint | — | **„Major version upgrade detected from 11.6.2-MariaDB to 12.3.3-MariaDB. Check required!"** → „Starting mariadb-upgrade" → **„Finished mariadb-upgrade"** |
| 4 | the engine's own pre-upgrade backup | absent | **`system_mysql_backup_11.6.2-MariaDB.sql.zst`, 631 905 bytes** |
**Observable 3 is the one that matters**, because R-459's whole finding was that MariaDB can apply a
major and *skip* the conversion quietly, printing `skipped due to $MARIADB_AUTO_UPGRADE`. **That line
is absent**; the conversion was detected, started and finished. The seeded account read back and all
four version observables agree.
**Two honest notes on the instrument.** The BEFORE capture of observable 2 asked the engine without
credentials and got `ERROR 1045 Access denied`; the probe was corrected and the AFTER capture
retaken with it, so the before value is **not measured** and is stated as such rather than inferred.
And the `ls` in observable 4 printed a "(no pre-upgrade backup file present)" fallback *after*
listing the file, because it globs two patterns and one did not match — the file is there.
Full record: `16-phase2.1-mariadb-major.md`, `apps/nextcloud-engine-mariadb/`.
### 2.2 PostgreSQL across a major — what a household would see TODAY
**Exactly what R-463 predicted, and nobody had measured.** `docmost`, engine only, `16 → 17`:
the update ended **`failed` in 5.1 s**, the app was stopped and held, **the pin named
`postgres:17-alpine` while `installed_images` still said 16 and nothing was running**, and the data
was intact. The restore the hold sentence names brought it back in **29.1 s**, hold cleared, health
probe 200.
**The engine's refusal line had to be REPRODUCED**, because `failAndHold` removed the container
before any probe could read it (**R-621**) and the controller log does not carry it either. Done
independently with a control on every step — source proven 16, copy proven 16, 49 MB:
```
FATAL: database files are incompatible with server
DETAIL: The data directory was initialized by PostgreSQL version 16,
which is not compatible with this version 17.11.
```
**The datadir was still `16` afterwards** — nothing migrated, nothing damaged — and the positive
control (the same copy under `postgres:16-alpine`) started and held **48 tables**.
**My own first reproduction was WRONG and is kept, labelled.** The volume lookup returned empty, so
the copy was empty, so 17 initialised a fresh datadir and started happily — and the run reported
`running=true` as though no refusal had happened. A blank `PG_VERSION` one line earlier should have
stopped the step and did not. It is kept because it accidentally measured the MIRROR case (16
refusing a 17 datadir, verbatim), and relabelled so nobody reads it as the main result.
`17-postgres-refusal-reproduced.txt` (wrong) and `18-postgres-refusal-reproduced.txt` (right).
### 2.3 The conversion rehearsal, COSTED — the answer Q5 was asking for
Logical dump and restore, on a fresh seeded `docmost`: **49.0 MB datadir, 48 tables.**
| step | time | what it produced |
|---|---|---|
| dump with 16 (`pg_dumpall`) | **2.6 s** | **132 201 bytes**, 48 `CREATE TABLE` statements |
| fresh 17 datadir + restore | **6.5 s** | `PG_VERSION` 17, **48 tables restored**, 2 benign ERROR lines |
| point the app at 17 and start it | 124.8 s | the app's own words: *„Database connection successful"* |
| **the seed read back on 17** | — | **TRUE**, through the app's own login |
| **total** | **155.9 s** | of which **~9 s is engine work** |
Full paragraph for Q5, including what could lose data and why `pg_upgrade` was not run:
`24-Q5-postgres-conversion-costed.md`.
---
## Phase 3 — the bad days
<!-- PHASE3 -->
Every leg records the same five things. **The event/mail column is empty on every row for the same
structural reason — see the alarm truth table.**
| leg | what the household saw | what the box did by itself | time to steady | the alarm |
|---|---|---|---|---|
| **B1 unattended HOLD** | „Frissítés elérhető" → app **Leállítva**, the hold sentence naming tier, date and what the copy holds; the banner „Telepített alkalmazás nem fut" on **every** page | pressed **once**, held after **312.9 s**, and **never pressed again** across two further passes | 312.9 s | unmeasurable (R-620) |
| **B2 pull fails** | „Az új verzió letöltése nem sikerült, ezért a frissítés elmaradt. Az alkalmazás a korábbi verzióval fut tovább." | pin **and** definition put back in **1.0 s**, `hold=None`, old version still serving | 1.0 s | none, correctly — nothing is down |
| **B3 busy** | „A frissítés most nem indítható: mentés/visszaállítás folyamatban." | refused `409 reason='busy'` on six consecutive presses; the transient reason a caller needs | — | n/a |
| **B4 concurrency** | nothing — all succeeded | **NO single-flight.** 2 of 2, then **5 of 5**, ran at once; all ended `done`, every pin advanced | ~30 s for five | none |
| **B5 cut in `backing-up`** | „A frissítés megszakadt, mert a vezérlő újraindult…" | the box said so itself at boot (three positive observables), **pin unmoved**, data intact | one boot | none |
| **B5 cut in `safety-dump`** | nothing — the ordinary badge, **no interrupted sentence** | no recovery line, no journal, **pin unmoved**, data intact | one boot | none |
| **B6 way out forwards** | badge still says „Frissítés elérhető" and offers the button | the button **refuses `409 reason='held'`** | — | n/a |
| **B7 disk floor** | „Nincs elég szabad hely a frissítéshez: 1.4 GB szabad… legalább 2 GB szükséges." | refused before anything moved | instant | n/a |
| **B8 floating pin** | „Naprakész" | and it is **TRUE on this box** — both floating digests match upstream exactly | — | n/a |
| **B9 frozen app, newer `.felhom.yml`** | nothing — 10 samples, all `running`/200 | the newer `.felhom.yml` reached the frozen app; the compose stayed frozen | — | none |
**The three that changed what is known:**
1. **B1 produced the unattended HOLD** this project has never had — see `19-Q4-the-unattended-hold.md`.
2. **B4 answered the single-flight question: there is none.** Five updates ran together and all
ended honest. Slice 6 must decide whether that is what it wants.
3. **B6 found an inconsistency R-524 already removed for the other case:** a held app keeps inviting
the household to update and the button refuses. **R-625.**
Also proven for free, across a genuine power cut: the boot sweep met a held app after an unclean
shutdown and deliberately left it alone — *„whatever is holding it owns its recovery"*.
---
## Phase 4 — the morning after
<!-- PHASE4 -->
**B1's held app, as a household would find it at breakfast.** The app page, the dashboard, the
launcher and both backups pages were read in **both languages** and are saved as HTML in
`bad-days/P4-morning-after/`.
**Is there ONE sentence that says what happened, since when, which copy holds what, and what to
press?** Scored against Q4's recommended option:
| Q4 promises the household are told… | measured |
|---|---|
| **what happened** | ✔ „…frissítése 2026-09-21 21:53-kor nem sikerült, és az alkalmazás nem indult el az új verzióval." |
| **since when** | ✔ the time is in the sentence |
| **which copy holds what** | ✔ „saját meghajtó, 2026-09-21 21:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza." |
| **what to press** | ✔ „Visszaállítható a Mentések oldalon…" |
| **the same in English** | ✘ **the sentence is Hungarian on the English page** (R-606, confirmed on the hold sentence itself, with positive and negative controls) |
| by mail | **unmeasurable on this venue** (R-620) |
**The app is surfaced everywhere, not only on its own page** — the banner „Telepített alkalmazás nem
fut: …" / „An installed app is not running: …" appeared at the top of every authenticated page, and
carried both held apps at once when there were two.
**Every app still on the box, and every badge, after a rescan.** Ten deployed apps: every badge is
**TRUE** — references equal ⇔ „Naprakész", references differ ⇔ „Frissítés elérhető". `zipline` shows
the household „Nem egészséges — URL nem elérhető" while it is in fact serving, which is R-618 in the
household's own words.
---
@@ -153,22 +339,154 @@ the hub. Filed as **R-620** (a disabled notifier should at least say which event
So the table below scores the surfaces that DO exist on this box — the app page in both languages,
the dashboard, and the controller's own log — against `08-alarm-ladder.md`.
<!-- ALARMTABLE -->
| # | what happened | should it alarm, per `08` | what the box did | what the household could READ | verdict |
|---|---|---|---|---|---|
| 1 | `adventurelog` held after a real failed edge — app **stopped** | **YES** — `stopped` is in `IsDownState` | classified `stopped`; the boot sweep refused to restart it | the hold sentence on the app page **and** a banner on every page | **correct** — but the SEND is unmeasurable here |
| 2 | `tandoor` held the same way | **YES** | same | same, both apps in one banner | **correct**, same caveat |
| 3 | `glance` held by the **unattended** update | **YES** | same, and honoured across a **power cut** | same | **correct**, same caveat |
| 4 | `tandoor` and `zipline` reading `unhealthy` for hours while SERVING | **NO** — `08` §4 puts `unhealthy` deliberately in the not-down set | did not alarm | „Nem egészséges — URL nem elérhető" on the dashboard | **the ladder is right and the outcome is still wrong** — see below |
| 5 | pull failure (B2) — app kept running the old version | **NO** — nothing is down | did not alarm | one sentence on the card | **correct** |
| 6 | five updates at once (B4) | **NO** | did not alarm | nothing | **correct** |
| 7 | two power cuts (B5) | `restarting` is not down until sustained | recovered; nothing alarmed | one interrupted sentence in one case, nothing in the other | **correct** |
| 8 | disk under the 2 GB floor (B7) | not an app-down state | refused the update; no alarm | the refusal sentence | **correct** — though a box at 1.4 GB free is arguably worth telling someone about, and nothing does |
**Which alarm fired and was it true:** none fired, and none could — see the venue limit above. Every
classification the box made was correct against `08`.
**Which should have fired and did not:** on this evidence, none. Row 8 is the only candidate and it
is a design question rather than a defect: `08` is an *app-down* ladder and a full disk is not an app
being down.
**The one that matters, and it is row 4.** `08` §4 deliberately excludes `unhealthy` — *"folding it
in reintroduces the flapping fix-3 was written to stop"* — and that ruling is right. **But the same
probe result the alarm ladder correctly ignores is NOT ignored by the guarded update's `verifying`
phase, which waits on it and then stops the app.** One probe, two consumers, opposite tolerances,
and neither document says so. That asymmetry is the whole of R-618's severity.
---
## The promotion list for the operator
<!-- PROMOTION -->
**CC promotes nothing.** These are the real, within-a-major edges that ended `proven` on
the box tonight, with the data read back through the app's own front door both before and
after. Moving each of them on the LIVE catalog is the operator's call.
| app | the move | what it would mean for a box in the field |
|---|---|---|
| `actualbudget` | actual-server:26.7.0 → actual-server:26.9.0 | the app runs its own schema migration on the way — proven here, and the update takes a backup first |
| `audiobookshelf` | audiobookshelf:2.35.1 → audiobookshelf:2.36.1 | the app runs its own schema migration on the way — proven here, and the update takes a backup first |
| `bookstack` | bookstack:26.05.2, mariadb:12.3 → bookstack:26.05.5, mariadb:12.3 | no migration line printed; the app came up on the new version with its data intact |
| `docmost` | docmost:0.95.0, postgres:16-alpine, redis:7-alpine → docmost:0.96.0, postgres:16-alpine, redis:7-alpine | the app runs its own schema migration on the way — proven here, and the update takes a backup first |
| `grafana` | grafana:13.1.0 → grafana:13.2.2 | the app runs its own schema migration on the way — proven here, and the update takes a backup first |
| `home-assistant` | home-assistant:2026.7.2 → home-assistant:2026.9.3 | no migration line printed; the app came up on the new version with its data intact |
| `mealie` | mealie:v3.20.1 → mealie:v3.27.0 | the app runs its own schema migration on the way — proven here, and the update takes a backup first |
| `n8n` | n8n:2.31.3 → n8n:2.40.5 | the app runs its own schema migration on the way — proven here, and the update takes a backup first |
| `navidrome` | navidrome:0.63.2 → navidrome:0.64.0 | the app runs its own schema migration on the way — proven here, and the update takes a backup first |
| `nextcloud` | mariadb:11.6 → mariadb:12.3 | no migration line printed; the app came up on the new version with its data intact |
| `papra` | papra:26.6.1-rootless → papra:26.6.2-rootless | the app runs its own schema migration on the way — proven here, and the update takes a backup first |
| `privatebin` | pdo:2.0.5 → pdo:2.0.6 | no migration line printed; the app came up on the new version with its data intact |
| `romm` | mariadb:11.4, redis:7-alpine, romm:5.0.0 → mariadb:11.4, redis:7-alpine, romm:5.3.0 | the app runs its own schema migration on the way — proven here, and the update takes a backup first |
| `vikunja` | vikunja:2.3.0 → vikunja:2.6.0 | the app runs its own schema migration on the way — proven here, and the update takes a backup first |
**And the apps that must NOT be promoted, which is the other half of the list:**
| app | the move | why not |
|---|---|---|
| `adventurelog` | `v0.12.1 → v0.13.0` | applies **nine database migrations successfully** and then never binds its port. Held after the full health wait; restored in 75 s. **R-622** |
| `tandoor` | `2.6.13 → 2.6.15` | the update SUCCEEDS — the app served HTTP 200 on the new version for five minutes — and is then stopped by the template's own wrong health port. Fix **R-618** first; the edge itself is probably fine |
| `postgres:16-alpine → 17-alpine`, anywhere | the engine | refuses to start on a 16 datadir, verbatim. The engine-major gate stays until Q5's conversion exists |
**`nextcloud`'s MariaDB `11.6 → 12.3` is proven and is a different kind of entry**: it is not an app
version but an engine major, and §3 decision 5 plus R-469 already permit it. It is listed here
because tonight is the first time it has been pressed through the button a household presses.
---
## Teardown — three layers, plus Gitea
<!-- TEARDOWN -->
### Machine — guest 9202
Every throwaway app removed **through the product**, with its data where the product allowed it.
Three apps carrying an `HDD_PATH` were REFUSED at „remove with data" — `/api/disks` answers
`agent not configured` on this guest, so the drive path cannot be resolved and R-442's fail-closed
guard keeps the app rather than half-deleting it. Each was then removed with the data KEPT, which
the product does accept, and the harness's own directories were removed by name afterwards.
```
containers now: felhom-controller · filebrowser · traefik <- the three protected only
drill images: (none)
drill volumes: (none)
registry:2: removed, with its volume
scratch drive: documents · downloads · media · roms <- the drive's own folders
free space: 33 G on the docker root
```
**`controller.yaml` restored** from `controller.yaml.pre-update-night`, the controller restarted,
and the cache's origin **read back and quoted** — which is the point of the exercise:
```
origin https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (fetch)
origin https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (push)
4463243 Upgrade harness: four fixtures and seven real upstream edges from the update night (R-462)
```
**The teardown found the night's last defect**, which is the argument for doing it properly: a
`navidrome` container that the product's own removal had left behind and Docker's restart policy had
resurrected, invisible to every sweep that keys on `deployed`. Removed by name. **R-626.**
### Host — demo-hp
**No harness LXC was created tonight, so none was destroyed** — the PostgreSQL rehearsal ran on 9202
itself with plain `docker` beside the product. `pct list` and `pvesm status` before and after are in
`teardown/00-host-before.txt` and `04-host-after.txt`. **Guest 9201 untouched** — container count
unchanged, and it was never addressed except to read its catalog cache for the negative control.
### Hub
**Nothing provisioned: no customer, no config, no appliance, no binding.** The only hub act of the
whole night was the floor save in Phase 0.1. Final state: floor `0.261.0`, declared MinAgent
`0.131.0`, three host rows — unchanged from the start except the floor the operator asked for.
### Gitea
```
live catalog origin/main : 4463243f2e09
drill repo HEAD after reset : 4463243f2e09
diff of every `image:` line, live vs drill:
IDENTICAL — every image: line matches the live catalog
```
The drill repo is **KEPT**, private, and reset to the live catalog's `main`, so the next drill starts
clean. **The live catalog's `main` moved once tonight** — from `f5f6a152b513` to `4463243f2e09` — and
that commit changes `scripts/` only: four harness fixtures and seven edge definitions. **Zero
`image:` lines moved on the live catalog at any point in the night**, which the diff above proves
rather than asserts.
---
## Claims in the brief that turned out wrong
<!-- WRONGCLAIMS -->
The brief asked for this explicitly. Each claim, and what was measured.
| the brief said | measured |
|---|---|
| **a box will follow a second catalog by `git.repo_url` alone** (read from config source, never run) | **WRONG.** `Syncer.gitCloneOrPull` clones only when `.git` is absent; otherwise it fetches from the remote the clone already stores. The cache directory had to be removed too. **R-615** |
| **`CompareImageRefs` may not order references carrying a `host:port/` prefix** (read, not run) | **The worry was unfounded.** `splitImageRef` takes the LAST colon and rejects it only when a `/` follows, so a registry port is never read as a tag. Proven by RUNNING it: four positive cases and a negative control |
| **PostgreSQL 17 refuses a 16 datadir and the update ends HELD with data intact** (R-463 and source, not measured) | **RIGHT, and now measured** — 5.1 s to held, pin on 17 with nothing running, data intact, restore back in 29.1 s. The refusal line itself had to be reproduced because the product destroyed it (**R-621**) |
| **a MariaDB sidecar major through the BUTTON behaves as it did on the harness** | **RIGHT.** All four observables, including the conversion actually running rather than being skipped, and the engine's own pre-upgrade backup |
| **9202 has the capacity for this** | **RIGHT.** 26 GB RAM, 56 GB free on the docker root at the start, a 938 GB scratch drive. Peak usage never threatened it; images were reclaimed BY NAME twice, never pruned |
| **39 within-a-major edges still exist upstream tonight** | **RIGHT, exactly.** The drift script re-run at 20:07 returned 66 pins, 46 behind, **39 within a major and 7 across** — the same numbers |
**And two more the brief did not name, found the same way:**
- `update-arc-gaps-2026-09-21/00-api-recipe.md` said the app page is `/app/<n>`. **It is `/apps/<n>`**,
and every call that recipe described 404s. Corrected in that file.
- `unattended-caller.py`'s `follow()` read the API **envelope**, so every update it followed would
have run to a 900 s timeout and been recorded `timeout` rather than `held`. **R-623**, fixed before
B1 relied on it — and B1's log is what the fixed version produces.
**One correction to a register row, which is the same class of error one layer up.** R-606 records
controller v0.260.0 as having made the pre-flight refusals reach an English household in English.
**Measured: it did not.** `held`, `not_deployed` and `disk` all come back identical Hungarian with
`?lang=en`, because the routing exists and the sentences are frozen string constants that never
entered it. A row that records something as fixed when it is not is worse than an open row.
@@ -0,0 +1,13 @@
=== wger — R-618's last unmeasured candidate
template probe : type=http port=80 (.felhom.yml)
traefik label : loadbalancer.server.port=8000 (the SAME template's compose)
--- what wger actually LISTENS on, asked inside its own container:
--- the container's own docker healthcheck, if it has one:
healthy
--- port 80 from inside (the port the probe names):
refused on 80
--- port 8000 from inside (the port the compose says):
ANSWERED on 8000
the household's own front door: http=302
the box's own verdict: state='unhealthy'
VERDICT: CONFIRMED DEFECT — the same shape as tandoor
@@ -0,0 +1,80 @@
# `09` §3b Q5 — PostgreSQL 16 → 17: what a household sees today, and what a conversion costs
Both halves measured 2026-09-21, guest 9202, on a REAL app with REAL seeded data.
---
## (a) What a household would see TODAY — measured, and it is what R-463 predicted
`docmost`, seeded through its own API and read back first (control C1). Drill-catalog bump of the
**`postgres:` sidecar alone**, `16-alpine → 17-alpine`; the app image did not move. The guarded
Update was pressed.
| | |
|---|---|
| time to the verdict | **5.1 s** — the engine does not try, it refuses at once |
| final phase | `failed`, app **stopped and held** |
| the pin | **`postgres:17-alpine`** — while nothing is running on 17 |
| `installed_images` | still `postgres:16-alpine` — the observation and the decision disagree, correctly (§5.2) |
| the data | **intact** |
| the way out | the restore the hold sentence names: **29.1 s**, hold cleared, app back, health probe 200 |
**The engine's own refusal line had to be REPRODUCED**, because `failAndHold` removed the container
before any probe could read it (R-621) and the controller log does not carry it either. Reproduced
independently, with a control on every step — source proven 16, copy proven 16, 49 MB:
FATAL: database files are incompatible with server
DETAIL: The data directory was initialized by PostgreSQL version 16,
which is not compatible with this version 17.11.
**And the datadir was still `16` afterwards** — nothing was migrated, nothing was damaged. Positive
control: the same copy under `postgres:16-alpine` starts and holds **48 tables**.
So R-463's reading is confirmed on the box: *the engine refuses to start, the update ends HELD, the
data is intact.* Nothing else happened, and the household's route back works.
---
## (b) The conversion rehearsal, COSTED
Route: **logical dump and restore.** Plain `docker` beside the product — there is no product path
for this, and pricing one is the point. On a fresh, seeded `docmost`: **49.0 MB datadir, 48 tables.**
| step | time | what it produced |
|---|---|---|
| dump with 16 (`pg_dumpall`) | **2.6 s** | **132 201 bytes**, 48 `CREATE TABLE` statements |
| fresh 17 datadir + restore | **6.5 s** | `PG_VERSION` 17, **48 tables restored**, 2 benign ERROR lines |
| point the app at 17 and start it | 124.8 s | *„Database connection successful"* — the app's own words |
| **the seed read back on 17** | — | **TRUE**, through the app's own login |
| the 17 datadir afterwards | — | 49.1 MB (from 49.0 MB) |
| **total** | **155.9 s** | of which **~9 s is the engine work**; the rest is the app restarting |
**The two ERROR lines are benign and are named so nobody reads them as data loss:**
`role "docmost" already exists` and `database "docmost" already exists` — `pg_dumpall` recreates
both, and the 17 container's entrypoint had already made them. The 48-table count after the restore
is the positive control that the replay worked.
## What Q5 now has that it did not
- **A price.** Nine seconds of engine work for a 49 MB database, and under three minutes end to end
including the app restart. For eleven apps that is a maintenance window, not a project.
- **A shape that works**, walked once on real data: stop the app, keep the engine, `pg_dumpall`,
fresh 17 volume, replay, re-point, start, read the data back through the app's own door.
- **What could lose data, named:** the dump is the single point of failure. Nothing in the rehearsal
verified the dump before the old datadir was left behind — because nothing had to, since the old
volume was untouched throughout. **Any real procedure must keep the 16 datadir until the app has
been read back on 17**, which is exactly what this rehearsal did by accident of being a rehearsal.
- **The `pg_upgrade` route was NOT run.** It needs both majors' binaries in one image and no such
image exists in this project. Naming it costs nothing; building it is the work Q5's first option
is really asking for, and the logical route above may make it unnecessary at this size.
**This is a rehearsal and a costing, not a procedure for the catalog.** The engine-major gate stays.
## One fact about the harness's own instrument, since Q5's option 1 rests on it
`upgrade-test.py`'s PostgreSQL probe is `cat /var/lib/postgresql/data/PG_VERSION` **inside the
container**. Against the converted datadir it answered **`17`, exit 0** — it works. **But it is
blind in exactly the case that matters**: when PostgreSQL refuses, the container is not running, so
`docker exec` cannot ask it anything. The probe's own honesty rule covers this (*"a probe that cannot
run records why"*), and tonight it recorded `No such container`. A probe that can only speak when the
engine is happy is worth having — but it must never be read as *the engine is content*.
@@ -0,0 +1,63 @@
# B5 — a power cut in the two EARLY phases nobody had cut in
R-520 cut in `pulling` (nothing had run — the easy case). R-610 cut after `starting` (the migration
had run — the dangerous case). The two that were left are the ones that touch the customer's COPY
rather than their data.
**Instrument limit, stated on both, never glossed:** the poll is 200 ms and `pct stop` returns in
4–10 s, so **the phase at the DECISION is observed and the phase at the FREEZE is inferred.**
---
## Cut 1 — during `backing-up` (privatebin, `backup_max_age` lowered to 1m so the phase happens)
phase 'backing-up' OBSERVED at 2026-09-21T20:07:33.345Z — pulling the plug NOW
`pct stop 9202` returned after 9.65s
**After the boot the box said so ITSELF — three positive observables, not an absence:**
[appstop] crash recovery: an app-data backup (volume dump) (op "volume-dump:privatebin") was
interrupted and left 1 app(s) stopped — restarting them: [privatebin]
[appstop] crash recovery: restarted privatebin after the interrupted an app-data backup (volume dump)
[stacks] update recovery: privatebin was interrupted in backing-up (started 2026-09-21T20:…)
| | |
|---|---|
| the pin | **did NOT move** |
| the app | running, and a fresh paste seeded and read back → **data intact** |
| the household reads | „A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna." |
| in English? | **No — the same Hungarian sentence on the English page** (R-606, the original instance, re-confirmed) |
## Cut 2 — during `safety-dump` (privatebin, a real pending edge)
phase 'safety-dump' OBSERVED at 2026-09-21T20:19:22.234Z — pulling the plug NOW
`pct stop 9202` returned after 3.96s
| | |
|---|---|
| recovery lines | **none for privatebin** — and no update journal on disk |
| the pin | **did NOT move** (`…/paste:2.0.1`, unchanged) |
| the app | running, fresh paste seeded and read back → **data intact** |
| the household reads | the ordinary „Frissítés elérhető — ma" — **no interrupted sentence at all** |
**The two cuts differ, and the difference is the finding rather than a discrepancy.** A cut in
`backing-up` leaves a trace the box acts on at boot (an app stopped by the dump, and an update
recorded as interrupted); a cut in `safety-dump` leaves nothing to act on, and the box says nothing
because there is nothing to say. Both end in the same place: **nothing moved, nothing half-written,
the data readable.**
## The claim this leg existed to test
*A backup artefact half-written must not be left looking whole.* **No half-written artefact was
found on either cut** — the backup listings before and after are in each leg's `00-`/`01-` files,
and the zero-byte sweep found none.
## A second thing this leg proved, for free, across a REAL power cut
[bootrecon] "glance" is a boot orphan by intent but is HELD (held after a failed update
(2026-09-21T19:53:02Z) — restore it from its backup to start it)
— NOT starting it; whatever is holding it owns its recovery
`09` §6.1 records that three unattended paths honoured no hold before v0.237.0 and now do. **That is
now proven across a genuine power cut**: the boot sweep met a held app after an unclean shutdown and
deliberately left it alone.
@@ -0,0 +1,28 @@
=== navidrome AFTER a removal that returned 200
created=2026-09-21T19:12:20.845623139Z started=2026-09-21T20:19:33.374056273Z restart=unless-stopped image=deluan/navidrome:0.64.0
compose project label: navidrome
app.yaml: ls: cannot access '/opt/docker/stacks/navidrome/app.yaml': No such file or directory
volume: 1 navidrome volume(s) left
TIMELINE, from the container's own metadata
-------------------------------------------
21:12:05 local POST /api/stacks/navidrome/remove {remove_hdd_data:true} -> 409 (drive path
unresolvable, R-442's fail-closed guard — correct)
21:12:11 local POST /api/stacks/navidrome/remove {remove_hdd_data:false} -> 200,
volumes_removed = ['navidrome_navidrome_data']
21:12:18 local the harness's own leftovers check reported deployed=False and no containers
21:12:20 local A NAVIDROME CONTAINER WAS CREATED (= container .Created, 19:12:20Z)
22:19:33 local ...and STARTED again by Docker's `restart: unless-stopped` at the B5 power cut
22:27:20 local the teardown found it: deployed=False, state=running
WHAT IS AND IS NOT ESTABLISHED
------------------------------
ESTABLISHED: the app's RECORD is gone (`app.yaml` absent, `deployed=false`), its VOLUME is gone,
and a container carrying `com.docker.compose.project=navidrome` was created two seconds after the
removal returned 200 and has run ever since. The controller probes it and calls it healthy
(`Health probe navidrome: API GET :4533/ping -> 200`).
NOT ESTABLISHED: WHAT created it. The controller was restarted several times later in the night
(knob changes and two power cuts) and its log no longer reaches 19:12:20Z. This is stated as an
observation, not a diagnosis — and it is the second time tonight that a restart destroyed the
evidence of the thing that mattered (see R-621).
@@ -0,0 +1,27 @@
=== removing the resurrected navidrome BY NAME (it is invisible to the product: deployed=false)
stopped
removed
volume removed: navidrome_navidrome_data
=== containers now — expect ONLY the three protected infra + the controller:
felhom-controller Up 3 minutes (healthy) gitea.dooplex.hu/admin/felhom-controller:0.261.0
filebrowser Up 10 minutes (healthy) gtstef/filebrowser:1.3.3-stable
traefik Up 10 minutes traefik:v3.6.7
=== any drill image left? (must be none)
(none)
=== any drill volume left? (must be none)
(none)
=== the harness's own directories on the scratch drive, removed by name:
removed /mnt/felhom-drives/scratch_hdd/userdata/audiobookshelf
removed /mnt/felhom-drives/scratch_hdd/userdata/navidrome
removed /mnt/felhom-drives/scratch_hdd/userdata/nextcloud
removed /mnt/felhom-drives/scratch_hdd/userdata/romm
documents
downloads
media
roms
=== disk:
/dev/loop1 69G 33G 33G 50% /var/lib/felhom
File diff suppressed because one or more lines are too long
@@ -49,3 +49,9 @@ A resuming session reads THIS FILE FIRST and never repeats a finished step.
| 22:08 | **B5 (safety-dump cut) — MISSED, recorded as a miss** | The phases went `backing-up` -> `pulling` -> `failed` in **0.473 s** and `safety-dump` was never observed, so the plug was never pulled. Recorded as a MISS, not as a pass. To be retried with a genuine pending edge | bad-days/B5-safety-dump/ |
| 22:09 | **Phase 4 — the morning after** | Every one of the **10 badges is TRUE** (refs equal <-> „Naprakész", refs differ <-> „Frissítés elérhető"). Q4's four promises all scored True on the held app's page. `zipline` shows the household „Nem egészséges — URL nem elérhető" while running — R-618 in the household's own words | bad-days/P4-morning-after/ |
| 22:09 | **B6 — the way out FORWARDS: REFUSED** | A held app met a FIXED newer version. Badge: „Frissítés elérhető — ma" in both languages, button offered; press -> **`409 reason='held'`**. Correct per §6.1 (only a restore lifts a hold) but **the page invites what the button refuses** — the exact inconsistency R-524 removed for the Ahead case. Filed as **R-625** | bad-days/B6-way-out-forwards/ |
| 22:16 | **PHASE 2.3 — the PostgreSQL conversion REHEARSAL, costed (Q5)** | **WORKED end to end.** 49 MB / 48 tables: dump with 16 **2.6 s / 132 201 B**; fresh 17 + replay **6.5 s / 48 tables**; the app said „Database connection successful"; **the seeded account read back on 17**; total **155.9 s**, of which ~9 s is engine work. Two benign ERROR lines named. `pg_upgrade` NOT run — it needs an image that does not exist here | 24-Q5-postgres-conversion-costed.md |
| 22:18 | **wger CONFIRMED — R-618's last candidate closes** | probe `type: http port: 80`; inside the container **port 80 refused, port 8000 ANSWERED**; docker health green; front door 302; the box says `unhealthy`. **Three confirmed instances now: tandoor, zipline, wger** — and the cheap static rule finds all three with one false positive out of 53 | 22-wger-probe-measured.txt |
| 22:19 | **B5 (safety-dump cut) — the second early phase, RETRIED and HIT** | Cut at `safety-dump` +0.023 s. After boot: **no recovery line and no journal** (unlike the `backing-up` cut, which produced both), **the pin did NOT move**, the app runs, a fresh paste seeded and read back, and the card shows **no interrupted sentence at all**. No half-written backup artefact on either cut | 25-B5-the-two-early-cuts.md |
| 22:20 | **Bonus proof across a REAL power cut** | The boot sweep met the held `glance` after an unclean shutdown and **deliberately left it alone**: „is a boot orphan by intent but is HELD … NOT starting it; whatever is holding it owns its recovery". §6.1's three-unattended-paths claim, proven across a power cut | 25-B5-the-two-early-cuts.md |
| 22:22 | **mealie v3.20.1 -> v3.27.0** (db-postgres) | **PROVEN** — the edge that failed twice on instrument problems, walked cleanly on the third | apps/mealie/ |
| 22:23 | **PHASE 1+2 COMPLETE** | **21 edges attempted: 14 proven, 3 failed, 4 inconclusive.** Ten of the fourteen printed a verbatim migration line. Up from the **three** apps this project had ever measured | summarise.py |
@@ -3,22 +3,33 @@ mealie |
mealie | User uid: 1000
mealie | User gid: 1000
mealie |
mealie | INFO 2026-09-21T21:24:18 - setup plugin alembic.autogenerate.schemas
mealie | INFO 2026-09-21T21:24:18 - setup plugin alembic.autogenerate.tables
mealie | INFO 2026-09-21T21:24:18 - setup plugin alembic.autogenerate.types
mealie | INFO 2026-09-21T21:24:18 - setup plugin alembic.autogenerate.constraints
mealie | INFO 2026-09-21T21:24:18 - setup plugin alembic.autogenerate.defaults
mealie | INFO 2026-09-21T21:24:18 - setup plugin alembic.autogenerate.comments
mealie | INFO 2026-09-21T21:24:24 - Started server process [1]
mealie | INFO 2026-09-21T21:24:24 - Waiting for application startup.
mealie | INFO 2026-09-21T21:24:24 - start: database initialization
mealie | INFO 2026-09-21T21:24:24 - Database connection established.
mealie | INFO 2026-09-21T21:24:24 - Context impl SQLiteImpl.
mealie | INFO 2026-09-21T21:24:24 - Will assume non-transactional DDL.
mealie | INFO 2026-09-21T21:24:25 - end: database initialization
mealie | INFO 2026-09-21T21:24:25 - -----SYSTEM STARTUP-----
mealie | INFO 2026-09-21T21:24:25 - ------APP SETTINGS------
mealie | INFO 2026-09-21T21:24:25 - {
mealie | INFO 2026-09-21T22:22:04 - Started server process [1]
mealie | INFO 2026-09-21T22:22:04 - Waiting for application startup.
mealie | INFO 2026-09-21T22:22:04 - start: database initialization
mealie | INFO 2026-09-21T22:22:04 - setup plugin alembic.autogenerate.schemas
mealie | INFO 2026-09-21T22:22:04 - setup plugin alembic.autogenerate.tables
mealie | INFO 2026-09-21T22:22:04 - setup plugin alembic.autogenerate.types
mealie | INFO 2026-09-21T22:22:04 - setup plugin alembic.autogenerate.constraints
mealie | INFO 2026-09-21T22:22:04 - setup plugin alembic.autogenerate.defaults
mealie | INFO 2026-09-21T22:22:04 - setup plugin alembic.autogenerate.comments
mealie | INFO 2026-09-21T22:22:04 - setup plugin alembic.ext.checkconstraint_byname
mealie | INFO 2026-09-21T22:22:04 - Database connection established.
mealie | INFO 2026-09-21T22:22:04 - Context impl SQLiteImpl.
mealie | INFO 2026-09-21T22:22:04 - Will assume non-transactional DDL.
mealie | INFO 2026-09-21T22:22:05 - Migration needed. Performing migration...
mealie | INFO 2026-09-21T22:22:05 - Context impl SQLiteImpl.
mealie | INFO 2026-09-21T22:22:05 - Will assume non-transactional DDL.
mealie | INFO 2026-09-21T22:22:05 - Running upgrade 2187537c52b8 -> 69e942bab3aa, add tokens valid after column to users
mealie | INFO 2026-09-21T22:22:05 - Running upgrade 69e942bab3aa -> b3f1c9a27d84, add external avatar hash to users
mealie | INFO 2026-09-21T22:22:05 - Running upgrade b3f1c9a27d84 -> f2191b69db2e, add ingredient substitutions
mealie | INFO 2026-09-21T22:22:05 - Running upgrade f2191b69db2e -> 4b91d3a7c0e2, backfill recipe image column from disk
mealie | INFO 2026-09-21T22:22:05 - Recipe image backfill checked 1 recipes: 0 image references restored, 0 cleared
mealie | INFO 2026-09-21T22:22:05 - Running upgrade 4b91d3a7c0e2 -> 3527efeeec34, 'add recipe_note_ref_link'
mealie | INFO 2026-09-21T22:22:05 - Checking for migration data fixes
mealie | INFO 2026-09-21T22:22:05 - end: database initialization
mealie | INFO 2026-09-21T22:22:05 - -----SYSTEM STARTUP-----
mealie | INFO 2026-09-21T22:22:05 - ------APP SETTINGS------
mealie | INFO 2026-09-21T22:22:05 - {
mealie | "TESTING": false,
mealie | "PRODUCTION": true,
mealie | "LOG_CONFIG_OVERRIDE": null,
@@ -47,10 +58,12 @@ mealie | "API_HOST": "0.0.0.0",
mealie | "API_PORT": 9000,
mealie | "API_DOCS": true,
mealie | "TOKEN_TIME": 48,
mealie | "GIT_COMMIT_HASH": "a562409e43964d3734eabd64e4a2fd64bffaab4d",
mealie | "GIT_COMMIT_HASH": "dedc6cc75f2bd8e89108ad998970e7bdf6d4444f",
mealie | "ALLOW_SIGNUP": false,
mealie | "ALLOW_PASSWORD_LOGIN": true,
mealie | "ALLOWED_IFRAME_HOSTS": "",
mealie | "HTTP_ALLOW_LIST": "",
mealie | "HTTP_DISALLOW_LIST": "",
mealie | "DAILY_SCHEDULE_TIME": "23:45",
mealie | "SECURITY_MAX_LOGIN_ATTEMPTS": 5,
mealie | "SECURITY_USER_LOCKOUT_TIME": 24,
@@ -83,6 +96,7 @@ mealie | "OIDC_CLIENT_ID": null,
mealie | "OIDC_CLIENT_SECRET": null,
mealie | "OIDC_CONFIGURATION_URL": null,
mealie | "OIDC_SIGNUP_ENABLED": true,
mealie | "OIDC_REQUIRES_EMAIL_VERIFICATION": true,
mealie | "OIDC_USER_GROUP": null,
mealie | "OIDC_ADMIN_GROUP": null,
mealie | "OIDC_AUTO_REDIRECT": false,
@@ -95,27 +109,32 @@ mealie | "OIDC_SCOPES_OVERRIDE": null,
mealie | "OIDC_TLS_CACERTFILE": null,
mealie | "OIDC_CLIENT_TIMEOUT": "default",
mealie | "OPENAI_CUSTOM_PROMPT_DIR": null,
mealie | "SCRAPER_PROXY_URL": null,
mealie | "SCRAPER_PROXY_MODE": "always",
mealie | "SCRAPER_FLARESOLVERR_URL": null,
mealie | "SCRAPER_FLARESOLVERR_TIMEOUT": 60,
mealie | "WORKER_PER_CORE": 1,
mealie | "UVICORN_WORKERS": 1,
mealie | "TLS_CERTIFICATE_PATH": null,
mealie | "TLS_PRIVATE_KEY_PATH": null
mealie | "TLS_PRIVATE_KEY_PATH": null,
mealie | "YTDLP_COOKIEFILE": null
mealie | }
mealie | INFO 2026-09-21T21:24:25 - ------APP FEATURES------
mealie | INFO 2026-09-21T21:24:25 - --------==SMTP==--------
mealie | INFO 2026-09-21T21:24:25 - Enabled: False
mealie | INFO 2026-09-21T21:24:25 - --------==LDAP==--------
mealie | INFO 2026-09-21T21:24:25 - Enabled: False
mealie | INFO 2026-09-21T22:22:05 - ------APP FEATURES------
mealie | INFO 2026-09-21T22:22:05 - --------==SMTP==--------
mealie | INFO 2026-09-21T22:22:05 - Enabled: False
mealie | INFO 2026-09-21T22:22:05 - --------==LDAP==--------
mealie | INFO 2026-09-21T22:22:05 - Enabled: False
mealie | Reason: LDAP_AUTH_ENABLED is false
mealie | INFO 2026-09-21T21:24:25 - --------==OIDC==--------
mealie | INFO 2026-09-21T21:24:25 - Enabled: False
mealie | INFO 2026-09-21T22:22:05 - --------==OIDC==--------
mealie | INFO 2026-09-21T22:22:05 - Enabled: False
mealie | Reason: OIDC_AUTH_ENABLED is false
mealie | INFO 2026-09-21T21:24:25 - ------------------------
mealie | INFO 2026-09-21T21:24:25 - Daily tasks scheduled for 2026-09-21 21:45:00+00:00
mealie | INFO 2026-09-21T21:24:25 - Application startup complete.
mealie | INFO 2026-09-21T21:24:25 - Uvicorn running on http://0.0.0.0:9000 (Press CTRL+C to quit)
mealie | INFO 2026-09-21T21:25:08 - [192.168.0.180:0] 200 OK "GET /api/app/about HTTP/1.1"
mealie | INFO 2026-09-21T21:25:08 - [192.168.0.180:0] 200 OK "POST /api/auth/token HTTP/1.1"
mealie | ERROR 2026-09-21T21:25:08 - No Entry Found on recipe controller action
mealie | ERROR 2026-09-21T21:25:08 - No Entry Found on recipe controller action
mealie | INFO 2026-09-21T21:25:08 - [192.168.0.180:0] 404 Not Found "GET /api/recipes/nopefec546301e HTTP/1.1"
mealie | INFO 2026-09-21T21:25:08 - [192.168.0.180:0] 200 OK "GET /api/recipes/drill-087e4023c5 HTTP/1.1"
mealie | INFO 2026-09-21T22:22:05 - ------------------------
mealie | INFO 2026-09-21T22:22:05 - Daily tasks scheduled for 2026-09-21 21:45:00+00:00
mealie | INFO 2026-09-21T22:22:05 - Application startup complete.
mealie | INFO 2026-09-21T22:22:05 - Uvicorn running on http://0.0.0.0:9000 (Press CTRL+C to quit)
mealie | INFO 2026-09-21T22:22:13 - [192.168.0.180:0] 200 OK "GET /api/app/about HTTP/1.1"
mealie | INFO 2026-09-21T22:22:14 - [192.168.0.180:0] 200 OK "POST /api/auth/token HTTP/1.1"
mealie | ERROR 2026-09-21T22:22:14 - No Entry Found on recipe controller action
mealie | ERROR 2026-09-21T22:22:14 - No Entry Found on recipe controller action
mealie | INFO 2026-09-21T22:22:14 - [192.168.0.180:0] 404 Not Found "GET /api/recipes/nope988c526fa6 HTTP/1.1"
mealie | INFO 2026-09-21T22:22:15 - [192.168.0.180:0] 200 OK "GET /api/recipes/drill-2dc46249e8 HTTP/1.1"
@@ -16,16 +16,16 @@
"after": {
"hu": [
{
"title": "Ez az alkalmazás a legfrissebb elérhető változatot futtatja.",
"text": "Naprakész"
"title": "Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot.",
"text": "Frissítés elérhető — ma"
}
],
"en": [
{
"title": "This app is running the newest version available.",
"text": "Up to date"
"title": "A newer version of this app is available. Select the Update button to start it.",
"text": "Update available — today"
}
]
},
"drill_commit": "91e4bb97212e"
"drill_commit": "004105af9a16"
}
@@ -1,11 +1,24 @@
21:31:06 ==== mealie: ghcr.io/mealie-recipes/mealie:v3.20.1 -> ghcr.io/mealie-recipes/mealie:v3.27.0 (sub=mealie, class=db-postgres)
21:31:07 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
21:32:17 [1] deployed, controller state=running, pinned={'mealie': 'ghcr.io/mealie-recipes/mealie:v3.27.0'}
21:32:20 [2] seeding through the app's own front door
21:32:21 mealie: create recipe http=201
21:32:21 [3] control C1 — reading the seed back BEFORE the update
21:32:21 mealie: readback of the seeded recipe http=200 ok=True
21:32:21 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'}
21:33:06 [4] backup idle; last=None
21:33:06 [5] FROM ref not found in compose: ghcr.io/mealie-recipes/mealie:v3.20.1
21:33:06 [9] verdict inconclusive -> /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/update-night-2026-09-21/apps/mealie/verdict.json
22:20:09 ==== mealie: ghcr.io/mealie-recipes/mealie:v3.20.1 -> ghcr.io/mealie-recipes/mealie:v3.27.0 (sub=mealie, class=db-postgres)
22:20:09 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
22:20:29 [1] deployed, controller state=running, pinned={'mealie': 'ghcr.io/mealie-recipes/mealie:v3.20.1'}
22:20:32 [2] seeding through the app's own front door
22:20:33 mealie: create recipe http=201
22:20:33 [3] control C1 — reading the seed back BEFORE the update
22:20:34 mealie: readback of the seeded recipe http=200 ok=True
22:20:34 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'}
22:21:49 [4] backup idle; last=None
22:21:50 [5] drill commit 004105af9a16: mealie ghcr.io/mealie-recipes/mealie:v3.20.1 -> ghcr.io/mealie-recipes/mealie:v3.27.0 (push rc=0)
22:21:55 [5] badge HU: [{'title': 'Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot.', 'text': 'Frissítés elérhető — ma'}]
22:21:55 [5] badge EN: [{'title': 'A newer version of this app is available. Select the Update button to start it.', 'text': 'Update available — today'}]
22:21:55 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
22:21:55 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None
22:21:56 + 1.1s phase=starting label=Indítás az új verzióval… err=None hold=None
22:21:58 + 3.1s phase=verifying label=Működés ellenőrzése… err=None hold=None
22:22:13 + 18.5s phase=done label=Frissítve err=None hold=None
22:22:13 [7] reading the seed back AFTER the update
22:22:15 mealie: readback of the seeded recipe http=200 ok=True
22:22:17 [8] pinned = {'mealie': 'ghcr.io/mealie-recipes/mealie:v3.27.0'}
22:22:17 [8] installed = {'mealie': 'ghcr.io/mealie-recipes/mealie:v3.27.0'}
22:22:17 [8] compose = ['image: ghcr.io/mealie-recipes/mealie:v3.27.0']
22:22:17 [8] inspect = ['mealie ghcr.io/mealie-recipes/mealie:v3.27.0 running=true restarts=0']
22:22:17 [9] verdict proven -> /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/update-night-2026-09-21/apps/mealie/verdict.json
@@ -1,15 +1,15 @@
{
"pinned_images": {
"mealie": "ghcr.io/mealie-recipes/mealie:v3.27.0"
"mealie": "ghcr.io/mealie-recipes/mealie:v3.20.1"
},
"installed_images": {
"mealie": "ghcr.io/mealie-recipes/mealie:v3.27.0"
"mealie": "ghcr.io/mealie-recipes/mealie:v3.20.1"
},
"catalog_images": null,
"live_compose_image_lines": [
"image: ghcr.io/mealie-recipes/mealie:v3.27.0"
"image: ghcr.io/mealie-recipes/mealie:v3.20.1"
],
"docker_inspect": [
"mealie ghcr.io/mealie-recipes/mealie:v3.27.0 running=true restarts=0"
"mealie ghcr.io/mealie-recipes/mealie:v3.20.1 running=true restarts=0"
]
}
@@ -6,9 +6,7 @@
"installed_images": {
"mealie": "ghcr.io/mealie-recipes/mealie:v3.20.1"
},
"catalog_images": {
"mealie": "ghcr.io/mealie-recipes/mealie:v3.20.1"
},
"catalog_images": null,
"live_compose_image_lines": [
"image: ghcr.io/mealie-recipes/mealie:v3.20.1"
],
@@ -18,19 +16,19 @@
},
"after": {
"pinned_images": {
"mealie": "ghcr.io/mealie-recipes/mealie:v3.20.1"
"mealie": "ghcr.io/mealie-recipes/mealie:v3.27.0"
},
"installed_images": {
"mealie": "ghcr.io/mealie-recipes/mealie:v3.20.1"
"mealie": "ghcr.io/mealie-recipes/mealie:v3.27.0"
},
"catalog_images": {
"mealie": "ghcr.io/mealie-recipes/mealie:v3.20.1"
"mealie": "ghcr.io/mealie-recipes/mealie:v3.27.0"
},
"live_compose_image_lines": [
"image: ghcr.io/mealie-recipes/mealie:v3.20.1"
"image: ghcr.io/mealie-recipes/mealie:v3.27.0"
],
"docker_inspect": [
"mealie ghcr.io/mealie-recipes/mealie:v3.20.1 running=true restarts=0"
"mealie ghcr.io/mealie-recipes/mealie:v3.27.0 running=true restarts=0"
]
}
}
@@ -12,6 +12,14 @@
},
{
"t": 1.1,
"phase": "starting",
"label": "Indítás az új verzióval…",
"updating": true,
"error": null,
"hold": null
},
{
"t": 3.1,
"phase": "verifying",
"label": "Működés ellenőrzése…",
"updating": true,
@@ -19,7 +27,7 @@
"hold": null
},
{
"t": 2.1,
"t": 18.5,
"phase": "done",
"label": "Frissítve",
"updating": false,
@@ -27,7 +35,7 @@
"hold": null
}
],
"duration_s": 3.1,
"duration_s": 18.5,
"final_phase": "done",
"update_error": null,
"hold_reason": null,
@@ -4,20 +4,41 @@
"venue": "guest 9202 demo-hp-scratch, controller 0.261.0",
"class": "db-postgres",
"from": {
"mealie": "ghcr.io/mealie-recipes/mealie:v3.20.1"
},
"to": {
"mealie": "ghcr.io/mealie-recipes/mealie:v3.27.0"
},
"to": {},
"verdict": "inconclusive",
"verdict": "proven",
"seed_read_before": true,
"seed_read_after": false,
"healthy_after": false,
"migration_observed": null,
"seed_read_after": true,
"healthy_after": true,
"migration_observed": "mealie | INFO 2026-09-21T22:22:05 - Migration needed. Performing migration...",
"abort": "not-attempted",
"abort_detail": null,
"duration_s": 120.0,
"measured_at": "2026-09-21T19:31:06.937002+00:00",
"duration_s": 18.5,
"measured_at": "2026-09-21T20:20:09.681676+00:00",
"evidence": "apps/mealie/",
"notes": [
"the drill bump could not be committed — the FROM ref did not match the template"
]
"notes": [],
"badge_catchup_seconds": 4.4,
"observables_after": {
"pinned_images": {
"mealie": "ghcr.io/mealie-recipes/mealie:v3.27.0"
},
"installed_images": {
"mealie": "ghcr.io/mealie-recipes/mealie:v3.27.0"
},
"catalog_images": {
"mealie": "ghcr.io/mealie-recipes/mealie:v3.27.0"
},
"live_compose_image_lines": [
"image: ghcr.io/mealie-recipes/mealie:v3.27.0"
],
"docker_inspect": [
"mealie ghcr.io/mealie-recipes/mealie:v3.27.0 running=true restarts=0"
]
},
"final_phase": "done",
"hold_reason": null,
"update_error": null
}
@@ -0,0 +1,30 @@
=== the app's own recovery unit + its db dumps, with sizes and times
2026-09-21 20:06 288 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/app.yaml
2026-09-21 20:06 1050 /mnt/sys_drive/felhom-data/backups/primary/privatebin/manifest.json
2026-09-21 20:06 1235 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/docker-compose.yml
2026-09-21 20:06 3125 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/.felhom.yml
2026-09-21 20:07 1536 /mnt/sys_drive/felhom-data/backups/primary/privatebin/volume-dumps/privatebin_privatebin_data.tar
=== any temp/partial names left behind
=== the unit manifest, if there is one
--- /mnt/sys_drive/felhom-data/backups/primary/privatebin/manifest.json
{
"schema_version": 2,
"app_name": "privatebin",
"display_name": "PrivateBin",
"controller_version": "0.261.0",
"created_at": "2026-09-21T20:06:16Z",
"drive": "/mnt/sys_drive",
"namespace_root": "/mnt/sys_drive/felhom-data",
"image_pins": [
"localhost:5000/drill/paste:2.0.1"
],
"secret_env_vars": null,
"data_key_env_vars": null,
"secret_source": "portable secrets (data keys, DB passwords, internal signing secrets) are IN this unit's compose/app.yaml (0600); internet-reachable admin logins are NOT, and come from the guest's app.yaml or are regenerated on restore",
"config_files": [
"docker-compose.yml",
".felhom.yml",
"app.yaml"
],
"db_dumps": [],
=== ZERO-BYTE files under this app's backups (a half-write that still looks like a file)
@@ -0,0 +1,30 @@
=== the app's own recovery unit + its db dumps, with sizes and times
2026-09-21 20:06 288 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/app.yaml
2026-09-21 20:06 1050 /mnt/sys_drive/felhom-data/backups/primary/privatebin/manifest.json
2026-09-21 20:06 1235 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/docker-compose.yml
2026-09-21 20:06 3125 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/.felhom.yml
2026-09-21 20:07 1536 /mnt/sys_drive/felhom-data/backups/primary/privatebin/volume-dumps/privatebin_privatebin_data.tar
=== any temp/partial names left behind
=== the unit manifest, if there is one
--- /mnt/sys_drive/felhom-data/backups/primary/privatebin/manifest.json
{
"schema_version": 2,
"app_name": "privatebin",
"display_name": "PrivateBin",
"controller_version": "0.261.0",
"created_at": "2026-09-21T20:06:16Z",
"drive": "/mnt/sys_drive",
"namespace_root": "/mnt/sys_drive/felhom-data",
"image_pins": [
"localhost:5000/drill/paste:2.0.1"
],
"secret_env_vars": null,
"data_key_env_vars": null,
"secret_source": "portable secrets (data keys, DB passwords, internal signing secrets) are IN this unit's compose/app.yaml (0600); internet-reachable admin logins are NOT, and come from the guest's app.yaml or are regenerated on restore",
"config_files": [
"docker-compose.yml",
".felhom.yml",
"app.yaml"
],
"db_dumps": [],
=== ZERO-BYTE files under this app's backups (a half-write that still looks like a file)
@@ -0,0 +1,16 @@
22:19:17 ==== B5: a power cut during `safety-dump` on privatebin
22:19:17 [0] pending edge for the cut: installed={'privatebin': 'localhost:5000/drill/paste:2.0.1'} catalog={'privatebin': 'localhost:5000/drill/paste:2.0.2'}
22:19:21 [0] pinned before = {'privatebin': 'localhost:5000/drill/paste:2.0.1'}
22:19:22 [cut] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követ
22:19:22 + 0.023s phase=safety-dump
22:19:22 [cut] phase 'safety-dump' OBSERVED at 2026-09-21T20:19:22.234Z — pulling the plug NOW
22:19:26 [cut] `pct stop 9202` returned after 3.96s ::
22:19:29 [boot] guest 9202 starting
22:19:42 [boot] controller back: Up 6 seconds (healthy)
22:20:05 [boot] recovery lines: 2026/09/21 20:19:36 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false) | 2026/09/21 20:20:01 bootrecon.go:236: [INFO] [bootrecon] "glance" is a boot orphan by intent but is HELD (held after a failed update (2026-09-21T19:53:02Z) — restore it from its backup to start it) — NOT starting it; whatever is holding it owns its recovery | --- journal file: | ls: cannot access '/var/lib/docker/volumes/f
22:20:09 [2] household sentence HU: ['PrivateBin — Felhom.eu Indítópult Vezérlőpult Alkalmazások Tárhely Meghajtók Hálózati tárhely Biztonsági mentés Áttekintés Távoli mentés Alkalmazások Visszaállítás Megosztás Hálózati megosztás Rendszermonitor Debug Beállítások Rendszer Értesítések Biztonság és hozzáférés 0.261.0 Magyar English Kijelentkezés ↗ Hub kapcsolat kikapcsolva — a központi monitoring nem aktív Rendszermonitor → ← Alkalmazások PrivateBin Fut Frissítés elérhető — ma Megnyitás ↗ Napló Exportálás Beállítások Titkosított szöveg megosztás - a szerver nem látja a tartalmat ~30M RAM security Pi kompatibilis Áthelyezés másik tárhelyre Ennek az alkalmazásnak az adatait másik csatlakoztatott tárhelyre helyezheted át.', 'Érzékeny szövegek biztonságos megosztása E2E titkosítás - a szerver nem fér hozzá a tartalomhoz Beállítható lejárati idő (5 perc - 1 év, vagy soha) Olvasás után automatikus törlés opció Jelszóvédelem a még nagyobb biztonságért Első lépések Nyisd meg a paste.DOMAIN címet a böngészőben Írd be a szöveget és kattints a Küldés gombra Oszd meg a generált linket - a titkosítási kulcs az URL-ben van Dokumentáció Hivatalos dokumentáció ↗']
22:20:09 [2] household sentence EN: ['PrivateBin — Felhom.eu Launcher Dashboard Apps Storage Drives Network storage Backup Overview Remote backup Apps Restore Sharing Network sharing System monitor Debug Settings System Notifications Security and access 0.261.0 Magyar English Sign out ↗ The hub connection is off — central monitoring is not running System monitor → ← Apps PrivateBin Running Update available — today Open ↗ Log Export Settings Encrypted text sharing - the server never sees the content ~30M RAM security Runs on Pi Move to another storage You can move this app’s data to another connected storage.', 'Share sensitive text safely End-to-end encryption - the server cannot reach the content You choose how long it lasts (5 minutes to 1 year, or never) Optional: delete it automatically after it is read Password protection for extra safety First steps Open paste.DOMAIN in your browser Type your text and select Send Share the link you get - the encryption key is part of the URL Documentation Official documentation ↗']
22:20:09 privatebin: seeded paste id=5d8afde576872b63
22:20:09 privatebin: readback http=200 marker_present=True
22:20:09 [3] the app works and holds data after the cut: True
22:20:09 [4] pinned after = {'privatebin': 'localhost:5000/drill/paste:2.0.1'} (moved: False)
@@ -0,0 +1,72 @@
{
"leg": "B5-safety-dump-retry",
"app": "privatebin",
"phase_targeted": "safety-dump",
"observables_before": {
"pinned_images": {
"privatebin": "localhost:5000/drill/paste:2.0.1"
},
"installed_images": {
"privatebin": "localhost:5000/drill/paste:2.0.1"
},
"catalog_images": {
"privatebin": "localhost:5000/drill/paste:2.0.2"
},
"live_compose_image_lines": [
"image: localhost:5000/drill/paste:2.0.1"
],
"docker_inspect": [
"privatebin localhost:5000/drill/paste:2.0.1 running=true restarts=0"
]
},
"backup_artefacts_before": "=== the app's own recovery unit + its db dumps, with sizes and times\n2026-09-21 20:06 288 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/app.yaml\n2026-09-21 20:06 1050 /mnt/sys_drive/felhom-data/backups/primary/privatebin/manifest.json\n2026-09-21 20:06 1235 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/docker-compose.yml\n2026-09-21 20:06 3125 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/.felhom.yml\n2026-09-21 20:07 1536 /mnt/sys_drive/felhom-data/backups/primary/privatebin/volume-dumps/privatebin_privatebin_data.tar\n=== any temp/partial names left behind\n=== the unit manifest, if there is one\n--- /mnt/sys_drive/felhom-data/backups/primary/privatebin/manifest.json\n{\n \"schema_version\": 2,\n \"app_name\": \"privatebin\",\n \"display_name\": \"PrivateBin\",\n \"controller_version\": \"0.261.0\",\n \"created_at\": \"2026-09-21T20:06:16Z\",\n \"drive\": \"/mnt/sys_drive\",\n \"namespace_root\": \"/mnt/sys_drive/felhom-data\",\n \"image_pins\": [\n \"localhost:5000/drill/paste:2.0.1\"\n ],\n \"secret_env_vars\": null,\n \"data_key_env_vars\": null,\n \"secret_source\": \"portable secrets (data keys, DB passwords, internal signing secrets) are IN this unit's compose/app.yaml (0600); internet-reachable admin logins are NOT, and come from the guest's app.yaml or are regenerated on restore\",\n \"config_files\": [\n \"docker-compose.yml\",\n \".felhom.yml\",\n \"app.yaml\"\n ],\n \"db_dumps\": [],\n=== ZERO-BYTE files under this app's backups (a half-write that still looks like a file)\n",
"cut": {
"pressed": true,
"phases_seen": [
{
"t": 0.023,
"phase": "safety-dump"
}
],
"cut_decided_at": "2026-09-21T20:19:22.234Z",
"pct_stop_returned_after_s": 3.96,
"instrument_limit": "the phase at the DECISION is observed; the phase at the FREEZE is inferred — pct stop is not instantaneous"
},
"after_boot": {
"recovery_lines": "2026/09/21 20:19:36 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false)\n2026/09/21 20:20:01 bootrecon.go:236: [INFO] [bootrecon] \"glance\" is a boot orphan by intent but is HELD (held after a failed update (2026-09-21T19:53:02Z) — restore it from its backup to start it) — NOT starting it; whatever is holding it owns its recovery\n--- journal file:\nls: cannot access '/var/lib/docker/volumes/felhom-controller-data/_data/data/update-journal.json': No such file or directory\n",
"state": "running",
"update_phase": null,
"update_error": null,
"hold_reason": null,
"observables": {
"pinned_images": {
"privatebin": "localhost:5000/drill/paste:2.0.1"
},
"installed_images": {
"privatebin": "localhost:5000/drill/paste:2.0.1"
},
"catalog_images": {
"privatebin": "localhost:5000/drill/paste:2.0.2"
},
"live_compose_image_lines": [
"image: localhost:5000/drill/paste:2.0.1"
],
"docker_inspect": [
"privatebin localhost:5000/drill/paste:2.0.1 running=true restarts=0"
]
}
},
"backup_artefacts_after": "=== the app's own recovery unit + its db dumps, with sizes and times\n2026-09-21 20:06 288 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/app.yaml\n2026-09-21 20:06 1050 /mnt/sys_drive/felhom-data/backups/primary/privatebin/manifest.json\n2026-09-21 20:06 1235 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/docker-compose.yml\n2026-09-21 20:06 3125 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/.felhom.yml\n2026-09-21 20:07 1536 /mnt/sys_drive/felhom-data/backups/primary/privatebin/volume-dumps/privatebin_privatebin_data.tar\n=== any temp/partial names left behind\n=== the unit manifest, if there is one\n--- /mnt/sys_drive/felhom-data/backups/primary/privatebin/manifest.json\n{\n \"schema_version\": 2,\n \"app_name\": \"privatebin\",\n \"display_name\": \"PrivateBin\",\n \"controller_version\": \"0.261.0\",\n \"created_at\": \"2026-09-21T20:06:16Z\",\n \"drive\": \"/mnt/sys_drive\",\n \"namespace_root\": \"/mnt/sys_drive/felhom-data\",\n \"image_pins\": [\n \"localhost:5000/drill/paste:2.0.1\"\n ],\n \"secret_env_vars\": null,\n \"data_key_env_vars\": null,\n \"secret_source\": \"portable secrets (data keys, DB passwords, internal signing secrets) are IN this unit's compose/app.yaml (0600); internet-reachable admin logins are NOT, and come from the guest's app.yaml or are regenerated on restore\",\n \"config_files\": [\n \"docker-compose.yml\",\n \".felhom.yml\",\n \"app.yaml\"\n ],\n \"db_dumps\": [],\n=== ZERO-BYTE files under this app's backups (a half-write that still looks like a file)\n",
"sentences": {
"hu": [
"PrivateBin — Felhom.eu Indítópult Vezérlőpult Alkalmazások Tárhely Meghajtók Hálózati tárhely Biztonsági mentés Áttekintés Távoli mentés Alkalmazások Visszaállítás Megosztás Hálózati megosztás Rendszermonitor Debug Beállítások Rendszer Értesítések Biztonság és hozzáférés 0.261.0 Magyar English Kijelentkezés ↗ Hub kapcsolat kikapcsolva — a központi monitoring nem aktív Rendszermonitor → ← Alkalmazások PrivateBin Fut Frissítés elérhető — ma Megnyitás ↗ Napló Exportálás Beállítások Titkosított szöveg megosztás - a szerver nem látja a tartalmat ~30M RAM security Pi kompatibilis Áthelyezés másik tárhelyre Ennek az alkalmazásnak az adatait másik csatlakoztatott tárhelyre helyezheted át.",
"Érzékeny szövegek biztonságos megosztása E2E titkosítás - a szerver nem fér hozzá a tartalomhoz Beállítható lejárati idő (5 perc - 1 év, vagy soha) Olvasás után automatikus törlés opció Jelszóvédelem a még nagyobb biztonságért Első lépések Nyisd meg a paste.DOMAIN címet a böngészőben Írd be a szöveget és kattints a Küldés gombra Oszd meg a generált linket - a titkosítási kulcs az URL-ben van Dokumentáció Hivatalos dokumentáció ↗"
],
"en": [
"PrivateBin — Felhom.eu Launcher Dashboard Apps Storage Drives Network storage Backup Overview Remote backup Apps Restore Sharing Network sharing System monitor Debug Settings System Notifications Security and access 0.261.0 Magyar English Sign out ↗ The hub connection is off — central monitoring is not running System monitor → ← Apps PrivateBin Running Update available — today Open ↗ Log Export Settings Encrypted text sharing - the server never sees the content ~30M RAM security Runs on Pi Move to another storage You can move this app’s data to another connected storage.",
"Share sensitive text safely End-to-end encryption - the server cannot reach the content You choose how long it lasts (5 minutes to 1 year, or never) Optional: delete it automatically after it is read Password protection for extra safety First steps Open paste.DOMAIN in your browser Type your text and select Send Share the link you get - the encryption key is part of the URL Documentation Official documentation ↗"
]
},
"app_usable_after": true,
"pin_moved": false
}
@@ -0,0 +1,70 @@
22:13:25 ==== Phase 2.3: the PostgreSQL 16 -> 17 conversion rehearsal (Q5)
22:13:25 removing the existing docmost so the rehearsal meets a FRESH workspace it can seed
22:13:27 [X] stop -> 200 {'ok': True, 'message': 'Stack docmost stop completed'}
22:13:32 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'docmost', 'volumes_removed': ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage'], 'hdd_
22:13:40 [X] after remove: deployed=False leftovers='/opt/docker/stacks/docmost'
22:13:40 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
22:14:05 [1] deployed, controller state=running, pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'}
22:14:05 docmost: /api/auth/setup http=200 rc=0
22:14:06 docmost: login as the seeded user http=200 ok=True
22:14:09 [00-state-before] 2.7s
22:14:09 PG_VERSION (the harness's own postgres probe, verbatim):
22:14:09 16
22:14:09 engine version:
22:14:09 postgres (PostgreSQL) 16.15
22:14:09 datadir size:
22:14:09 49.0M /var/lib/postgresql/data
22:14:09 volume:
22:14:09 docmost_docmost_postgres_data
22:14:12 [01-stop-the-app-keep-the-engine] 3.2s
22:14:12 app stopped (the engine stays up to be dumped)
22:14:12 docmost-postgres Up 31 seconds (healthy)
22:14:12 docmost-redis Up 31 seconds (healthy)
22:14:15 [02-dump-with-16] 2.6s
22:14:15 dumping as user=docmost
22:14:15 rc=0
22:14:15 dump bytes: 132201
22:14:15 CREATE TABLE statements: 48
22:14:21 [03-fresh-17-datadir-and-restore] 6.5s
22:14:21 17 up: postgres (PostgreSQL) 17.11
22:14:21 PG_VERSION on the fresh datadir: 17
22:14:21 restore rc=0
22:14:21 ERROR lines in the restore: 2
22:14:21 ERROR: role "docmost" already exists
22:14:21 ERROR: database "docmost" already exists
22:14:21 tables restored:
22:14:21 48
22:16:26 [04-point-the-app-at-17-and-start-it] 124.8s
22:16:26 DRILL-pg17 now answers to the name docmost-postgres on docmost_docmost-internal
22:16:26 app started
22:16:26 docmost Up 2 minutes (healthy)
22:16:26 docmost-postgres Exited (0) 2 minutes ago
22:16:26 docmost-redis Up 2 minutes (healthy)
22:16:26 {"level":"info","time":"2026-09-21T20:14:00.698Z","pid":45,"hostname":"24a35fe9b998","context":"NestApplication","msg":"Listening on http://127.0.0.1:3000 / https://docs.enkisfelhom.hu"}
22:16:26 [ELIFECYCLE] Command failed.
22:16:26 $ pnpm --filter ./apps/server run start:prod
22:16:26 $ cross-env NODE_ENV=production node dist/main
22:16:26 (node:45) ExperimentalWarning: localStorage is not available because --localstorage-file was not provided.
22:16:26 (Use `node --trace-warnings ...` to show where the warning was created)
22:16:26 {"level":"info","time":"2026-09-21T20:14:31.727Z","pid":45,"hostname":"24a35fe9b998","context":"RedisModule","msg":"default: the connection was successfully established"}
22:16:26 {"level":"info","time":"2026-09-21T20:14:32.033Z","pid":45,"hostname":"24a35fe9b998","context":"DatabaseModule","msg":"Establishing database connection"}
22:16:26 {"level":"info","time":"2026-09-21T20:14:32.065Z","pid":45,"hostname":"24a35fe9b998","context":"DatabaseModule","msg":"Database connection successful"}
22:16:27 docmost: login as the seeded user http=200 ok=True
22:16:27 [05] the seed read back on PostgreSQL 17: True
22:16:29 [06-engine-state-after] 2.3s
22:16:29 the harness's own postgres probe against the CONVERTED datadir:
22:16:29 17
22:16:29 [exit=0]
22:16:29 postgres (PostgreSQL) 17.11
22:16:29 size of the 17 datadir:
22:16:29 49.1M /var/lib/postgresql/data
22:16:43 [99-put-everything-back] 13.8s
22:16:43 docmost docmost/docmost:0.96.0 Up 5 seconds (health: starting)
22:16:43 docmost-postgres postgres:16-alpine Up 10 seconds (healthy)
22:16:43 docmost-redis redis:7-alpine Up 3 minutes (healthy)
22:16:43 docmost-postgres postgres:16-alpine Up 10 seconds (healthy)
22:16:43 PG_VERSION back on the original datadir: 16
22:16:43 docmost: login as the seeded user http=404 ok=False
22:16:43 docmost: login body 404 page not found
22:16:43 [99] the seed still reads on the ORIGINAL 16 datadir after teardown: False
@@ -2,9 +2,47 @@
"leg": "P2.3",
"app": "docmost",
"route": "logical dump and restore",
"measured_at": "2026-09-21T20:11:46.041578+00:00",
"steps": [],
"notes": [
"could not seed before the rehearsal — see log"
]
"measured_at": "2026-09-21T20:13:25.874392+00:00",
"steps": [
{
"label": "00-state-before",
"seconds": 2.7,
"output": "PG_VERSION (the harness's own postgres probe, verbatim):\n16\nengine version:\npostgres (PostgreSQL) 16.15\ndatadir size:\n49.0M\t/var/lib/postgresql/data\nvolume:\ndocmost_docmost_postgres_data\n"
},
{
"label": "01-stop-the-app-keep-the-engine",
"seconds": 3.2,
"output": "app stopped (the engine stays up to be dumped)\ndocmost-postgres Up 31 seconds (healthy)\ndocmost-redis Up 31 seconds (healthy)\n"
},
{
"label": "02-dump-with-16",
"seconds": 2.6,
"output": "dumping as user=docmost\nrc=0\ndump bytes: 132201\nCREATE TABLE statements: 48\n"
},
{
"label": "03-fresh-17-datadir-and-restore",
"seconds": 6.5,
"output": "17 up: postgres (PostgreSQL) 17.11\nPG_VERSION on the fresh datadir: 17\nrestore rc=0\nERROR lines in the restore: 2\nERROR: role \"docmost\" already exists\nERROR: database \"docmost\" already exists\ntables restored:\n48\n"
},
{
"label": "04-point-the-app-at-17-and-start-it",
"seconds": 124.8,
"output": "DRILL-pg17 now answers to the name docmost-postgres on docmost_docmost-internal\napp started\ndocmost Up 2 minutes (healthy)\ndocmost-postgres Exited (0) 2 minutes ago\ndocmost-redis Up 2 minutes (healthy)\n{\"level\":\"info\",\"time\":\"2026-09-21T20:14:00.698Z\",\"pid\":45,\"hostname\":\"24a35fe9b998\",\"context\":\"NestApplication\",\"msg\":\"Listening on http://127.0.0.1:3000 / https://docs.enkisfelhom.hu\"}\n[ELIFECYCLE] Command failed.\n$ pnpm --filter ./apps/server run start:prod\n$ cross-env NODE_ENV=production node dist/main\n(node:45) ExperimentalWarning: localStorage is not available because --localstorage-file was not provided.\n(Use `node --trace-warnings ...` to show where the warning was created)\n{\"level\":\"info\",\"time\":\"2026-09-21T20:14:31.727Z\",\"pid\":45,\"hostname\":\"24a35fe9b998\",\"context\":\"RedisModule\",\"msg\":\"default: the connection was successfully established\"}\n{\"level\":\"info\",\"time\":\"2026-09-21T20:14:32.033Z\",\"pid\":45,\"hostname\":\"24a35fe9b998\",\"context\":\"DatabaseModule\",\"msg\":\"Establishing database connection\"}\n{\"level\":\"info\",\"time\":\"2026-09-21T20:14:32.065Z\",\"pid\":45,\"hostname\":\"24a35fe9b998\",\"context\":\"DatabaseModule\",\"msg\":\"Database connection successful\"}\n{\"level\":\"info\",\"time\":\"2026-09-21T20:14:32.222Z\",\"pid\":45,\"hostname\":\"24a35fe9b998\",\"context\":\"DatabaseMigrationService\",\"msg\":\"No pending database migrations\"}\n{\"level\":\"info\",\"time\":\"2026-09-21T20:14:32.266Z\",\"pid\":45,\"hostname\":\"24a35fe9b998\",\"context\":\"NestApplication\",\"msg\":\"Nest application successfully started\"}\n{\"level\":\"info\",\"time\":\"2026-09-21T20:14:32.282Z\",\"pid\":45,\"hostname\":\"24a35fe9b998\",\"context\":\"NestApplication\",\"msg\":\"Listening on http://127.0.0.1:3000 / https://docs.enkisfelhom.hu\"}\n"
},
{
"label": "06-engine-state-after",
"seconds": 2.3,
"output": "the harness's own postgres probe against the CONVERTED datadir:\n17\n[exit=0]\npostgres (PostgreSQL) 17.11\nsize of the 17 datadir:\n49.1M\t/var/lib/postgresql/data\n"
},
{
"label": "99-put-everything-back",
"seconds": 13.8,
"output": "docmost docmost/docmost:0.96.0 Up 5 seconds (health: starting)\ndocmost-postgres postgres:16-alpine Up 10 seconds (healthy)\ndocmost-redis redis:7-alpine Up 3 minutes (healthy)\ndocmost-postgres postgres:16-alpine Up 10 seconds (healthy)\nPG_VERSION back on the original datadir: 16\n"
}
],
"notes": [],
"seed_read_before": true,
"seed_read_after_on_17": true,
"seed_read_back_on_16_after_teardown": false,
"total_seconds": 155.9
}
@@ -0,0 +1,74 @@
############################################## 22:16:43 L1 wger — the LAST unmeasured candidate in R-618, and the one that would close it
22:16:43 === wger — R-618's last unmeasured candidate
22:16:43 template probe : type=http port=80 (.felhom.yml)
22:16:43 traefik label : loadbalancer.server.port=8000 (the SAME template's compose)
22:16:45 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
22:18:26 [1] deployed, controller state=unhealthy, pinned={'wger': 'wger/server:2.6'}
22:18:26 [1] NOTE: the controller's own state is 'unhealthy', not 'running' — recorded, not treated as a failure; the fixture's front-door wait is the real gate
22:18:29 --- what wger actually LISTENS on, asked inside its own container:
22:18:29 --- the container's own docker healthcheck, if it has one:
22:18:29 healthy
22:18:29 --- port 80 from inside (the port the probe names):
22:18:29 refused on 80
22:18:29 --- port 8000 from inside (the port the compose says):
22:18:29 ANSWERED on 8000
22:18:29 the household's own front door: http=302
22:18:29 the box's own verdict: state='unhealthy'
22:18:29 VERDICT: CONFIRMED DEFECT — the same shape as tandoor
22:18:29 written -> /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/update-night-2026-09-21/22-wger-probe-measured.txt
22:18:39 [X] stop -> 200 {'ok': True, 'message': 'Stack wger stop completed'}
22:18:45 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'wger', 'volumes_removed': ['wger_wger_data', 'wger_wger_media'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note
22:18:52 [X] after remove: deployed=False leftovers='/opt/docker/stacks/wger'
############################################## 22:18:52 L2 B5 retry — the cut in safety-dump, with a REAL pending edge this time
22:19:17 [knob] backup_max_age=24h :: update: | backup_max_age: 24h | Up 22 seconds (healthy)
22:19:17 ==== B5: a power cut during `safety-dump` on privatebin
22:19:17 [0] pending edge for the cut: installed={'privatebin': 'localhost:5000/drill/paste:2.0.1'} catalog={'privatebin': 'localhost:5000/drill/paste:2.0.2'}
22:19:21 [0] pinned before = {'privatebin': 'localhost:5000/drill/paste:2.0.1'}
22:19:22 [cut] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követ
22:19:22 + 0.023s phase=safety-dump
/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/update-night-2026-09-21/phase3_b5.py:82: DeprecationWarning: datetime.datetime.utcnow() is deprecated and scheduled for removal in a future version. Use timezone-aware objects to represent datetimes in UTC: datetime.datetime.now(datetime.UTC).
decided = datetime.utcnow().isoformat(timespec="milliseconds") + "Z"
22:19:22 [cut] phase 'safety-dump' OBSERVED at 2026-09-21T20:19:22.234Z — pulling the plug NOW
22:19:26 [cut] `pct stop 9202` returned after 3.96s ::
22:19:29 [boot] guest 9202 starting
22:19:42 [boot] controller back: Up 6 seconds (healthy)
22:20:05 [boot] recovery lines: 2026/09/21 20:19:36 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false) | 2026/09/21 20:20:01 bootrecon.go:236: [INFO] [bootrecon] "glance" is a boot orphan by intent but is HELD (held after a failed update (2026-09-21T19:53:02Z) — restore it from its backup to start it) — NOT starting it; whatever is holding it owns its recovery | --- journal file: | ls: cannot access '/var/lib/docker/volumes/f
22:20:09 [2] household sentence HU: ['PrivateBin — Felhom.eu Indítópult Vezérlőpult Alkalmazások Tárhely Meghajtók Hálózati tárhely Biztonsági mentés Áttekintés Távoli mentés Alkalmazások Visszaállítás Megosztás Hálózati megosztás Rendszermonitor Debug Beállítások Rendszer Értesítések Biztonság és hozzáférés 0.261.0 Magyar English Kijelentkezés ↗ Hub kapcsolat kikapcsolva — a központi monitoring nem aktív Rendszermonitor → ← Alkalmazások PrivateBin Fut Frissítés elérhető — ma Megnyitás ↗ Napló Exportálás Beállítások Titkosított szöveg megosztás - a szerver nem látja a tartalmat ~30M RAM security Pi kompatibilis Áthelyezés másik tárhelyre Ennek az alkalmazásnak az adatait másik csatlakoztatott tárhelyre helyezheted át.', 'Érzékeny szövegek biztonságos megosztása E2E titkosítás - a szerver nem fér hozzá a tartalomhoz Beállítható lejárati idő (5 perc - 1 év, vagy soha) Olvasás után automatikus törlés opció Jelszóvédelem a még nagyobb biztonságért Első lépések Nyisd meg a paste.DOMAIN címet a böngészőben Írd be a szöveget és kattints a Küldés gombra Oszd meg a generált linket - a titkosítási kulcs az URL-ben van Dokumentáció Hivatalos dokumentáció ↗']
22:20:09 [2] household sentence EN: ['PrivateBin — Felhom.eu Launcher Dashboard Apps Storage Drives Network storage Backup Overview Remote backup Apps Restore Sharing Network sharing System monitor Debug Settings System Notifications Security and access 0.261.0 Magyar English Sign out ↗ The hub connection is off — central monitoring is not running System monitor → ← Apps PrivateBin Running Update available — today Open ↗ Log Export Settings Encrypted text sharing - the server never sees the content ~30M RAM security Runs on Pi Move to another storage You can move this app’s data to another connected storage.', 'Share sensitive text safely End-to-end encryption - the server cannot reach the content You choose how long it lasts (5 minutes to 1 year, or never) Optional: delete it automatically after it is read Password protection for extra safety First steps Open paste.DOMAIN in your browser Type your text and select Send Share the link you get - the encryption key is part of the URL Documentation Official documentation ↗']
22:20:09 privatebin: seeded paste id=5d8afde576872b63
22:20:09 privatebin: readback http=200 marker_present=True
22:20:09 [3] the app works and holds data after the cut: True
22:20:09 [4] pinned after = {'privatebin': 'localhost:5000/drill/paste:2.0.1'} (moved: False)
############################################## 22:20:09 L3 mealie — the edge that was never walked because its drill pin had already moved
22:20:09 ==== mealie: ghcr.io/mealie-recipes/mealie:v3.20.1 -> ghcr.io/mealie-recipes/mealie:v3.27.0 (sub=mealie, class=db-postgres)
22:20:09 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
22:20:29 [1] deployed, controller state=running, pinned={'mealie': 'ghcr.io/mealie-recipes/mealie:v3.20.1'}
22:20:32 [2] seeding through the app's own front door
22:20:33 mealie: create recipe http=201
22:20:33 [3] control C1 — reading the seed back BEFORE the update
22:20:34 mealie: readback of the seeded recipe http=200 ok=True
22:20:34 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'}
22:21:49 [4] backup idle; last=None
22:21:50 [5] drill commit 004105af9a16: mealie ghcr.io/mealie-recipes/mealie:v3.20.1 -> ghcr.io/mealie-recipes/mealie:v3.27.0 (push rc=0)
22:21:55 [5] badge HU: [{'title': 'Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot.', 'text': 'Frissítés elérhető — ma'}]
22:21:55 [5] badge EN: [{'title': 'A newer version of this app is available. Select the Update button to start it.', 'text': 'Update available — today'}]
22:21:55 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'}
22:21:55 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None
22:21:56 + 1.1s phase=starting label=Indítás az új verzióval… err=None hold=None
22:21:58 + 3.1s phase=verifying label=Működés ellenőrzése… err=None hold=None
22:22:13 + 18.5s phase=done label=Frissítve err=None hold=None
22:22:13 [7] reading the seed back AFTER the update
22:22:15 mealie: readback of the seeded recipe http=200 ok=True
22:22:17 [8] pinned = {'mealie': 'ghcr.io/mealie-recipes/mealie:v3.27.0'}
22:22:17 [8] installed = {'mealie': 'ghcr.io/mealie-recipes/mealie:v3.27.0'}
22:22:17 [8] compose = ['image: ghcr.io/mealie-recipes/mealie:v3.27.0']
22:22:17 [8] inspect = ['mealie ghcr.io/mealie-recipes/mealie:v3.27.0 running=true restarts=0']
22:22:17 [9] verdict proven -> /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/update-night-2026-09-21/apps/mealie/verdict.json
22:22:19 [X] stop -> 200 {'ok': True, 'message': 'Stack mealie stop completed'}
22:22:24 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'mealie', 'volumes_removed': ['mealie_mealie_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az alkalm
22:22:32 [X] after remove: deployed=False leftovers='/opt/docker/stacks/mealie'
22:22:32 leftovers done
@@ -4,3 +4,68 @@
22:13:32 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'docmost', 'volumes_removed': ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage'], 'hdd_
22:13:40 [X] after remove: deployed=False leftovers='/opt/docker/stacks/docmost'
22:13:40 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
22:14:05 [1] deployed, controller state=running, pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'}
22:14:05 docmost: /api/auth/setup http=200 rc=0
22:14:06 docmost: login as the seeded user http=200 ok=True
22:14:09 [00-state-before] 2.7s
22:14:09 PG_VERSION (the harness's own postgres probe, verbatim):
22:14:09 16
22:14:09 engine version:
22:14:09 postgres (PostgreSQL) 16.15
22:14:09 datadir size:
22:14:09 49.0M /var/lib/postgresql/data
22:14:09 volume:
22:14:09 docmost_docmost_postgres_data
22:14:12 [01-stop-the-app-keep-the-engine] 3.2s
22:14:12 app stopped (the engine stays up to be dumped)
22:14:12 docmost-postgres Up 31 seconds (healthy)
22:14:12 docmost-redis Up 31 seconds (healthy)
22:14:15 [02-dump-with-16] 2.6s
22:14:15 dumping as user=docmost
22:14:15 rc=0
22:14:15 dump bytes: 132201
22:14:15 CREATE TABLE statements: 48
22:14:21 [03-fresh-17-datadir-and-restore] 6.5s
22:14:21 17 up: postgres (PostgreSQL) 17.11
22:14:21 PG_VERSION on the fresh datadir: 17
22:14:21 restore rc=0
22:14:21 ERROR lines in the restore: 2
22:14:21 ERROR: role "docmost" already exists
22:14:21 ERROR: database "docmost" already exists
22:14:21 tables restored:
22:14:21 48
22:16:26 [04-point-the-app-at-17-and-start-it] 124.8s
22:16:26 DRILL-pg17 now answers to the name docmost-postgres on docmost_docmost-internal
22:16:26 app started
22:16:26 docmost Up 2 minutes (healthy)
22:16:26 docmost-postgres Exited (0) 2 minutes ago
22:16:26 docmost-redis Up 2 minutes (healthy)
22:16:26 {"level":"info","time":"2026-09-21T20:14:00.698Z","pid":45,"hostname":"24a35fe9b998","context":"NestApplication","msg":"Listening on http://127.0.0.1:3000 / https://docs.enkisfelhom.hu"}
22:16:26 [ELIFECYCLE] Command failed.
22:16:26 $ pnpm --filter ./apps/server run start:prod
22:16:26 $ cross-env NODE_ENV=production node dist/main
22:16:26 (node:45) ExperimentalWarning: localStorage is not available because --localstorage-file was not provided.
22:16:26 (Use `node --trace-warnings ...` to show where the warning was created)
22:16:26 {"level":"info","time":"2026-09-21T20:14:31.727Z","pid":45,"hostname":"24a35fe9b998","context":"RedisModule","msg":"default: the connection was successfully established"}
22:16:26 {"level":"info","time":"2026-09-21T20:14:32.033Z","pid":45,"hostname":"24a35fe9b998","context":"DatabaseModule","msg":"Establishing database connection"}
22:16:26 {"level":"info","time":"2026-09-21T20:14:32.065Z","pid":45,"hostname":"24a35fe9b998","context":"DatabaseModule","msg":"Database connection successful"}
22:16:27 docmost: login as the seeded user http=200 ok=True
22:16:27 [05] the seed read back on PostgreSQL 17: True
22:16:29 [06-engine-state-after] 2.3s
22:16:29 the harness's own postgres probe against the CONVERTED datadir:
22:16:29 17
22:16:29 [exit=0]
22:16:29 postgres (PostgreSQL) 17.11
22:16:29 size of the 17 datadir:
22:16:29 49.1M /var/lib/postgresql/data
22:16:43 [99-put-everything-back] 13.8s
22:16:43 docmost docmost/docmost:0.96.0 Up 5 seconds (health: starting)
22:16:43 docmost-postgres postgres:16-alpine Up 10 seconds (healthy)
22:16:43 docmost-redis redis:7-alpine Up 3 minutes (healthy)
22:16:43 docmost-postgres postgres:16-alpine Up 10 seconds (healthy)
22:16:43 PG_VERSION back on the original datadir: 16
22:16:43 docmost: login as the seeded user http=404 ok=False
22:16:43 docmost: login body 404 page not found
22:16:43 [99] the seed still reads on the ORIGINAL 16 datadir after teardown: False
22:16:43 rehearsal total 155.9s -> /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/update-night-2026-09-21/bad-days/P2.3-pg-rehearsal/result.json
@@ -0,0 +1,53 @@
22:23:52 ==== Phase 5: teardown, three layers
22:23:55 -> 00-host-before.txt
22:23:55 [M] throwaway apps still deployed: ['bentopdf', 'docmost', 'gitea', 'glance', 'nextcloud', 'opengist', 'privatebin', 'uptime-kuma', 'vaultwarden', 'wishlist', 'zipline']
22:23:55 [X] stop -> 200 {'ok': True, 'message': 'Stack bentopdf stop completed'}
22:24:01 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'bentopdf', 'volumes_removed': [], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az alkalmazás nem tárolt sa
22:24:08 [X] after remove: deployed=False leftovers='/opt/docker/stacks/bentopdf'
22:24:09 [X] stop -> 200 {'ok': True, 'message': 'Stack docmost stop completed'}
22:24:15 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'docmost', 'volumes_removed': ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage'], 'hdd_
22:24:22 [X] after remove: deployed=False leftovers='/opt/docker/stacks/docmost'
22:24:22 [X] stop -> 200 {'ok': True, 'message': 'Stack gitea stop completed'}
22:24:28 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'gitea', 'volumes_removed': ['gitea_gitea_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az alkalmazá
22:24:35 [X] after remove: deployed=False leftovers='/opt/docker/stacks/gitea'
22:24:35 [X] stop -> 200 {'ok': True, 'message': 'Stack glance stop completed'}
22:24:41 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'glance', 'volumes_removed': ['glance_glance_config'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az alka
22:24:48 [X] after remove: deployed=False leftovers='/opt/docker/stacks/glance'
22:24:50 [X] stop -> 200 {'ok': True, 'message': 'Stack nextcloud stop completed'}
22:24:55 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/nextcloud tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó
22:24:55 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead
22:24:56 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'nextcloud', 'volumes_removed': ['nextcloud_nextcloud_db_data', 'nextcloud_nextcloud_html', 'nextcloud_nextcloud_redis_data'],
22:25:04 [X] after remove: deployed=False leftovers='/opt/docker/stacks/nextcloud'
22:25:04 [X] stop -> 200 {'ok': True, 'message': 'Stack opengist stop completed'}
22:25:09 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'opengist', 'volumes_removed': ['opengist_opengist_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az
22:25:17 [X] after remove: deployed=False leftovers='/opt/docker/stacks/opengist'
22:25:17 [X] stop -> 200 {'ok': True, 'message': 'Stack privatebin stop completed'}
22:25:22 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'privatebin', 'volumes_removed': ['privatebin_privatebin_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note'
22:25:30 [X] after remove: deployed=False leftovers='/opt/docker/stacks/privatebin'
22:25:30 [X] stop -> 200 {'ok': True, 'message': 'Stack uptime-kuma stop completed'}
22:25:35 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'uptime-kuma', 'volumes_removed': ['uptime-kuma_uptime_kuma_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_no
22:25:43 [X] after remove: deployed=False leftovers='/opt/docker/stacks/uptime-kuma'
22:25:43 [X] stop -> 200 {'ok': True, 'message': 'Stack vaultwarden stop completed'}
22:25:48 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'vaultwarden', 'volumes_removed': ['vaultwarden_vaultwarden_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_no
22:25:56 [X] after remove: deployed=False leftovers='/opt/docker/stacks/vaultwarden'
22:26:06 [X] stop -> 200 {'ok': True, 'message': 'Stack wishlist stop completed'}
22:26:12 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'wishlist', 'volumes_removed': ['wishlist_wishlist_data', 'wishlist_wishlist_uploads'], 'hdd_paths_removed': [], 'hdd_paths_pre
22:26:19 [X] after remove: deployed=False leftovers='/opt/docker/stacks/wishlist'
22:26:30 [X] stop -> 200 {'ok': True, 'message': 'Stack zipline stop completed'}
22:26:35 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'zipline', 'volumes_removed': ['zipline_zipline_postgres_data', 'zipline_zipline_public', 'zipline_zipline_uploads'], 'hdd_path
22:26:42 [X] after remove: deployed=False leftovers='/opt/docker/stacks/zipline'
22:26:45 -> 01-image-store-removed.txt
22:27:13 -> 02-config-restored.txt
22:27:13 [M] git.repo_url reads back as the LIVE catalog: True
22:27:20 [M] filebrowser deployed=False state=running badge=[]
22:27:20 [M] navidrome deployed=False state=running badge=[]
22:27:20 [M] traefik deployed=False state=running badge=[]
22:27:22 -> 03-containers-after.txt
22:27:22 [M] containers left: ['felhom-controller', 'filebrowser', 'navidrome', 'traefik'] only protected infra: False
22:27:25 -> 04-host-after.txt
22:27:25 [H] no harness LXC was created tonight, so none was destroyed — the PostgreSQL rehearsal ran on 9202 itself (stated in the audit)
22:27:28 -> 05-hub.txt
22:27:28 [U] floor = 0.261.0 | min_agent = 0.131.0 | host rows seen = 3 | nothing was provisioned at the hub tonight: no customer, no config, no appliance, | no binding. The only hub act of the whole night was the floor save in Phase 0.1. |
22:27:30 -> 06-gitea.txt
22:27:30 [G] live main=4463243f2e09 drill main=4463243f2e09 image lines identical: True
22:27:30 teardown written -> /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/update-night-2026-09-21/teardown/result.json
@@ -0,0 +1,9 @@
VMID Status Lock Name
9201 running demo-hp
9202 running demo-hp-scratch
--- pvesm
Name Type Status Total (KiB) Used (KiB) Available (KiB) %
felhom-pbs pbs active 0 0 0 0.00%
local dir active 40453376 32767900 5598360 81.00%
local-lvm lvmthin active 56487936 27018179 29469756 47.83%
nvme-scratch dir active 983379700 51991128 881361960 5.29%
@@ -0,0 +1,65 @@
=== registry container + volume, removed BY NAME (never a prune)
drill-registry
drill-registry-data
=== drill images, removed BY NAME
localhost:5000/drill/glance:1.0.1 -> removed
localhost:5000/drill/wishes:2.0.0 -> removed
localhost:5000/drill/wishes:2.0.1 -> removed
localhost:5000/drill/glance:1.0.0 -> removed
localhost:5000/drill/glance:1.0.2 -> removed
localhost:5000/drill/status:2.0.0 -> removed
localhost:5000/drill/status:2.0.1 -> removed
localhost:5000/drill/paste:2.0.0 -> removed
localhost:5000/drill/paste:2.0.1 -> removed
localhost:5000/drill/paste:2.0.2 -> removed
localhost:5000/drill/pdf:1.0.0 -> removed
localhost:5000/drill/pdf:2.0.0 -> removed
localhost:5000/drill/pdf:2.0.1 -> removed
localhost:5000/drill/pdf:2.0.2 -> removed
localhost:5000/drill/gist:2.0.0 -> removed
localhost:5000/drill/gist:2.0.1 -> removed
registry:2 removed
=== anything left that says drill?
(none)
(no drill containers)
=== NOTHING WAS PRUNED — this is the full image list, for the record
alpine:3.20 7.81MB
alpine:latest 8.42MB
deluan/navidrome:0.64.0 250MB
docmost/docmost:0.96.0 753MB
ghcr.io/advplyr/audiobookshelf:2.36.1 320MB
ghcr.io/alam00000/bentopdf:v2.8.6 287MB
ghcr.io/cmintey/wishlist:v0.66.0 835MB
ghcr.io/cmintey/wishlist:v0.67.0 900MB
ghcr.io/diced/zipline:4.6.1 817MB
ghcr.io/home-assistant/home-assistant:2026.7.2 2.36GB
ghcr.io/home-assistant/home-assistant:2026.9.3 2.34GB
ghcr.io/mealie-recipes/mealie:v3.20.1 1.18GB
ghcr.io/mealie-recipes/mealie:v3.27.0 1.2GB
ghcr.io/papra-hq/papra:26.6.1-rootless 647MB
ghcr.io/papra-hq/papra:26.6.2-rootless 894MB
ghcr.io/seanmorley15/adventurelog-backend:v0.12.1 1.18GB
ghcr.io/seanmorley15/adventurelog-frontend:v0.12.1 345MB
ghcr.io/tandoorrecipes/recipes:2.6.13 868MB
ghcr.io/tandoorrecipes/recipes:2.6.15 919MB
ghcr.io/thomiceli/opengist:1.13 91.7MB
ghcr.io/thomiceli/opengist:1.15 94.4MB
gitea.dooplex.hu/admin/felhom-controller:0.241.0 423MB
gitea.dooplex.hu/admin/felhom-controller:0.242.0 423MB
gitea.dooplex.hu/admin/felhom-controller:0.243.0 423MB
gitea.dooplex.hu/admin/felhom-controller:0.245.0 423MB
gitea.dooplex.hu/admin/felhom-controller:0.260.0 423MB
gitea.dooplex.hu/admin/felhom-controller:0.261.0 423MB
gitea/gitea:1.27.0 191MB
glanceapp/glance:v0.8.6 24.1MB
grafana/grafana:13.1.0 1.16GB
grafana/grafana:13.2.2 1.4GB
gtstef/filebrowser:1.3.3-stable 215MB
louislam/uptime-kuma:2.5.1 1.71GB
louislam/uptime-kuma:2.5.5 1.72GB
mariadb:11.4 327MB
mariadb:11.6 415MB
mariadb:12.3 334MB
n8nio/n8n:2.31.3 1.53GB
n8nio/n8n:2.40.5 1.04GB
nextcloud:34.0.1-apache 1.44GB
@@ -0,0 +1,18 @@
=== restoring controller.yaml from the pre-update-night copy
-rw------- 1 root root 2018 Sep 21 20:18 /var/lib/docker/volumes/felhom-controller-data/_data/controller.yaml
-rw------- 1 root root 1944 Sep 21 18:09 /var/lib/docker/volumes/felhom-controller-data/_data/controller.yaml.pre-update-night
restored
=== the git section, read back (token redacted by this script, not by the box)
git:
branch: main
repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git
sync_interval: 15m
token: "<redacted>"
username: ""
hub:
=== removing the drill catalog cache so the next sync clones the LIVE repo (R-615)
gitea.dooplex.hu/admin/felhom-controller:0.261.0 Up 25 seconds (healthy)
=== the cache's origin, READ BACK — this is the quote the brief asks for
origin https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (fetch)
origin https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (push)
4463243 Upgrade harness: four fixtures and seven real upstream edges from the update night (R-462)
@@ -0,0 +1,4 @@
felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.261.0 Up 33 seconds (healthy)
filebrowser gtstef/filebrowser:1.3.3-stable Up 7 minutes (healthy)
navidrome deluan/navidrome:0.64.0 Up 7 minutes (healthy)
traefik traefik:v3.6.7 Up 7 minutes
@@ -0,0 +1,30 @@
VMID Status Lock Name
9201 running demo-hp
9202 running demo-hp-scratch
--- pvesm
Name Type Status Total (KiB) Used (KiB) Available (KiB) %
felhom-pbs pbs active 0 0 0 0.00%
local dir active 40453376 32767980 5598280 81.00%
local-lvm lvmthin active 56487936 27018179 29469756 47.83%
nvme-scratch dir active 983379700 51059184 882293904 5.19%
--- 9201 untouched:
adventurelog
adventurelog-frontend
adventurelog-postgres
bentopdf
bookstack
bookstack-db
calibre-web
cloudflared
docmost
docmost-postgres
docmost-redis
felhom-controller
filebrowser
kimai
kimai-db
opengist
paperless-postgres
paperless-redis
paperless-webserver
privatebin
@@ -0,0 +1,5 @@
floor = 0.261.0
min_agent = 0.131.0
host rows seen = 3
nothing was provisioned at the hub tonight: no customer, no config, no appliance,
no binding. The only hub act of the whole night was the floor save in Phase 0.1.
@@ -0,0 +1,7 @@
live catalog origin/main : 4463243f2e09
drill repo HEAD after reset: 4463243f2e09
reset+force-push rc=0
diff of every `image:` line, live vs drill:
IDENTICAL — every image: line matches the live catalog
@@ -0,0 +1,52 @@
22:23:52 ==== Phase 5: teardown, three layers
22:23:55 -> 00-host-before.txt
22:23:55 [M] throwaway apps still deployed: ['bentopdf', 'docmost', 'gitea', 'glance', 'nextcloud', 'opengist', 'privatebin', 'uptime-kuma', 'vaultwarden', 'wishlist', 'zipline']
22:23:55 [X] stop -> 200 {'ok': True, 'message': 'Stack bentopdf stop completed'}
22:24:01 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'bentopdf', 'volumes_removed': [], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az alkalmazás nem tárolt sa
22:24:08 [X] after remove: deployed=False leftovers='/opt/docker/stacks/bentopdf'
22:24:09 [X] stop -> 200 {'ok': True, 'message': 'Stack docmost stop completed'}
22:24:15 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'docmost', 'volumes_removed': ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage'], 'hdd_
22:24:22 [X] after remove: deployed=False leftovers='/opt/docker/stacks/docmost'
22:24:22 [X] stop -> 200 {'ok': True, 'message': 'Stack gitea stop completed'}
22:24:28 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'gitea', 'volumes_removed': ['gitea_gitea_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az alkalmazá
22:24:35 [X] after remove: deployed=False leftovers='/opt/docker/stacks/gitea'
22:24:35 [X] stop -> 200 {'ok': True, 'message': 'Stack glance stop completed'}
22:24:41 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'glance', 'volumes_removed': ['glance_glance_config'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az alka
22:24:48 [X] after remove: deployed=False leftovers='/opt/docker/stacks/glance'
22:24:50 [X] stop -> 200 {'ok': True, 'message': 'Stack nextcloud stop completed'}
22:24:55 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/nextcloud tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó
22:24:55 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead
22:24:56 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'nextcloud', 'volumes_removed': ['nextcloud_nextcloud_db_data', 'nextcloud_nextcloud_html', 'nextcloud_nextcloud_redis_data'],
22:25:04 [X] after remove: deployed=False leftovers='/opt/docker/stacks/nextcloud'
22:25:04 [X] stop -> 200 {'ok': True, 'message': 'Stack opengist stop completed'}
22:25:09 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'opengist', 'volumes_removed': ['opengist_opengist_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az
22:25:17 [X] after remove: deployed=False leftovers='/opt/docker/stacks/opengist'
22:25:17 [X] stop -> 200 {'ok': True, 'message': 'Stack privatebin stop completed'}
22:25:22 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'privatebin', 'volumes_removed': ['privatebin_privatebin_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note'
22:25:30 [X] after remove: deployed=False leftovers='/opt/docker/stacks/privatebin'
22:25:30 [X] stop -> 200 {'ok': True, 'message': 'Stack uptime-kuma stop completed'}
22:25:35 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'uptime-kuma', 'volumes_removed': ['uptime-kuma_uptime_kuma_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_no
22:25:43 [X] after remove: deployed=False leftovers='/opt/docker/stacks/uptime-kuma'
22:25:43 [X] stop -> 200 {'ok': True, 'message': 'Stack vaultwarden stop completed'}
22:25:48 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'vaultwarden', 'volumes_removed': ['vaultwarden_vaultwarden_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_no
22:25:56 [X] after remove: deployed=False leftovers='/opt/docker/stacks/vaultwarden'
22:26:06 [X] stop -> 200 {'ok': True, 'message': 'Stack wishlist stop completed'}
22:26:12 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'wishlist', 'volumes_removed': ['wishlist_wishlist_data', 'wishlist_wishlist_uploads'], 'hdd_paths_removed': [], 'hdd_paths_pre
22:26:19 [X] after remove: deployed=False leftovers='/opt/docker/stacks/wishlist'
22:26:30 [X] stop -> 200 {'ok': True, 'message': 'Stack zipline stop completed'}
22:26:35 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'zipline', 'volumes_removed': ['zipline_zipline_postgres_data', 'zipline_zipline_public', 'zipline_zipline_uploads'], 'hdd_path
22:26:42 [X] after remove: deployed=False leftovers='/opt/docker/stacks/zipline'
22:26:45 -> 01-image-store-removed.txt
22:27:13 -> 02-config-restored.txt
22:27:13 [M] git.repo_url reads back as the LIVE catalog: True
22:27:20 [M] filebrowser deployed=False state=running badge=[]
22:27:20 [M] navidrome deployed=False state=running badge=[]
22:27:20 [M] traefik deployed=False state=running badge=[]
22:27:22 -> 03-containers-after.txt
22:27:22 [M] containers left: ['felhom-controller', 'filebrowser', 'navidrome', 'traefik'] only protected infra: False
22:27:25 -> 04-host-after.txt
22:27:25 [H] no harness LXC was created tonight, so none was destroyed — the PostgreSQL rehearsal ran on 9202 itself (stated in the audit)
22:27:28 -> 05-hub.txt
22:27:28 [U] floor = 0.261.0 | min_agent = 0.131.0 | host rows seen = 3 | nothing was provisioned at the hub tonight: no customer, no config, no appliance, | no binding. The only hub act of the whole night was the floor save in Phase 0.1. |
22:27:30 -> 06-gitea.txt
22:27:30 [G] live main=4463243f2e09 drill main=4463243f2e09 image lines identical: True
@@ -0,0 +1,43 @@
{
"host_before": "VMID Status Lock Name \n9201 running demo-hp \n9202 running demo-hp-scratch \n--- pvesm\nName Type Status Total (KiB) Used (KiB) Available (KiB) %\nfelhom-pbs pbs active 0 0 0 0.00%\nlocal dir active 40453376 32767900 5598360 81.00%\nlocal-lvm lvmthin active 56487936 27018179 29469756 47.83%\nnvme-scratch dir active 983379700 51991128 881361960 5.29%\n",
"apps_removed": {
"bentopdf": "200",
"docmost": "200",
"gitea": "200",
"glance": "200",
"nextcloud": "200",
"opengist": "200",
"privatebin": "200",
"uptime-kuma": "200",
"vaultwarden": "200",
"wishlist": "200",
"zipline": "200"
},
"repo_url_read_back": true,
"still_on_the_box": [
{
"name": "filebrowser",
"state": "running",
"deployed": false,
"badge_hu": []
},
{
"name": "navidrome",
"state": "running",
"deployed": false,
"badge_hu": []
},
{
"name": "traefik",
"state": "running",
"deployed": false,
"badge_hu": []
}
],
"only_protected_infra_left": false,
"host_after": "VMID Status Lock Name \n9201 running demo-hp \n9202 running demo-hp-scratch \n--- pvesm\nName Type Status Total (KiB) Used (KiB) Available (KiB) %\nfelhom-pbs pbs active 0 0 0 0.00%\nlocal dir active 40453376 32767980 5598280 81.00%\nlocal-lvm lvmthin active 56487936 27018179 29469756 47.83%\nnvme-scratch dir active 983379700 51059184 882293904 5.19%\n--- 9201 untouched:\nadventurelog\nadventurelog-frontend\nadventurelog-postgres\nbentopdf\nbookstack\nbookstack-db\ncalibre-web\ncloudflared\ndocmost\ndocmost-postgres\ndocmost-redis\nfelhom-controller\nfilebrowser\nkimai\nkimai-db\nopengist\npaperless-postgres\npaperless-redis\npaperless-webserver\nprivatebin\n",
"hub": "floor = 0.261.0\nmin_agent = 0.131.0\nhost rows seen = 3\nnothing was provisioned at the hub tonight: no customer, no config, no appliance,\nno binding. The only hub act of the whole night was the floor save in Phase 0.1.\n",
"live_main": "4463243f2e09",
"drill_main": "4463243f2e09",
"image_lines_identical": true
}
File diff suppressed because one or more lines are too long
+20 -21
View File
@@ -14,44 +14,43 @@ customer, hub off, tunnel off, `https://192.168.0.114` with the same Host names
## Standing on demo-hp — walked on the scratch guest 9202
- [x] bentopdf — 2026-09-13 (night 2, scratch guest 9202): install, front door 200, no data route (browser-side toolbox — recorded, not faked), backup now + Tier 2, remove-keep → Tier-2 restore back, same-version Update (backed up first: the restore's deploy time is newer than the copies, R-478 by design), remove-all clean. No new rows. `audits/nightly-2026-09-13b-bentopdf/`.
- [ ] bookstack
- [x] bookstack — 2026-09-21 (update night, scratch guest 9202): install, seeded through `php artisan` with its own negative control, backup, real upstream Update 26.05.2->26.05.5 (45.1 s), read back, removed. DATABASE HALF ONLY (R-460). `audits/DRILL-update-night-2026-09-21.md`.
- [ ] calibre-web
- [ ] docmost
- [x] docmost — 2026-09-21 (update night, scratch guest 9202): install, seeded a workspace+user through its own API, read back, backup, real upstream Update 0.95.0->0.96.0 (103.6 s) with its own migration line quoted, read back, removed with data. `audits/DRILL-update-night-2026-09-21.md`.
- [ ] kimai
- [ ] opengist
- [ ] opengist — NOT A FULL WALK; attempted 2026-09-21 (update night, scratch guest 9202): install, seeded an account through its own sign-up form, read back, backup, real upstream Update 1.13->1.15 — the UPDATE SUCCEEDED (14.4 s, all four observables agree) but the readback ran before /login was serving, so the edge is **inconclusive**, not proven. `audits/DRILL-update-night-2026-09-21.md`.
- [ ] paperless-ngx
- [ ] privatebin
- [ ] romm
- [x] privatebin — 2026-09-21 (update night, scratch guest 9202): install, seeded a paste through its own JSON API, read back, „Mentés most", real upstream Update 2.0.5->2.0.6 (15.4 s), read back again, removed with data. `audits/DRILL-update-night-2026-09-21.md`.
- [x] romm — 2026-09-21 (update night, scratch guest 9202): install, seeded through its own CSRF-protected user API, backup, real upstream Update 5.0.0->5.3.0 (60.6 s), logged in again, removed. DATABASE HALF ONLY. `audits/DRILL-update-night-2026-09-21.md`.
## The rest of the catalog
- [x] actualbudget — 2026-09-13: no non-browser front door (the client talks a sync protocol, no REST); deploy / update / remove were exercised the same day for R-475 (`audits/rulings-r472-r475-2026-09-13/`). Recorded, not faked.
- [x] adventurelog — 2026-09-13: full walk (install, sign-up + trip data through the API, backup now + Tier 2, remove-with-data-kept → second-drive restore REFUSED (R-486) → own-copy restore byte-identical, same-version guarded Update, remove-with-backups left 484 MB behind (R-474)). Photo upload needs a browser (R-483). Rows: R-482..R-487. `audits/nightly-2026-09-13-adventurelog/`.
- [ ] audiobookshelf
- [x] audiobookshelf — 2026-09-21 (update night, scratch guest 9202): install with its drive path, seeded a root account through /init, read back, backup, real upstream Update 2.35.1->2.36.1 (23.6 s), read back, removed. DATABASE HALF ONLY. `audits/DRILL-update-night-2026-09-21.md`.
- [ ] calcom
- [ ] claper
- [ ] code-server
- [ ] crafty-controller
- [ ] emby
- [ ] ghost
- [ ] gitea
- [ ] glance
- [ ] gitea — NOT A FULL WALK; attempted 2026-09-21 (update night, scratch guest 9202): attempted 2026-09-21 and **inconclusive**: a fresh Gitea sits in its web-installer state so `gitea admin user create` refuses. Needs an installer-form fixture (R-624). `audits/DRILL-update-night-2026-09-21.md`.
- [ ] glance — NOT A FULL WALK; attempted 2026-09-21 (update night, scratch guest 9202): used 2026-09-21 as the UNATTENDED HOLD subject on drill images (R-618's method), not as a catalog walk — deliberately not ticked. `audits/DRILL-update-night-2026-09-21.md`.
- [ ] gokapi
- [ ] grafana
- [x] grafana — 2026-09-21 (update night, scratch guest 9202): install, seeded a folder through its own API as the deploy-generated admin, read back, backup, real upstream Update 13.1.0->13.2.2 (26.7 s), read back, removed with data. `audits/DRILL-update-night-2026-09-21.md`.
- [ ] gramps-web
- [ ] home-assistant
- [x] home-assistant — 2026-09-21 (update night, scratch guest 9202): install, seeded the owner through its own onboarding API, logged in through its own login flow, backup, real upstream Update 2026.7.2->2026.9.3 (103.6 s), read back, removed with data. `audits/DRILL-update-night-2026-09-21.md`.
- [ ] homebox
- [ ] homepage
- [ ] immich
- [ ] jellyfin
- [ ] komga
- [ ] mealie
- [ ] n8n
- [ ] navidrome
- [ ] nextcloud
- [x] mealie — 2026-09-21 (update night, scratch guest 9202): install, seeded a recipe through its own API as the first-run admin, read back, backup, real upstream Update v3.20.1->v3.27.0 (18.5 s), read back, removed with data. `audits/DRILL-update-night-2026-09-21.md`.
- [x] n8n — 2026-09-21 (update night, scratch guest 9202): install, seeded the owner through its own setup API, read back, backup, real upstream Update 2.31.3->2.40.5 (117.9 s), read back, removed with data. `audits/DRILL-update-night-2026-09-21.md`.
- [x] navidrome — 2026-09-21 (update night, scratch guest 9202): install with its drive path, seeded the first admin through its own API, read back, backup, real upstream Update 0.63.2->0.64.0 (11.3 s), read back, removed. **The removal left a container behind that a later reboot resurrected — R-626**. `audits/DRILL-update-night-2026-09-21.md`.
- [x] nextcloud — 2026-09-21 (update night, scratch guest 9202): install, seeded a user through `occ`, read back, backup, **MariaDB engine major 11.6->12.3 through the real Update button** (217.4 s) with all four SPIKE-r459 observables, read back, removed. `audits/DRILL-update-night-2026-09-21.md`.
- [ ] onlyoffice
- [ ] outline
- [ ] papra
- [x] papra — 2026-09-21 (update night, scratch guest 9202): install, seeded an account through its own sign-up API, signed in again, backup, real upstream Update 26.6.1->26.6.2 (60.5 s), removed with data. `audits/DRILL-update-night-2026-09-21.md`.
- [ ] plant-it
- [ ] plex
- [ ] radarr
@@ -60,12 +59,12 @@ customer, hub off, tunnel off, `https://192.168.0.114` with the same Host names
- [ ] seerr
- [ ] sonarr
- [ ] sparkyfitness
- [ ] tandoor
- [ ] tandoor — NOT A FULL WALK; attempted 2026-09-21 (update night, scratch guest 9202): install, seeded a superuser through its own Django CLI, read back, backup, real upstream Update 2.6.13->2.6.15 — **the update SUCCEEDED and the app served 200 on the new version, then the box stopped it because the template probes the wrong port (R-618, P1)**. Restored in 32.4 s. **failed**, and the failure is ours. `audits/DRILL-update-night-2026-09-21.md`.
- [ ] termix
- [ ] uptime-kuma
- [ ] vaultwarden
- [ ] vikunja
- [ ] vaultwarden — NOT A FULL WALK; attempted 2026-09-21 (update night, scratch guest 9202): attempted 2026-09-21 and **inconclusive BY DESIGN**: the catalog closes self-registration on purpose (R-512), so no headless account route exists (R-624). `audits/DRILL-update-night-2026-09-21.md`.
- [x] vikunja — 2026-09-21 (update night, scratch guest 9202): install, registered a user and created a project through its own API, read back, backup, real upstream Update 2.3.0->2.6.0 (24.6 s), read back, removed with data. `audits/DRILL-update-night-2026-09-21.md`.
- [ ] wanderer
- [ ] wger
- [ ] wger — NOT A FULL WALK; attempted 2026-09-21 (update night, scratch guest 9202): deployed 2026-09-21 ONLY to measure its health probe: it listens on 8000 and the template probes 80, so the box reads it unhealthy while it serves (R-618). Removed. NOT a catalog walk. `audits/DRILL-update-night-2026-09-21.md`.
- [ ] wishlist
- [ ] zipline
- [ ] zipline — NOT A FULL WALK; attempted 2026-09-21 (update night, scratch guest 9202): attempted 2026-09-21 and **inconclusive**: registration disabled by the template; and its probe path is wrong (R-618), so an update would have been held by the probe (R-624). `audits/DRILL-update-night-2026-09-21.md`.