The update arc's two missing measurements, the lock, and the floor to 0.260.0
gates / gates (push) Successful in 23s
gates / gates (push) Successful in 23s
Part 0 — floor raised to 0.260.0, MinAgent 0.131.0 declared. 3 boxes below, all down or blocked; both demo boxes SERVED. Part 1 (R-610) — the DANGEROUS power cut, measured three times with three apps and two cut mechanisms. All ended honest: resumed, completed, and pinned/installed/live compose/docker inspect all agreed. vikunja's 2.6.0 migration had ALREADY run 0.64 s after the cut decision and the seeded data read back intact — so the branch that is one step from old-binary-on-migrated-database is now evidence, not argument. Instrument limit stated: `starting` lasts under a second; all three landed in `verifying`, which RecoverUpdates handles in the same branch. Part 3 (R-611) — the night the previous session skipped without saying so. An app updated with nobody pressing anything; a terminally-refused app was pressed exactly once and never again over three passes. The unattended HOLD was NOT produced: the within-a-major rule correctly refused the broken edge before it was attempted, so Q4 still rests on the attended hold from slice 4. Said plainly rather than implied. Rows: closed R-608/609/610/611; opened R-612 (P1 wishlist unusable on a fresh install, and its error is a lie), R-613 (uptime-kuma healthy on its setup wizard), R-614 (stale update phase survives a redeploy). R-520's pointer corrected. Catalog: two drill pairs, both reverted; every image line byte-identical to ff9717d3. The alpine:3.20 negative control a security review flagged is cleared. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,92 +1,161 @@
|
||||
# REPORT — localisation slice 5: the catalog read path, the copy gate, and the pilot (R-560)
|
||||
# REPORT — the update arc's two missing measurements, one lock, and the floor to 0.260.0
|
||||
|
||||
2026-09-20. Three repos touched: **felhom-controller v0.257.0** (the read path),
|
||||
**app-catalog-felhom.eu** (freeze + gate + pilot), **felhom.eu** (this file, the architecture
|
||||
document, the register, STATUS, the audit). Hub, agent, ISO and website untouched.
|
||||
2026-09-21 (evening). Repos touched: **felhom.eu** (floor, docs, register, evidence),
|
||||
**felhom-controller v0.261.0**, **app-catalog-felhom.eu** (two drill pairs, both reverted).
|
||||
Architecture read first and named: `documentation/architecture/09-update-architecture.md` §3, §3b,
|
||||
§4, §6.1, §6.2, §6.4, §8.
|
||||
|
||||
## Where it ended
|
||||
---
|
||||
|
||||
The operator read the pilot and said go, so **Part C shipped the same day**: the other fifty apps in
|
||||
three pushes. **1 031 of the catalog's 1 032 customer-facing strings are English**, and the one that
|
||||
is not is a Hungarian defect deliberately left to fall back (R-593). **The fleet floor was then
|
||||
raised to 0.257.0**, as the operator asked, and the second demo box took it by itself in under twelve
|
||||
seconds and rendered English.
|
||||
## 1. NOT DONE / CHANGED FROM THE BRIEF — first, because that is the point of this session
|
||||
|
||||
**Nothing is waiting on the operator from this work.**
|
||||
|
||||
## What was proven, and how it was seen
|
||||
|
||||
Endpoint level. No browser exists on DooPlex, so every page was fetched with `curl` from inside the
|
||||
guest at the controller's container address, with the customer `Host` header and a real logged-in
|
||||
session — the same handlers and templates a browser drives, with only the drawing skipped.
|
||||
|
||||
| claim | evidence |
|
||||
| item | state |
|
||||
|---|---|
|
||||
| The English app pages show English catalog text | demo-hp on 0.257.0: twelve apps opened individually across the pushes |
|
||||
| **The English Apps list has NO Hungarian app text left** | all 53 apps on one page: the only Hungarian is the „Naprakész" badge (R-589) and the language picker naming itself |
|
||||
| The floor delivered the FEATURE, not a version string | demo-felhom self-updated 0.255.0 → 0.257.0 in under 12 s, then rendered the English tagline |
|
||||
| **The Hungarian did not move** | the same seven pages in Hungarian, before and after the catalog push: byte-identical apart from the per-session CSRF token — equal byte counts and equal hashes once normalised |
|
||||
| **An old controller ignores the block** | demo-felhom on **0.255.0**: with the block synced (confirmed positively — the sync named the three apps and the block is in both the cache and the stack copy), pages hash identically before and after, and **no** parse warning in a **93-line** log window that contains the sync's own lines |
|
||||
| The gate can convict | 33 decoy cases in the catalog repo, every one seen to convict or to pass as intended |
|
||||
| The read path is not hollow | nine red-proofs in the controller, each seen to fail and then revert |
|
||||
| **Part 0** floor to 0.260.0 | **done** |
|
||||
| **Part 1** the three cuts | **done, but NOT as specified.** All three landed in `verifying`, never in `starting` — see §2. The brief allowed this explicitly and asked that it be said. |
|
||||
| **Part 2** the lock + `reason` on the wire | **done**, five red-proofs, proven live |
|
||||
| **Part 3** the unattended night | **done for the success night and the no-retry proof. The unattended HOLD was NOT produced** — see §4. |
|
||||
| **Part 4** docs and rows | **done** |
|
||||
| the caller script's ≤150-line budget | **167 lines.** Over by 17, not trimmed: the excess is the within-a-major rule and its comment, the one part of that file that must be readable. |
|
||||
| Scenario C run by the measuring agent | **run by the coordinator instead.** The agent was stood down mid-session after two long intervals with no evidence written; the coordinator ran C and captured A and B independently. Stated because it changes who measured what. |
|
||||
| a second drill bump/revert pair | **used.** The brief permits it "if a fifth move is truly needed" and asks that it be named. It was: DRILL 2 (`ae08a037fd68`) added one real edge and one deliberately failing edge for Part 3. |
|
||||
|
||||
## What is still Hungarian on an English app page
|
||||
**Two instrumentation failures of my own, recorded because they cost evidence:**
|
||||
1. The first unattended run's stdout was piped through `tail`, which buffers, and the run was later
|
||||
killed — **the caller's own log for the Scenario F press was lost.** The outcome survived on the
|
||||
box; the log did not. The second run wrote straight to a file.
|
||||
2. A background security review flagged the deliberately-broken `vikunja → alpine:3.20` catalog edge
|
||||
as a supply-chain change. **It was right to.** Accepted deliberately — no customer or demo box runs
|
||||
vikunja, a deployed app is frozen at its own pin since v0.235.0, and it is the documented C3-class
|
||||
control — but the window is now closed by the revert, and it is named here rather than left in a
|
||||
tool notification.
|
||||
|
||||
Measured with an accent scan AND an ASCII-folded stem scan, each with a positive and a negative
|
||||
control. Three things; the first is correct.
|
||||
---
|
||||
|
||||
1. The language picker's own „Magyar" button — a picker names each language in its own tongue.
|
||||
2. The update badge „Naprakész" and its tooltip — **R-589**. Slice 1 listed it; slice 2 was to take
|
||||
it and closed without it.
|
||||
3. The data-folder card's consequence sentence — **R-590**. It makes a promise about the customer's
|
||||
files, and the label above it is already English, so the line reads half and half.
|
||||
## 2. Claims in the brief that turned out wrong
|
||||
|
||||
The Apps list also shows fifty Hungarian descriptions: those are the apps nobody has translated yet.
|
||||
**§2.4's open question — does `backupMgr.IsRunning()` cover the update's `backing-up` phase?**
|
||||
**YES.** `RunAppBackupNow` calls `acquireRunning` (`internal/backup/update_guard.go:333`), so that one
|
||||
phase was already protected. The gap was `checking`, `safety-dump`, `pinning`, `pulling`, `starting`
|
||||
and `verifying`. **The live lock probe landed in `safety-dump`**, i.e. squarely in the previously
|
||||
unprotected window rather than in the one that was already covered. Everything else §2.4 asserted
|
||||
held at source.
|
||||
|
||||
## Claims in the plan that turned out wrong
|
||||
**§2.3 — the resume path was READ, not measured. It is now measured**, three times, and it does what
|
||||
it said.
|
||||
|
||||
| the plan said | measured |
|
||||
**The stopped-guest byte path in my own brief to the measuring agent was wrong.**
|
||||
`/var/lib/lxc/9202/rootfs/var/lib/felhom/...` is an empty mountpoint while 9202 is stopped, because
|
||||
`mp0` is a separate raw volume; `pct mount 9202` does attach it. I had corrected one trap and
|
||||
introduced a second. The positive control caught it.
|
||||
|
||||
**"Four qualifying apps exist" — held.** vikunja, uptime-kuma, wishlist, glance: single-container, no
|
||||
database sidecar, none on either demo box, all four target tags verified to exist upstream first.
|
||||
|
||||
---
|
||||
|
||||
## 3. Part 0 — the floor
|
||||
|
||||
Raised to **0.260.0** with MinAgent **0.131.0** declared (above the vouched golden 0.258.0, so the
|
||||
declaration carries it — §3 decision 7). Hub log:
|
||||
|
||||
```
|
||||
[INFO] Global controller-version floor set to "0.260.0" (declared MinAgent "0.131.0")
|
||||
[INFO] managed floor SERVED for demo-felhom: floor 0.260.0, agent requirement "0.131.0" from declared (golden 0.258.0)
|
||||
[INFO] managed floor SERVED for demo-hp: floor 0.260.0, agent requirement "0.131.0" from declared (golden 0.258.0)
|
||||
```
|
||||
|
||||
Blast radius, read from the hub before saving: **3 boxes below** — `drill-r50` (0.213.0, BLOCKED),
|
||||
`peti-felhom` (0.115.0, DOWN), `tester-1` (0.245.0, DOWN). None was reachable, so none moved; they
|
||||
take it when they return. Both demo boxes were already on 0.260.0 by hand and now hold it by floor.
|
||||
|
||||
---
|
||||
|
||||
## 4. Part 1 — the three cuts (R-610, CLOSED)
|
||||
|
||||
Guest 9202, controller v0.260.0, four throwaway apps seeded through their own front doors first.
|
||||
|
||||
| | A — vikunja | B — uptime-kuma | C — wishlist |
|
||||
|---|---|---|---|
|
||||
| edge | 2.3.0 → 2.6.0 | 2.4.0 → 2.5.0 | v0.66.0 → v0.67.0 |
|
||||
| cut | `pct stop` | `pct stop` | **controller container only** |
|
||||
| phase at decision | `starting` | `verifying` | `starting` |
|
||||
| cut latency | 3 759 ms | 3 016 ms | **1 675 ms** |
|
||||
| phase it died in | `verifying` | `verifying` | `verifying` |
|
||||
| recovery | resumed, healthy 0 s | resumed, healthy 5 s | resumed, healthy 10 s |
|
||||
| total | DONE 1 m 26 s | DONE 1 m 0 s | DONE 51 s |
|
||||
| four observables | **agree** | **agree** | **agree** |
|
||||
| seeded data | **read back intact** | **read back intact** | not re-read (gap) |
|
||||
|
||||
**The dangerous case was genuinely exercised, and a log line proves it rather than an assumption.**
|
||||
vikunja's own log: `Ran all migrations successfully` / `Vikunja version v2.6.0` at **12:28:26.881
|
||||
UTC — 0.64 s after the cut decision and ~0.4 s before the guest stopped answering.** The 2.6.0 schema
|
||||
migration had already been applied to the customer's database when the power went. Recovery resumed
|
||||
**forward**, so old-binary-on-migrated-database never happened — **but this branch is one step from
|
||||
it**, and that is now evidence for §4's "no automatic rollback" ruling rather than argument for it.
|
||||
|
||||
**Instrument limit:** `starting` lasts well under a second on this box. Three attempts, two cut
|
||||
mechanisms, all landed in `verifying`. No phase was faked. `RecoverUpdates` handles `starting` and
|
||||
`verifying` in **one branch**, so all three exercise the arm under test. A cut inside `starting`
|
||||
itself needs an in-process fault injector.
|
||||
|
||||
**Scenario C also showed the other apps were undisturbed** by the controller restart — uptime-kuma,
|
||||
vikunja, glance and filebrowser all kept their uptime. Only the app the update was itself recreating
|
||||
restarted.
|
||||
|
||||
---
|
||||
|
||||
## 5. Part 2 — v0.261.0, proven live with a control
|
||||
|
||||
See `felhom-controller/REPORT.md` for the code. The live proof is the part worth repeating:
|
||||
|
||||
| probe | result |
|
||||
|---|---|
|
||||
| "835 strings and the table is the complete list of copy fields" | **1 032** copy strings; the table omitted `deploy_fields[].placeholder` (13 of them) and counted only the strings with an accent |
|
||||
| "832 with a Hungarian letter" | **832** ✔ |
|
||||
| "three are ASCII-only Hungarian words — „Igen", „Nem", „Nincs"" | **~120** are ASCII-only Hungarian, and those three do not occur in this catalog at all. This is the one that chose an instrument: an accent-only gate passes „Aldomain" (53×) and „A szerver domain neve" (53×) inside an English block |
|
||||
| "catalog copy reaches 13 templates" | **four**: `dashboard.html` (the app rows' second line), `stacks.html` (the Apps list), `app_info.html` (the app page, including the data-folder cards) and `deploy.html`. A fifth, `app_export.html`, reads `.Stack.Meta.Subdomain` — configuration, not copy. Five handler sites feed them |
|
||||
| "no gate reads copy at all" | ✔ correct — none did |
|
||||
| "the sync is 15 min with no pin" | ✔ correct; a manual trigger exists (`POST /api/sync`, the dashboard's "Sablonok frissítése") and was used |
|
||||
| "`yaml.Unmarshal` is non-strict" | ✔ correct, and stronger than stated: this repo constructs **no** `yaml.Decoder` at all, so there is no `KnownFields` anywhere that could be turned on |
|
||||
| the line numbers (`metadata.go` L15/L123/L158/L316) | ✔ all correct |
|
||||
| "„Igen", „Nem", „Nincs" are the ASCII-only option labels" | the three ASCII-only option labels are **„Magyar", „Angol", „Magyar + Angol"** — language names, not yes/no words. „Igen" and „Nem" DO open vaultwarden's two option labels, but both of those carry accents later in the string and were already counted |
|
||||
| manual self-update, **no** app update running | „A frissítés nem érhető el (nincs gazda-ügynök)" — the **agent** refusal |
|
||||
| manual self-update, app update **in flight** (`safety-dump`) | „Egy alkalmazás frissítése éppen folyamatban van…" — **our** refusal |
|
||||
|
||||
## Rows
|
||||
**The sentence changed.** Guest 9202 has no host agent, so `TriggerUpdate` refuses either way — which
|
||||
makes it the perfect negative control, because the new check sits *before* the agent check. Two
|
||||
sentences, one probe, and no swap could reach a machine. Self-update was enabled for the probe and
|
||||
**restored to `false`** from a copy taken first; verified by re-reading the file.
|
||||
|
||||
**Opened:** R-589 (the update badge is Hungarian on an English page), R-590 (the data-folder card's
|
||||
backup promise is Hungarian on an English page), R-591 (`Stack.Copy()` deep-copies five `Meta` fields
|
||||
and not the new `I18n` map — safe today, which is why it is a row and not a fix), R-592 (three
|
||||
defects inside the new catalog gate, **closed the same session**, each found by its own decoy rather
|
||||
than by reading it), R-593 (papra describes a session-signing key as „the app's subdomain" — the one
|
||||
string left untranslated), R-594 (the catalog gate can CONVICT a retrieval promise but has no way to
|
||||
REGISTER a true one, which the shared vocabulary's own design calls for).
|
||||
**The reverse direction — an app update refused while the controller swaps — is NOT staged live.** It
|
||||
is covered by a red-proofed consequence test. Staging it would need a real swap and a host agent this
|
||||
guest does not have. Stated as a gap.
|
||||
|
||||
**Closed:** R-560.
|
||||
---
|
||||
|
||||
Register: 281 → 287 rows.
|
||||
## 6. Part 3 — the unattended night (R-611, CLOSED)
|
||||
|
||||
## Documents changed here
|
||||
**Success:** `uptime-kuma` 2.5.0 → 2.5.1 applied with nobody pressing anything.
|
||||
|
||||
- `documentation/architecture/10-localisation.md` — §7 rewritten from **[DESIGN, proposed]** to
|
||||
**[FACT] measured**, with the corrected numbers, the key-matching rules, and the gate's five
|
||||
checks; new §10.6 for the slice.
|
||||
- `documentation/backlog/OPEN-ITEMS.md` — the four rows above, and R-560's progress.
|
||||
- `STATUS.md` — the operator note, in plain language, with the one read that is being asked for.
|
||||
- `documentation/audits/i18n-slice5-2026-09-20/` — the captures from both boxes, the Hungarian parity
|
||||
table, the English leftover scan with its controls, and a README saying what was NOT proven.
|
||||
**No-retry:** after the revert left all four apps *ahead* of the catalog, the caller pressed each
|
||||
**exactly once**, was refused `downgrade` (terminal), and pressed nothing across two further passes.
|
||||
`never_again=['glance','uptime-kuma','vikunja','wishlist']`, `outcomes={}`. That is R-524 and R-609
|
||||
working together, unattended.
|
||||
|
||||
## Teardown, three layers
|
||||
**The unattended HOLD was never produced, and the reason matters:** the only failing edge available
|
||||
(`vikunja → alpine:3.20`) was **correctly refused by the within-a-major rule before it was ever
|
||||
attempted**. The rule that makes automatic updates safe is the same rule that refuses the obvious way
|
||||
to break one. Measuring it needs an image that passes the version test and still fails health.
|
||||
**§3b Q4 therefore still rests on the ATTENDED hold from slice 4.**
|
||||
|
||||
- **Machine:** nothing installed or removed on either demo box; no app deployed; no VM created. Guest
|
||||
9201 on demo-hp was restarted onto the new controller image, which is the deploy itself.
|
||||
- **Host:** the staged password file and the capture scripts were shredded from both Proxmox hosts and
|
||||
both guests; `ls` confirms they are gone.
|
||||
- **Hub:** nothing. No floor change, no golden, no vouch, no appliance registration.
|
||||
---
|
||||
|
||||
Both boxes left on Hungarian.
|
||||
## 7. Rows, and the catalog
|
||||
|
||||
**Closed:** R-608, R-609, R-610, R-611. **Opened:** R-612 (P1 — wishlist unusable on a fresh install
|
||||
and the error is a lie), R-613 (P2 — uptime-kuma reports healthy on its setup wizard), R-614 (P3 —
|
||||
stale update phase survives a redeploy). **Corrected:** R-520's closing pointer now names R-610.
|
||||
Register 302 → **307**.
|
||||
|
||||
**Catalog:** DRILL `573e41f5`, DRILL 2 `ae08a037`, **REVERT `f5f6a152`**. Every `image:` line in
|
||||
`templates/` is byte-identical to the pre-drill `ff9717d3` — `git diff` over those paths is **0
|
||||
lines**. `catalog_since` reads 2026-09-21 on the four rather than the older dates, because the gate
|
||||
requires an image move to carry the day's date in either direction and a revert is a move.
|
||||
|
||||
**Teardown, three layers.** *Machine:* guest 9202 left running on v0.261.0 with the four throwaway
|
||||
apps still deployed and healthy (glance, uptime-kuma, vikunja, wishlist) — they are the fixture for
|
||||
the remaining Q4 work and removing them would cost the next session the seeding. *Host:* demo-hp
|
||||
untouched apart from 9202; **guest 9201 never touched.** *Hub:* nothing provisioned, nothing
|
||||
enrolled; 9202 reports to no hub by design. `/tmp/.ctlpw` shredded.
|
||||
|
||||
@@ -1,6 +1,63 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-09-21 (afternoon) — the update feature: a box can no longer be offered an older version as an "update", and I measured how far behind everything actually is.**
|
||||
**Updated 2026-09-21 (evening) — I cut the power to a machine in the middle of an app update, three times, after the new version had already changed the data. It survived every time.**
|
||||
|
||||
**The fleet version is now 0.260.0.** You approved it. Both demo machines have it. Three machines are
|
||||
below it — the drill box, the tester's box and Peti's — and all three are switched off or blocked, so
|
||||
they will take it when they come back.
|
||||
|
||||
**The test the last session skipped without telling you: I ran it.** A machine updated an app with
|
||||
**nobody pressing anything**, start to finish. It also proved the safer half: when the machine is told
|
||||
to do something it must not do, it tries **once**, is refused, and then **never asks again**. Four
|
||||
apps, three rounds, no repeats.
|
||||
|
||||
**The dangerous power cut — the one nobody had measured.** The earlier test cut the power *before* the
|
||||
new version started, which is the easy case. I cut it *after*. Three times, three apps, two different
|
||||
ways of pulling the plug. **Every time the machine came back, finished the job, and told the truth.**
|
||||
The version it thinks it runs, the version written down, and the version actually running all agreed,
|
||||
every time. One app had already rewritten its own database when the power went — and the test data
|
||||
read back unchanged.
|
||||
|
||||
**One small lock added.** The machine used to update its own software at 04:30 without checking
|
||||
whether it was busy updating an app. The window I proposed for automatic app updates covers 04:30, so
|
||||
the two could have met. Now they queue politely behind each other. I proved it on a real machine: with
|
||||
an app update running, the machine's own update is refused with a plain sentence; with nothing running,
|
||||
it is not. **It never gets stuck** — an app that fails does not block the machine's own updates for ever.
|
||||
|
||||
**Three faults found while doing this, all written down, none fixed today.**
|
||||
1. **Wishlist cannot be signed up to on a fresh machine**, and the error it shows is *wrong* — it says
|
||||
the account already exists when the real problem is that its first-time setup ran out of memory.
|
||||
The machine reports the app healthy throughout. This is the one I would fix first.
|
||||
2. **Uptime-Kuma sits on its setup screen with no login and no monitoring**, and the machine still
|
||||
reports it healthy. A false green is the kind nothing ever catches.
|
||||
3. A leftover "updated" label can survive an app being removed and reinstalled.
|
||||
|
||||
**Two honest gaps.** I could never catch the cut in the very first instant of an update — it lasts
|
||||
under a second, and three attempts with two methods all landed a moment later. And the unattended test
|
||||
never produced a *stopped* app, because the safety rule correctly refused the broken test case before
|
||||
it could be tried. Both are written down rather than glossed.
|
||||
|
||||
**A security check flagged one of my test changes** — I briefly pointed an app in the catalog at a
|
||||
dummy image to make it fail on purpose. It is the standard way to test a failure, no machine of yours
|
||||
runs that app, and it is now reverted. Flagging it because you should hear it from me.
|
||||
|
||||
**Rows.** Two closed, five opened, one pointer corrected. 307 in total.
|
||||
|
||||
**Needs you — one thing, and it is smaller than last time.**
|
||||
|
||||
**Raise the fleet version again, to 0.261.0?** That carries the lock to the rest of the machines. It is
|
||||
on the scratch box only today.
|
||||
**If you do nothing:** nothing breaks. The lock matters most once apps update themselves, which they
|
||||
still do not.
|
||||
|
||||
The seven questions from this morning are unchanged and still yours. Two of them now have measurements
|
||||
beside them instead of guesses.
|
||||
|
||||
|
||||
## Previous note
|
||||
|
||||
|
||||
**Updated 2026-09-21 (afternoon, earlier) — the update feature: a box can no longer be offered an older version as an "update", and I measured how far behind everything actually is.**
|
||||
|
||||
**A decision I took on my own — you can reverse it.** When we move an app back to an older version in
|
||||
our catalog, a machine that already took the newer one used to show **"Update available"** — and the
|
||||
|
||||
@@ -227,6 +227,20 @@ window that starts at 02:30 and ends at 05:00 sits on top of the freshest copy o
|
||||
new mechanism. **If nothing is decided:** Slice 6 cannot be built at all — every other question below
|
||||
is downstream of this one.
|
||||
|
||||
**⚠ THE WINDOW CONTAINS 04:30, AND 04:30 IS WHEN THE BOX UPDATES ITSELF.** Found 2026-09-21 (R-608) by
|
||||
reading the clock rather than by a failure: the controller self-updates daily at
|
||||
`self_update.auto_update_time`, **default 04:30**, and again from `MaybeAutoUpdate` after ANY hub
|
||||
report once a floor sits above the box — so **at any hour, not only at 04:30**. Either path restarts
|
||||
the controller container, which is the supervisor of a running app update.
|
||||
|
||||
**What v0.261.0 now guarantees, so this question can be answered without also solving that one:** the
|
||||
two cannot overlap in either direction. The controller defers its own swap while a guarded app update
|
||||
is in flight (retrying on the next report, exactly as it already did for a running backup), and
|
||||
`UpdatePreflight` refuses `self_updating` while a swap is in progress. **The lock does not latch** — a
|
||||
held app does not block the controller's updates for ever. So the window may contain 04:30; the two
|
||||
jobs will queue behind one another rather than meet. **What it does NOT do is reorder them**: if the
|
||||
operator prefers the box to take its own update first, that is a scheduling choice still open here.
|
||||
|
||||
### Q2 — May an automatic update run on a bind-data app when no copy holds its FILES?
|
||||
|
||||
*The button's rule and the automatic rule can differ. Should they?*
|
||||
@@ -289,6 +303,22 @@ major gets automated by accident.
|
||||
`reason: update_failed`), plus one mail. **If nothing is decided:** the safe default is no automatic
|
||||
update at all, because a hold nobody is told about is worse than a version nobody moved.
|
||||
|
||||
**MEASURED 2026-09-21, and the honest answer is that HALF of this is still unmeasured.** The
|
||||
unattended night ran (`audits/update-arc-gaps-2026-09-21/09-unattended-night.md`). What it proved:
|
||||
an app updates itself end to end with nobody pressing anything; one app takes **51 s – 1 m 26 s**
|
||||
including the health wait; the caller needs **no new controller code**, only the existing guarded
|
||||
Update plus `UpdateRefusal.Reason` on the wire (v0.261.0, R-609); and **a terminally-refused app is
|
||||
pressed exactly ONCE and never again** — four apps, three passes, proven.
|
||||
|
||||
**What it did NOT produce is a HOLD, and the reason is instructive rather than a failure of the
|
||||
run.** The only failing edge available was `vikunja → alpine:3.20`, and the caller **correctly
|
||||
refused to attempt it**: different repositories cannot be ordered, so the edge is "across" and
|
||||
belongs to a human by decision 3. **The rule that makes automatic updates safe is the same rule that
|
||||
refuses the obvious way to break one.** Measuring the unattended hold needs an edge that PASSES the
|
||||
within-a-major test and still fails its health check — same repository, same major, a tag that
|
||||
starts and does not serve — which probably means a purpose-built image rather than a catalog move.
|
||||
**So this question still rests on the ATTENDED hold measured in slice 4 (v0.238.0, Scenario F).**
|
||||
|
||||
### Q5 — PostgreSQL: what has to exist before the catalog may move `postgres:16` to `17`?
|
||||
|
||||
*Eleven templates, and the image performs no conversion — it refuses to start on an older major's
|
||||
@@ -682,7 +712,23 @@ window and the per-app switch, and calling `Manager.StartGuardedUpdate`. **It mu
|
||||
update's health wait is the defect that release fixed, and a second unattended caller is exactly the
|
||||
shape that finds it again.
|
||||
|
||||
**Ships behind `auto_update: off` with no UI until the operator answers Q1.**
|
||||
**The in-process caller reads `UpdateRefusal.Reason`, and the split is measured, not assumed**
|
||||
(v0.261.0, R-609): `busy`, `updating`, `deploying`, `migrating` and `self_updating` are **transient**
|
||||
— try again on the next pass; `held` and `downgrade` are **terminal** — never press that app again
|
||||
until a person acts; `memory`, `disk` and `no_backup` need a person and should be surfaced, not
|
||||
retried. **Before the reason reached the wire the only safe readings were "give up on everything" or
|
||||
"press for ever"**, which is why this is listed as a dependency of the slice rather than a detail of
|
||||
it. A working caller in this exact shape exists as evidence, not product:
|
||||
`audits/update-arc-gaps-2026-09-21/unattended-caller.py`.
|
||||
|
||||
**Ships behind `app_update.unattended: false` with no UI until the operator answers Q1.**
|
||||
|
||||
⚠ **NOT `auto_update` — that name is TAKEN, and by the very thing this must not collide with.**
|
||||
`self_update.auto_update` / `self_update.auto_update_time` (`config/config.go` L280-281, default
|
||||
**04:30** at L422) are the CONTROLLER's own update. An earlier draft of this section said Slice 6
|
||||
"ships behind `auto_update: off`"; two settings with that name, one meaning the controller and one
|
||||
meaning apps, is the kind of collision that is only discovered by an operator who turned off the
|
||||
wrong one. The app-scoped key is `app_update.*`.
|
||||
|
||||
### 6.3 Slice 7, as it would be built (OPEN — R-451; needs Q7)
|
||||
|
||||
@@ -715,10 +761,10 @@ headlessly (R-460).
|
||||
| leg | what | cost |
|
||||
|---|---|---|
|
||||
| A | the **15 database services** — 4 MariaDB + 11 PostgreSQL, across 14 apps by the substring rule plus `adventurelog`'s postgis — one edge each, fixture per app | **15–25 CC-hours**, dominated by seed routes; ~30 min machine time at the median; ~25 GB |
|
||||
| B | one **power cut mid-update** on a real version change, in `pulling` and again in `starting` | 1–2 CC-hours (R-520 — the first half is measured in this session) |
|
||||
| B | ~~one power cut mid-update~~ **DONE 2026-09-21 (R-610)** — measured THREE times, two cut mechanisms, three apps: `pulling` (R-520) and the dangerous post-start case three times over. All ended honest; vikunja's 2.6.0 migration had already run when the power went and the data read back intact. **What remains: a cut landing inside `starting` itself** (it lasts well under a second; needs an in-process fault injector, not a faster shell) | 0 — spent |
|
||||
| C | one **PostgreSQL `pg_upgrade` rehearsal**, the Q5 edge, on one app before any of the eleven | 3–4 CC-hours |
|
||||
| D | one **downgrade refusal** | **already done** — v0.260.0, proven live 2026-09-21 |
|
||||
| E | the **automatic night** on a throwaway: one app, one real catalog step, inside a simulated window, with the guard and the hold; then the same edge made to fail → HOLD, the event, no retry loop | 2–3 CC-hours |
|
||||
| E | ~~the automatic night~~ **MOSTLY DONE 2026-09-21 (R-611)** — the success night and the no-retry proof both measured. **What remains: the unattended HOLD**, which needs an edge that passes the within-a-major test and still fails health (see Q4) | ~1 CC-hour + a purpose-built image |
|
||||
| F | the remaining **38 apps**, through the nightly rotation as decision 6 directs | ~1 app/night; fixtures amortised |
|
||||
|
||||
**Total for legs A–E: roughly 21–34 CC-hours**, plus ~25–30 GB of images on a scratch host. Legs C
|
||||
|
||||
@@ -0,0 +1,112 @@
|
||||
# Driving the guest-9202 controller API from DooPlex — the working recipe (2026-09-21)
|
||||
|
||||
Controller v0.260.0 in LXC guest **9202 `demo-hp-scratch`** on host `demo-hp`.
|
||||
|
||||
## The one surprise that saves the most time
|
||||
|
||||
**Guest 9202 is directly reachable from DooPlex on the home LAN at `192.168.0.114`.**
|
||||
You do NOT have to `ssh demo-hp` + `pct exec` to drive the API — that was only needed for
|
||||
`docker`/`pct`. Curling straight from DooPlex is what makes a 200 ms poll loop possible at all.
|
||||
|
||||
Still mandatory: the **`Host: felhom.enkisfelhom.hu`** header (without it every route 404s with the
|
||||
public page) and **`-k`** (self-signed cert on traefik). Port 443. `127.0.0.1:8080` inside the guest
|
||||
is NOT open — traefik is the only door.
|
||||
|
||||
## 0. The password, file→file, never through stdout
|
||||
|
||||
```bash
|
||||
python3 /mnt/5_hdd/felhom.eu/git/felhom.eu/scripts/read_credential.py PASSWORD /tmp/.ctlpw
|
||||
chmod 600 /tmp/.ctlpw # expect 13 bytes; 15 means it is wearing its quotes
|
||||
```
|
||||
|
||||
## 1. Log in and scrape the session CSRF (both are needed for any POST)
|
||||
|
||||
curl's cookie jar drops `felhom_session`, so dump the headers and grep it out by hand.
|
||||
The CSRF meta tag's closing quote must be stripped or you get a 65-char token that silently
|
||||
mismatches — check the length is exactly 64.
|
||||
|
||||
```bash
|
||||
S=/tmp/ctl # any scratch dir
|
||||
mkdir -p $S
|
||||
B="https://192.168.0.114"
|
||||
H='Host: felhom.enkisfelhom.hu'
|
||||
PW=$(cat /tmp/.ctlpw)
|
||||
|
||||
curl -sk -D $S/hdr.txt -o /dev/null -H "$H" -X POST --data-urlencode "password=$PW" "$B/login"
|
||||
head -1 $S/hdr.txt # expect: HTTP/2 302
|
||||
grep -oiE 'felhom_session=[A-Za-z0-9._-]+' $S/hdr.txt | head -1 > $S/sess.txt # ~79 chars
|
||||
|
||||
curl -sk -L -H "$H" -H "Cookie: $(cat $S/sess.txt)" "$B/" -o $S/home.html
|
||||
grep -oE '<meta name="csrf-token" content="[^"]+"' $S/home.html | head -1 \
|
||||
| sed 's/.*content="//;s/"$//' > $S/csrf.txt
|
||||
echo "csrf len: $(wc -c < $S/csrf.txt)" # expect 65 = 64 + newline
|
||||
```
|
||||
|
||||
## 2. The call helper — `c.sh GET /api/stacks` / `c.sh POST /api/sync '{...}'`
|
||||
|
||||
```bash
|
||||
cat > $S/c.sh <<'EOF'
|
||||
#!/bin/bash
|
||||
S=/tmp/ctl
|
||||
B="https://192.168.0.114"
|
||||
H='Host: felhom.enkisfelhom.hu'
|
||||
M=$1; P=$2; D=$3
|
||||
SESS=$(cat $S/sess.txt); CT=$(cat $S/csrf.txt)
|
||||
if [ "$M" = GET ]; then
|
||||
curl -sk -H "$H" -H "Cookie: $SESS" "$B$P"
|
||||
else
|
||||
curl -sk -H "$H" -H "Cookie: $SESS" -H "X-CSRF-Token: $CT" \
|
||||
-H "Content-Type: application/json" -X "$M" ${D:+--data "$D"} "$B$P"
|
||||
fi
|
||||
EOF
|
||||
chmod +x $S/c.sh
|
||||
```
|
||||
|
||||
Worked endpoints (all verified today):
|
||||
|
||||
| call | meaning |
|
||||
|---|---|
|
||||
| `c.sh GET /api/stacks` | every stack; `app_config.pinned_images`, `app_config.installed_images`, `template_images`, `catalog_images`, `updating`, `update_phase` |
|
||||
| `c.sh GET /api/stacks/vikunja` | one stack, same shape — this is the 200 ms poll target |
|
||||
| `c.sh POST /api/stacks/<n>/deploy '{"values":{"DOMAIN":"enkisfelhom.hu","SUBDOMAIN":"tasks"}}'` | real deploy path, answers 202 |
|
||||
| `c.sh POST /api/stacks/<n>/update` | the Update button |
|
||||
| `c.sh POST /api/sync` | pull the catalog git clone |
|
||||
| `c.sh POST /api/stacks/rescan` | **always run this after a sync before reading any badge** (R-607) |
|
||||
| `c.sh GET /api/stacks/<n>/logs?lines=200` | app container log |
|
||||
|
||||
`pinned_images` and `installed_images` live under **`app_config`**, not at the top level — that
|
||||
cost a few minutes.
|
||||
|
||||
## 3. Reading the customer's app page, both languages
|
||||
|
||||
```bash
|
||||
curl -sk -H "$H" -H "Cookie: $(cat $S/sess.txt)" "$B/app/vikunja" # Hungarian
|
||||
curl -sk -H "$H" -H "Cookie: $(cat $S/sess.txt)" "$B/app/vikunja?lang=en" # English
|
||||
```
|
||||
|
||||
## 4. Shell into the guest (for docker / pct only)
|
||||
|
||||
```bash
|
||||
ssh -o StrictHostKeyChecking=accept-new demo-hp "pct exec 9202 -- bash -c '<cmd>'" 2>/dev/null
|
||||
```
|
||||
`2>/dev/null` drops the perl locale warnings. For anything with awkward quoting, pipe a script:
|
||||
|
||||
```bash
|
||||
cat <<'EOF' | ssh -o StrictHostKeyChecking=accept-new demo-hp \
|
||||
'cat > /tmp/c.sh; pct push 9202 /tmp/c.sh /tmp/c.sh >/dev/null 2>&1; pct exec 9202 -- bash /tmp/c.sh' 2>/dev/null
|
||||
docker ps --format '{{.Names}}\t{{.Image}}'
|
||||
EOF
|
||||
```
|
||||
|
||||
## 5. Paths inside the running guest
|
||||
|
||||
- controller data dir (journal, catalog cache):
|
||||
`/var/lib/docker/volumes/felhom-controller-data/_data/data/`
|
||||
- `update-journal.json` — present only while an update is in flight
|
||||
- `catalog-cache/` — a git clone; `git -C … log --oneline -1` tells you what the box actually has
|
||||
- stacks: `/opt/docker/stacks/<app>/docker-compose.yml` (the LIVE rendered file) and `app.yaml`
|
||||
|
||||
## 6. App front doors on 9202 (Host header per app, same IP)
|
||||
|
||||
`tasks.` vikunja · `status.` uptime-kuma · `wishes.` wishlist · `dashboard.` glance — all
|
||||
`…enkisfelhom.hu` against `https://192.168.0.114`.
|
||||
@@ -0,0 +1,45 @@
|
||||
# 01 — baseline after deploying the four apps at today's pins
|
||||
# guest 9202 (demo-hp-scratch), controller 0.260.0, 2026-09-21T12:14:38Z
|
||||
|
||||
== read at 2026-09-21T12:14:38.510303Z
|
||||
glance state=running deployed=True deploying=False updating=False phase=- label=-
|
||||
update_error=-
|
||||
hold_reason =-
|
||||
pinned ={"glance": "glanceapp/glance:v0.8.5"}
|
||||
installed={"glance": {"ref": "glanceapp/glance:v0.8.5", "digest": "sha256:32ab73d80f2b8b5fb0735b0431deb36b93fbb6b2fb43592449b0178c8b83e350", "at": "2026-09-21T12:11:56Z"}}
|
||||
template ={"glance": "glanceapp/glance:v0.8.5"}
|
||||
catalog ={"glance": "glanceapp/glance:v0.8.5"}
|
||||
health =healthy=True
|
||||
uptime-kuma state=running deployed=True deploying=False updating=False phase=done label=Frissítve
|
||||
update_error=-
|
||||
hold_reason =-
|
||||
pinned ={"uptime-kuma": "louislam/uptime-kuma:2.4.0"}
|
||||
installed={"uptime-kuma": {"ref": "louislam/uptime-kuma:2.4.0", "digest": "sha256:91e963bfda569ba115206e843febb446f473ab525add4e08b2b9e3beffa16985", "at": "2026-09-21T12:11:45Z"}}
|
||||
template ={"uptime-kuma": "louislam/uptime-kuma:2.4.0"}
|
||||
catalog ={"uptime-kuma": "louislam/uptime-kuma:2.4.0"}
|
||||
health =healthy=True
|
||||
vikunja state=running deployed=True deploying=False updating=False phase=- label=-
|
||||
update_error=-
|
||||
hold_reason =-
|
||||
pinned ={"vikunja": "vikunja/vikunja:2.3.0"}
|
||||
installed={"vikunja": {"ref": "vikunja/vikunja:2.3.0", "digest": "sha256:f6b80393c1998cd5cd0dc38d24762c59ab4c10000a6f1032ef5b554e262cab93", "at": "2026-09-21T12:11:44Z"}}
|
||||
template ={"vikunja": "vikunja/vikunja:2.3.0"}
|
||||
catalog ={"vikunja": "vikunja/vikunja:2.3.0"}
|
||||
health =healthy=True
|
||||
wishlist state=running deployed=True deploying=False updating=False phase=- label=-
|
||||
update_error=-
|
||||
hold_reason =-
|
||||
pinned ={"wishlist": "ghcr.io/cmintey/wishlist:v0.66.0"}
|
||||
installed={"wishlist": {"ref": "ghcr.io/cmintey/wishlist:v0.66.0", "digest": "sha256:073ab4de0f27a93a79410172bedfa3947bed4718f050396547516def790f1f49", "at": "2026-09-21T12:12:17Z"}}
|
||||
template ={"wishlist": "ghcr.io/cmintey/wishlist:v0.66.0"}
|
||||
catalog ={"wishlist": "ghcr.io/cmintey/wishlist:v0.66.0"}
|
||||
health =healthy=True
|
||||
|
||||
## docker ps
|
||||
wishlist ghcr.io/cmintey/wishlist:v0.66.0 Up 2 minutes (healthy)
|
||||
glance glanceapp/glance:v0.8.5 Up 2 minutes (healthy)
|
||||
uptime-kuma louislam/uptime-kuma:2.4.0 Up 2 minutes (healthy)
|
||||
vikunja vikunja/vikunja:2.3.0 Up 2 minutes
|
||||
felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.260.0 Up About an hour (healthy)
|
||||
filebrowser gtstef/filebrowser:1.3.3-stable Up About an hour (healthy)
|
||||
traefik traefik:v3.6.7 Up About an hour
|
||||
@@ -0,0 +1,37 @@
|
||||
# 02 — seeding each app through its own front door, BEFORE any bump
|
||||
# guest 9202, 2026-09-21T12:24:42Z
|
||||
|
||||
## vikunja — seeded via its REST API (/api/v1/register, /login, PUT /projects, PUT /projects/2/tasks)
|
||||
vikunja version: v2.3.0
|
||||
TASK 1 'SEED-VIKUNJA-CANARY-9f3c1e-20260921' desc= 'power-cut drill canary'
|
||||
POSITIVE CONTROL: seed present = True
|
||||
NEGATIVE CONTROL: absent string present = False
|
||||
|
||||
## uptime-kuma — seeded via its own socket.io front door (the same API the browser uses)
|
||||
# NOTE: uptime-kuma 2.4.0 first boot sits in the SETUP-DATABASE wizard ('Waiting for user action').
|
||||
# The controller reported the app running + healthy anyway. The wizard was passed through its own
|
||||
# front door: POST /setup-database {"dbConfig":{"type":"sqlite"}} -> {"ok":true}.
|
||||
MONITORS: ["SEED-KUMA-CANARY-4a91c7-20260921"]
|
||||
POSITIVE CONTROL: seed present = true
|
||||
NEGATIVE CONTROL: absent name present = false
|
||||
read-back finished
|
||||
|
||||
## wishlist — seeded via POST /signup then POST /lists/<id>/create-item
|
||||
# NOTE: on first boot the image's own 'pnpm prisma db seed' was OOM-Killed at mem_limit 128M,
|
||||
# so the Role/Group rows were missing and EVERY signup failed with a misleading
|
||||
# 'User with username or email already exists' (real cause: FOREIGN KEY constraint).
|
||||
# Repaired by running the image's own seed once (docker update --memory 512m, run seed, back to 128m).
|
||||
wishlist version banner: 308
|
||||
list page code=200
|
||||
POSITIVE CONTROL occurrences of seed name: 2
|
||||
NEGATIVE CONTROL occurrences of absent name: 0
|
||||
context: '-[--><!--[-1--><div data-scope="dialog" data-part="title" id="dialog:s11:title" class="truncate text-xl font-bold text-wrap wrap-break-word md:text-2xl"><!---->SEED-WISHLIST-CANARY-7b2d44-20260921<!----></div><!--]--><!--]--> <!--[--><!--[-1--><button data-scope="dialog" data-par'
|
||||
|
||||
## glance — no user data
|
||||
glance has NO user data (catalog app_info declares none; the only persistent state is the
|
||||
first-boot seeded /app/config/glance.yml). Observable recorded instead:
|
||||
d9104faa1890d25cd77ed62eb2271da5 /app/config/glance.yml
|
||||
1
|
||||
glance HTTP 200 size=7842
|
||||
POSITIVE CONTROL (ASCII fragment "Kezd" on the page): 4
|
||||
NEGATIVE CONTROL: 0
|
||||
@@ -0,0 +1,105 @@
|
||||
# 03 — DRILL commit pushed, catalog synced, all four read 'update available'
|
||||
# 2026-09-21T12:27:11Z
|
||||
|
||||
## catalog commit
|
||||
573e41f DRILL: move four app pins for the power-cut update-arc measurement
|
||||
573e41f DRILL: move four app pins for the power-cut update-arc measurement
|
||||
templates/glance/.felhom.yml | 2 +-
|
||||
templates/glance/docker-compose.yml | 2 +-
|
||||
templates/uptime-kuma/docker-compose.yml | 2 +-
|
||||
templates/vikunja/.felhom.yml | 2 +-
|
||||
templates/vikunja/docker-compose.yml | 2 +-
|
||||
templates/wishlist/.felhom.yml | 2 +-
|
||||
templates/wishlist/docker-compose.yml | 2 +-
|
||||
7 files changed, 7 insertions(+), 7 deletions(-)
|
||||
|
||||
## the box's catalog clone after POST /api/sync
|
||||
573e41f DRILL: move four app pins for the power-cut update-arc measurement
|
||||
|
||||
## POST /api/sync said: Sablonok frissitve - frissitve: glance, vikunja, wishlist
|
||||
## (uptime-kuma NOT listed - its .felhom.yml was already at catalog_since 2026-09-21 from the
|
||||
## earlier session, and the syncer names only files whose content it rewrote. R-607 trap:
|
||||
## a POST /api/stacks/rescan was run straight after, and only then were the badges read.)
|
||||
## POST /api/stacks/rescan said: Rescan completed: 55 stacks found
|
||||
|
||||
## the four observables per app, AFTER sync+rescan
|
||||
== read at 2026-09-21T12:27:12.574916Z
|
||||
glance state=running deployed=True deploying=False updating=False phase=- label=-
|
||||
update_error=-
|
||||
hold_reason =-
|
||||
pinned ={"glance": "glanceapp/glance:v0.8.5"}
|
||||
installed={"glance": {"ref": "glanceapp/glance:v0.8.5", "digest": "sha256:32ab73d80f2b8b5fb0735b0431deb36b93fbb6b2fb43592449b0178c8b83e350", "at": "2026-09-21T12:11:56Z"}}
|
||||
template ={"glance": "glanceapp/glance:v0.8.5"}
|
||||
catalog ={"glance": "glanceapp/glance:v0.8.6"}
|
||||
health =healthy=True
|
||||
uptime-kuma state=running deployed=True deploying=False updating=False phase=done label=Frissítve
|
||||
update_error=-
|
||||
hold_reason =-
|
||||
pinned ={"uptime-kuma": "louislam/uptime-kuma:2.4.0"}
|
||||
installed={"uptime-kuma": {"ref": "louislam/uptime-kuma:2.4.0", "digest": "sha256:91e963bfda569ba115206e843febb446f473ab525add4e08b2b9e3beffa16985", "at": "2026-09-21T12:11:45Z"}}
|
||||
template ={"uptime-kuma": "louislam/uptime-kuma:2.4.0"}
|
||||
catalog ={"uptime-kuma": "louislam/uptime-kuma:2.5.0"}
|
||||
health =healthy=True
|
||||
vikunja state=running deployed=True deploying=False updating=False phase=- label=-
|
||||
update_error=-
|
||||
hold_reason =-
|
||||
pinned ={"vikunja": "vikunja/vikunja:2.3.0"}
|
||||
installed={"vikunja": {"ref": "vikunja/vikunja:2.3.0", "digest": "sha256:f6b80393c1998cd5cd0dc38d24762c59ab4c10000a6f1032ef5b554e262cab93", "at": "2026-09-21T12:11:44Z"}}
|
||||
template ={"vikunja": "vikunja/vikunja:2.3.0"}
|
||||
catalog ={"vikunja": "vikunja/vikunja:2.6.0"}
|
||||
health =healthy=True
|
||||
wishlist state=running deployed=True deploying=False updating=False phase=- label=-
|
||||
update_error=-
|
||||
hold_reason =-
|
||||
pinned ={"wishlist": "ghcr.io/cmintey/wishlist:v0.66.0"}
|
||||
installed={"wishlist": {"ref": "ghcr.io/cmintey/wishlist:v0.66.0", "digest": "sha256:073ab4de0f27a93a79410172bedfa3947bed4718f050396547516def790f1f49", "at": "2026-09-21T12:12:17Z"}}
|
||||
template ={"wishlist": "ghcr.io/cmintey/wishlist:v0.66.0"}
|
||||
catalog ={"wishlist": "ghcr.io/cmintey/wishlist:v0.67.0"}
|
||||
health =healthy=True
|
||||
|
||||
## the badge, both languages — ASCII-only fragments, with a negative control
|
||||
vikunja[hu] http=200 size=43698 | fragments hit: ['Friss'] | NEGATIVE CONTROL ZZZ-NOT-IN-ANY-PAGE-ZZZ hit: False
|
||||
BADGE: 'Fut'
|
||||
BADGE: '~50M RAM'
|
||||
BADGE: 'productivity'
|
||||
BADGE: 'Pi kompatibilis'
|
||||
vikunja[en] http=200 size=42909 | fragments hit: ['Update available'] | NEGATIVE CONTROL ZZZ-NOT-IN-ANY-PAGE-ZZZ hit: False
|
||||
BADGE: 'Running'
|
||||
BADGE: '~50M RAM'
|
||||
BADGE: 'productivity'
|
||||
BADGE: 'Runs on Pi'
|
||||
uptime-kuma[hu] http=200 size=43881 | fragments hit: ['Friss'] | NEGATIVE CONTROL ZZZ-NOT-IN-ANY-PAGE-ZZZ hit: False
|
||||
BADGE: 'Fut'
|
||||
BADGE: '~50M RAM'
|
||||
BADGE: 'dashboard'
|
||||
BADGE: 'Pi kompatibilis'
|
||||
uptime-kuma[en] http=200 size=43092 | fragments hit: ['Update available'] | NEGATIVE CONTROL ZZZ-NOT-IN-ANY-PAGE-ZZZ hit: False
|
||||
BADGE: 'Running'
|
||||
BADGE: '~50M RAM'
|
||||
BADGE: 'dashboard'
|
||||
BADGE: 'Runs on Pi'
|
||||
wishlist[hu] http=200 size=43666 | fragments hit: ['Friss'] | NEGATIVE CONTROL ZZZ-NOT-IN-ANY-PAGE-ZZZ hit: False
|
||||
BADGE: 'Fut'
|
||||
BADGE: '~30M RAM'
|
||||
BADGE: 'home'
|
||||
BADGE: 'Pi kompatibilis'
|
||||
wishlist[en] http=200 size=42878 | fragments hit: ['Update available'] | NEGATIVE CONTROL ZZZ-NOT-IN-ANY-PAGE-ZZZ hit: False
|
||||
BADGE: 'Running'
|
||||
BADGE: '~30M RAM'
|
||||
BADGE: 'home'
|
||||
BADGE: 'Runs on Pi'
|
||||
glance[hu] http=200 size=43734 | fragments hit: ['Friss'] | NEGATIVE CONTROL ZZZ-NOT-IN-ANY-PAGE-ZZZ hit: False
|
||||
BADGE: 'Fut'
|
||||
BADGE: '~20M RAM'
|
||||
BADGE: 'dashboard'
|
||||
BADGE: 'Pi kompatibilis'
|
||||
glance[en] http=200 size=42966 | fragments hit: ['Update available'] | NEGATIVE CONTROL ZZZ-NOT-IN-ANY-PAGE-ZZZ hit: False
|
||||
BADGE: 'Running'
|
||||
BADGE: '~20M RAM'
|
||||
BADGE: 'dashboard'
|
||||
BADGE: 'Runs on Pi'
|
||||
|
||||
## the exact sentence, verbatim (vikunja)
|
||||
HU: <span class="tag tag-warn" title="Ujabb valtozat erheto el ehhez az alkalmazashoz. A frissites inditasahoz nyomd meg a Frissites gombot.">Frissites elerheto — ma</span>
|
||||
(accents intact in the page; transliterated here only to keep this file ASCII-searchable)
|
||||
EN: <span class="tag tag-warn" title="A newer version of this app is available. Select the Update button to start it.">Update available — today</span>
|
||||
+29
@@ -0,0 +1,29 @@
|
||||
# 04b — INDEPENDENT CORROBORATION of scenario A, captured by the COORDINATOR
|
||||
#
|
||||
# WHY THIS FILE EXISTS, stated plainly: 25 minutes passed with no evidence written while the box
|
||||
# already showed the update complete, so the main session took its own read-only capture rather than
|
||||
# risk losing the measurement (R-320 — evidence off before it can be lost). The measuring agent's own
|
||||
# file `04-scenarioA-starting-cut.txt` then landed and is FULLER and BETTER than this one: it has the
|
||||
# timestamp table, the journal read off the stopped guest, and the seeded-data read-back.
|
||||
#
|
||||
# THIS FILE IS NOT A SECOND MEASUREMENT. It is the same box, read independently ~90 minutes later by
|
||||
# a different reader with a different session. Its only value is that it agrees.
|
||||
|
||||
READ AT: 2026-09-21, after the scenario completed. Source: live guest 9202, read-only.
|
||||
|
||||
AGREES WITH 04-scenarioA-starting-cut.txt ON:
|
||||
* the recovery line — interrupted in `verifying`, "the new version may have run; marking it
|
||||
Updating and RESUMING the health wait", resumed, `up -d`, healthy after 0s, DONE in 1m26s;
|
||||
* all FOUR version observables reading vikunja/vikunja:2.6.0 —
|
||||
pinned_images, installed_images, the live compose `image:` line, and `docker inspect`
|
||||
(created 2026-09-21T12:28:26Z, state running);
|
||||
* end state: state=running, updating=False, update_phase=done, label 'Frissitve',
|
||||
no update_error, no hold_reason, health_probe.healthy=True;
|
||||
* the update journal ABSENT (cleared), with a positive control that the directory searched is the
|
||||
real one — catalog-cache, debug-ring.log, encryption.key, metrics.db*, settings.json* beside it.
|
||||
|
||||
WHAT THIS FILE COULD NOT DO, and why — because a gap recorded is worth more than a gap hidden:
|
||||
the seeded-data read-back. 02-seeding.txt records the canary task name but not the account it was
|
||||
created under, and the coordinator would not guess credentials or read vikunja's database behind
|
||||
the app's back. "Through the front door" is the condition the measurement is worth anything under.
|
||||
The measuring agent had the credentials and did it: the task read back unchanged.
|
||||
@@ -0,0 +1,129 @@
|
||||
# 04 — SCENARIO A: vikunja 2.3.0 -> 2.6.0, cut decided in `starting`
|
||||
# guest 9202 (demo-hp-scratch), controller 0.260.0, 2026-09-21
|
||||
#
|
||||
# VERDICT: the box ended HONEST. The update resumed after the reboot and finished `done`;
|
||||
# all four version observables agree on 2.6.0; the seeded task read back unchanged; the
|
||||
# journal is gone; the household sentence is „Naprakesz" / "Up to date".
|
||||
#
|
||||
# THE ONE THING THAT DID NOT GO AS THE BRIEF ASSUMED — stated first because it changes
|
||||
# how capture (1) should be read:
|
||||
# The DECISION was taken in `starting` (12:28:26.244 UTC, poll cadence ~21 ms).
|
||||
# But `pct stop 9202` took 3 759 ms to return, and the update moved on inside that window.
|
||||
# The journal read OFF the stopped guest proves where the box actually died:
|
||||
# phase "verifying", not "starting".
|
||||
# `starting` on this box lasts about 0.6 s (12:28:26.244 -> vikunja's own 14:28:26.881+02:00
|
||||
# migration line). NO instrument available here can land a guest kill inside it: the poll is
|
||||
# fast enough (21 ms), the CUT COMMAND is not (3.8 s).
|
||||
# This costs nothing for the measurement, because RecoverUpdates handles `starting` and
|
||||
# `verifying` in the SAME branch (update.go:909 `case UpdatePhaseStarting, UpdatePhaseVerifying`).
|
||||
# Scenario A and Scenario B therefore exercise one recovery arm, not two. Recorded, not hidden.
|
||||
|
||||
## (1) TIMESTAMP TABLE
|
||||
poll target : GET /api/stacks/vikunja, HTTPS keep-alive from DooPlex
|
||||
measured poll cadence : ~21 ms (per-sample HTTP cost 1.1-1.3 ms + 20 ms sleep)
|
||||
Update pressed : 12:28:23.075 UTC (POST /api/stacks/vikunja/update -> {"accepted":true})
|
||||
phase AT THE DECISION : "starting" observed 12:28:26.244 UTC
|
||||
cut command : ssh -S <prewarmed master> demo-hp 'pct stop 9202'
|
||||
cut command LATENCY : 3 759 ms (returned 12:28:30.006 UTC, rc=0, no output)
|
||||
phase the box DIED in : "verifying" (from update-journal.json read off the stopped guest)
|
||||
guest restarted : pct start 9202 returned after 3.6 s; controller up 12:29:48 UTC
|
||||
total app downtime : ~79 s (vikunja stopped 12:28:26.9 .. restarted 12:29:47.3 UTC)
|
||||
|
||||
## (1b) the phase trace, verbatim from the poller
|
||||
12:28:21.860 updating=False phase='' label='' state=running err=''
|
||||
12:28:23.088 updating=True phase='safety-dump' label='Adatbázis pillanatkép…' state=running err=''
|
||||
12:28:23.110 updating=True phase='pinning' label='Új verzió letöltése…' state=running err=''
|
||||
12:28:23.132 updating=True phase='pulling' label='Új verzió letöltése…' state=running err=''
|
||||
12:28:26.244 updating=True phase='starting' label='Indítás az új verzióval…' state=running err=''
|
||||
--- DECISION at 12:28:26.244: phase='starting' -> firing cut: ssh -S /tmp/cm-demohp demo-hp 'pct stop 9202'
|
||||
--- CUT returned rc=0 after 3759 ms
|
||||
--- CUT stdout:
|
||||
--- CUT stderr:
|
||||
|
||||
=== TIMESTAMP TABLE (vikunja) ===
|
||||
poll cadence : measured below
|
||||
phase at decision : 'starting' at 12:28:26.244 UTC
|
||||
cut command : ssh -S /tmp/cm-demohp demo-hp 'pct stop 9202'
|
||||
cut command latency : 3759 ms (returned 12:28:30.006 UTC)
|
||||
|
||||
## (1c) update-journal.json read from the STOPPED guest
|
||||
# The brief's path was WRONG for this guest: /var/lib/lxc/9202/rootfs/var/lib/felhom is an
|
||||
# EMPTY MOUNTPOINT while the guest is stopped, because mp0 is a separate raw volume
|
||||
# (nvme-scratch:9202/vm-9202-disk-1.raw). The POSITIVE CONTROL failed there: no catalog-cache/.
|
||||
/bin/bash: line 131: pct: command not found
|
||||
# DOES attach mp0, and then the same path is real (catalog-cache/ present).
|
||||
/bin/bash: line 132: pct: command not found
|
||||
/bin/bash: line 132: pct: command not found
|
||||
# Remember to afterwards or refuses.
|
||||
{
|
||||
"updates": {
|
||||
"vikunja": {
|
||||
"phase": "verifying",
|
||||
"started_at": "2026-09-21T12:28:23.06699273Z",
|
||||
"prev_pin": {
|
||||
"vikunja": "vikunja/vikunja:2.3.0"
|
||||
},
|
||||
"prev_compose": "/opt/docker/stacks/vikunja/pre-update-compose.yml",
|
||||
"prev_applied": "/opt/docker/stacks/vikunja/pre-update-applied.yml",
|
||||
"proven_copy_at": "2026-09-21T12:26:12Z",
|
||||
"proven_tier": 1
|
||||
}
|
||||
}
|
||||
}
|
||||
## (2) RecoverUpdates log lines after the restart, VERBATIM
|
||||
2026/09/21 12:29:48 update.go:909: [WARN] [stacks] update recovery: vikunja was interrupted in verifying (started 2026-09-21T12:28:23Z) — the new version may have run; marking it Updating and RESUMING the health wait
|
||||
2026/09/21 12:29:48 update.go:951: [INFO] [stacks] update vikunja: resuming after a controller restart — `up -d` then the health wait
|
||||
2026/09/21 12:29:48 update.go:858: [INFO] [stacks] update vikunja: phase verifying
|
||||
2026/09/21 12:29:48 update.go:652: [INFO] [stacks] update vikunja: healthy after 0s (the app's health check passed)
|
||||
2026/09/21 12:29:49 update.go:658: [INFO] [stacks] update vikunja: DONE in 1m26s
|
||||
|
||||
## (3) THE FOUR VERSION OBSERVABLES, SIDE BY SIDE (after recovery)
|
||||
1_pinned_images : {"vikunja": "vikunja/vikunja:2.6.0"}
|
||||
2_installed_images : {"vikunja": "vikunja/vikunja:2.6.0"} | digest: {"vikunja": "sha256:417ada6f94e81f0267a"}
|
||||
(updating=False phase=done label=Frissítve err=- hold=-)
|
||||
3_live compose line: image: vikunja/vikunja:2.6.0
|
||||
4_docker inspect : vikunja/vikunja:2.6.0 | RepoDigest: sha256:417ada6f94e81f0267a | started 2026-09-21T12:29:47.287254768Z
|
||||
-> all four name vikunja/vikunja:2.6.0; installed digest and the running container's
|
||||
RepoDigest are the same sha256:417ada6f94e81f0267a...
|
||||
|
||||
## (4) the seeded data, read back through vikunja's own front door
|
||||
vikunja version: v2.6.0
|
||||
TASK 1 'SEED-VIKUNJA-CANARY-9f3c1e-20260921' desc= 'power-cut drill canary'
|
||||
POSITIVE CONTROL: seed present = True
|
||||
NEGATIVE CONTROL: absent string present = False
|
||||
|
||||
## (4b) the vikunja MIGRATION line, VERBATIM from its container log
|
||||
3:time=2026-09-21T14:28:26.881+02:00 level=INFO msg="Running migrations…"
|
||||
5:time=2026-09-21T14:28:26.915+02:00 level=INFO msg="Ran all migrations successfully."
|
||||
9:time=2026-09-21T14:28:26.922+02:00 level=INFO msg="Vikunja version v2.6.0"
|
||||
14:time=2026-09-21T14:29:47.713+02:00 level=INFO msg="Running migrations…"
|
||||
16:time=2026-09-21T14:29:47.734+02:00 level=INFO msg="Ran all migrations successfully."
|
||||
20:time=2026-09-21T14:29:47.745+02:00 level=INFO msg="Vikunja version v2.6.0"
|
||||
-> the FIRST migration ran at 14:28:26.881+02:00 = 12:28:26.881 UTC — 0.64 s AFTER the cut
|
||||
decision and 3.1 s BEFORE pct stop returned. The 2.6.0 schema migration had ALREADY
|
||||
been applied to the customer's SQLite database when the power went. The pin did not
|
||||
roll back (that only happens in the pinning/pulling branch), so old binary vs migrated
|
||||
database never happened here — but it is the failure this branch is one step away from.
|
||||
|
||||
## (5) the app page's sentence to the household, both languages
|
||||
vikunja [hu] http=200
|
||||
[Naprak] ...tems:center;gap:.5rem"> <span class="stack-state-badge state-run">Fut</span> <span class="tag tag-ok" title="Ez az alkalmazás a legfrissebb elérhető változatot futtatja.">Naprakész</span> <a href="https://tasks.enkisfelhom.hu" target="_blank" class="btn btn-sm btn-outline">Megnyitás ↗</a> <a href="/stacks/vikunja/logs" class="btn btn-sm btn-outl
|
||||
NEGATIVE CONTROL 'ZZZNOTPRESENT' found: False
|
||||
vikunja [en] http=200
|
||||
[Up to date] ...align-items:center;gap:.5rem"> <span class="stack-state-badge state-run">Running</span> <span class="tag tag-ok" title="This app is running the newest version available.">Up to date</span> <a href="https://tasks.enkisfelhom.hu" target="_blank" class="btn btn-sm btn-outline">Open ↗</a> <a href="/stacks/vikunja/logs" class="btn btn-sm btn-outline"
|
||||
NEGATIVE CONTROL 'ZZZNOTPRESENT' found: False
|
||||
|
||||
## (6) done or HELD, and did a journal entry survive?
|
||||
ended: update_phase=done, update_phase_label='Frissitve', updating=false, hold_reason=none
|
||||
controller log: 'update vikunja: DONE in 1m26s'
|
||||
POSITIVE CONTROL: catalog-cache present -> this IS the controller data dir
|
||||
ls: cannot access '/var/lib/docker/volumes/felhom-controller-data/_data/data/update-journal.json': No such file or directory
|
||||
-> no journal entry survived the reboot.
|
||||
|
||||
## STOP CONDITIONS — none tripped
|
||||
seeded data gone/unreadable : NO (read back byte-identical)
|
||||
updating:true that never clears : NO (cleared 1.3 s after boot)
|
||||
journal surviving a 2nd reboot : N/A, no journal survived the 1st
|
||||
pin naming one version, container another : NO
|
||||
resumed update retrying in a loop : NO (one resume, one success)
|
||||
HOLD with no household sentence : N/A, no hold
|
||||
@@ -0,0 +1,128 @@
|
||||
# 05 — SCENARIO B: uptime-kuma 2.4.0 -> 2.5.0, cut during `verifying` (the health wait)
|
||||
# guest 9202 (demo-hp-scratch), controller 0.260.0, 2026-09-21
|
||||
#
|
||||
# VERDICT: the box ended HONEST. The update resumed after the reboot and finished `done`;
|
||||
# all four version observables agree on 2.5.0; the seeded monitor read back through
|
||||
# uptime-kuma's own socket.io front door; no journal survived; the household sentence is
|
||||
# „Naprakesz" / "Up to date".
|
||||
#
|
||||
# INSTRUMENT FINDING (this one matters for reading BOTH A and B):
|
||||
# `pct stop 9202` RETURNS after ~3.0-3.8 s, but the guest stops answering after ~1.1 s.
|
||||
# Measured here by firing the cut in a THREAD and continuing to poll through it:
|
||||
# decision 12:32:32.271, last successful API sample 12:32:33.345 (+1 075 ms),
|
||||
# command return 12:32:35.303 (+3 032 ms).
|
||||
# So the kill lands early and the rest of the 3 s is teardown. The practical consequence,
|
||||
# proven in Scenario A: `starting` lasts about 0.6 s on this box, so NO cut driven this way
|
||||
# can land inside `starting` — by the time the guest dies the update is in `verifying`.
|
||||
# A and B therefore exercise the SAME RecoverUpdates branch
|
||||
# (update.go:909 `case UpdatePhaseStarting, UpdatePhaseVerifying`). Stated, not hidden.
|
||||
|
||||
## (1) TIMESTAMP TABLE
|
||||
poll target : GET /api/stacks/uptime-kuma, HTTPS keep-alive from DooPlex
|
||||
measured poll cadence : ~21 ms (per-sample HTTP cost 1.1-1.3 ms + 20 ms sleep)
|
||||
Update pressed : 12:32:26.374 UTC (POST /api/stacks/uptime-kuma/update -> accepted)
|
||||
phase AT THE DECISION : "verifying" observed 12:32:32.271 UTC
|
||||
cut command : ssh -S <prewarmed master> demo-hp 'pct stop 9202'
|
||||
cut command LATENCY : 3 016 ms (rc=0, returned 12:32:35.303 UTC, no stdout/stderr)
|
||||
box stopped ANSWERING : 12:32:33.345 UTC was the LAST successful sample (+1 075 ms)
|
||||
phase the box DIED in : "verifying" (update-journal.json read off the stopped guest)
|
||||
guest restarted : pct start 9202; controller up 12:33:21 UTC
|
||||
update concluded : 12:33:26 UTC, "DONE in 1m0s"
|
||||
|
||||
## (1b) the phase trace, verbatim from the poller
|
||||
12:32:25.147 updating=False phase='' state=running err='' hold=''
|
||||
12:32:26.373 updating=True phase='checking' state=running err='' hold=''
|
||||
12:32:26.394 updating=True phase='safety-dump' state=running err='' hold=''
|
||||
12:32:26.416 updating=True phase='pulling' state=running err='' hold=''
|
||||
12:32:27.627 updating=True phase='starting' state=running err='' hold=''
|
||||
12:32:28.748 updating=True phase='starting' state=degraded err='' hold=''
|
||||
12:32:32.271 updating=True phase='verifying' state=starting err='' hold=''
|
||||
--- DECISION at 12:32:32.271: phase='verifying' -> firing (in a thread): ssh -S /tmp/cm-demohp demo-hp 'pct stop 9202'
|
||||
12:32:33.367 ERR ConnectionResetError: [Errno 104] Connection reset by peer
|
||||
12:32:33.393 ERR TimeoutError: timed out
|
||||
|
||||
=== TIMESTAMP TABLE (uptime-kuma) ===
|
||||
phase at decision : 'verifying' at 12:32:32.271 UTC
|
||||
cut command : ssh -S /tmp/cm-demohp demo-hp 'pct stop 9202'
|
||||
cut command latency : 3016 ms (rc=0, returned 12:32:35.303 UTC)
|
||||
cut stdout/stderr : '' / ''
|
||||
LAST SUCCESSFUL SAMPLE : 12:32:33.345 UTC <- the box was still answering here
|
||||
i.e. 1075 ms after the decision
|
||||
|
||||
## (1c) update-journal.json read from the STOPPED guest (pct mount 9202 first; pct unmount after)
|
||||
POSITIVE CONTROL: catalog-cache present -> real dir
|
||||
--- journal ---
|
||||
{
|
||||
"updates": {
|
||||
"uptime-kuma": {
|
||||
"phase": "verifying",
|
||||
"started_at": "2026-09-21T12:32:26.364656602Z",
|
||||
"prev_pin": {
|
||||
"uptime-kuma": "louislam/uptime-kuma:2.4.0"
|
||||
},
|
||||
"prev_compose": "/opt/docker/stacks/uptime-kuma/pre-update-compose.yml",
|
||||
"prev_applied": "/opt/docker/stacks/uptime-kuma/pre-update-applied.yml",
|
||||
"proven_copy_at": "2026-09-21T12:16:12Z",
|
||||
"proven_tier": 1
|
||||
}
|
||||
}
|
||||
}
|
||||
## (2) RecoverUpdates log lines after the restart, VERBATIM
|
||||
2026/09/21 12:33:21 update.go:909: [WARN] [stacks] update recovery: uptime-kuma was interrupted in verifying (started 2026-09-21T12:32:26Z) — the new version may have run; marking it Updating and RESUMING the health wait
|
||||
2026/09/21 12:33:21 update.go:951: [INFO] [stacks] update uptime-kuma: resuming after a controller restart — `up -d` then the health wait
|
||||
2026/09/21 12:33:21 update.go:858: [INFO] [stacks] update uptime-kuma: phase verifying
|
||||
2026/09/21 12:33:26 update.go:652: [INFO] [stacks] update uptime-kuma: healthy after 5s (the app's health check passed)
|
||||
2026/09/21 12:33:26 update.go:658: [INFO] [stacks] update uptime-kuma: DONE in 1m0s
|
||||
|
||||
## (3) THE FOUR VERSION OBSERVABLES, SIDE BY SIDE (after recovery)
|
||||
1_pinned_images : {"uptime-kuma": "louislam/uptime-kuma:2.5.0"}
|
||||
2_installed_images : {"uptime-kuma": "louislam/uptime-kuma:2.5.0"} | digest: {"uptime-kuma": "sha256:a8610b3b4c38077922b"}
|
||||
(updating=False phase=done label=Frissítve err=- hold=-)
|
||||
3_live compose line: image: louislam/uptime-kuma:2.5.0
|
||||
4_docker inspect : louislam/uptime-kuma:2.5.0 | RepoDigest: sha256:a8610b3b4c38077922b | started 2026-09-21T12:33:19.417787988Z
|
||||
-> all four name louislam/uptime-kuma:2.5.0; installed digest and the running
|
||||
container's RepoDigest are the same sha256:a8610b3b4c38077922b...
|
||||
|
||||
## (4) the seeded data, read back through uptime-kuma's OWN front door (socket.io, the same
|
||||
## API the browser uses), run from inside the container with its own socket.io-client
|
||||
MONITORS: ["SEED-KUMA-CANARY-4a91c7-20260921"]
|
||||
POSITIVE CONTROL: seed present = true
|
||||
NEGATIVE CONTROL: absent name present = false
|
||||
read-back finished
|
||||
NOTE: the FIRST read-back attempt, run 12 s after the app came up, TIMED OUT waiting for
|
||||
the monitorList event — the app was listening but not yet serving. Re-run 20 s later it
|
||||
answered. Recorded because a single timeout here reads exactly like data loss and is not.
|
||||
|
||||
## (5) the app page's sentence to the household, both languages
|
||||
uptime-kuma [hu] http=200
|
||||
[Naprak] ...tems:center;gap:.5rem"> <span class="stack-state-badge state-run">Fut</span> <span class="tag tag-ok" title="Ez az alkalmazás a legfrissebb elérhető változatot futtatja.">Naprakész</span> <a href="https://status.enkisfelhom.hu" target="_blank" class="btn btn-sm btn-outline">Megnyitás ↗</a> <a href="/stacks/uptime-kuma/logs" class="btn btn-sm btn
|
||||
NEGATIVE CONTROL 'ZZZNOTPRESENT' found: False
|
||||
uptime-kuma [en] http=200
|
||||
[Up to date] ...align-items:center;gap:.5rem"> <span class="stack-state-badge state-run">Running</span> <span class="tag tag-ok" title="This app is running the newest version available.">Up to date</span> <a href="https://status.enkisfelhom.hu" target="_blank" class="btn btn-sm btn-outline">Open ↗</a> <a href="/stacks/uptime-kuma/logs" class="btn btn-sm btn-out
|
||||
NEGATIVE CONTROL 'ZZZNOTPRESENT' found: False
|
||||
|
||||
## (6) done or HELD, and did a journal entry survive?
|
||||
ended: update_phase=done, update_phase_label='Frissitve', updating=false, hold_reason=none
|
||||
controller log: 'update uptime-kuma: DONE in 1m0s'
|
||||
POSITIVE CONTROL: catalog-cache present -> this IS the controller data dir
|
||||
ls: cannot access '/var/lib/docker/volumes/felhom-controller-data/_data/data/update-journal.json': No such file or directory
|
||||
-> no journal entry survived the reboot.
|
||||
|
||||
## STOP CONDITIONS — none tripped
|
||||
seeded data gone/unreadable : NO (monitor read back by name)
|
||||
updating:true that never clears : NO (cleared 5 s after boot)
|
||||
journal surviving a 2nd reboot : N/A, no journal survived the 1st
|
||||
pin naming one version, container another : NO
|
||||
resumed update retrying in a loop : NO (one resume, one success)
|
||||
HOLD with no household sentence : N/A, no hold
|
||||
|
||||
## SEPARATE DEFECT FOUND WHILE SEEDING (not an update-arc finding, filed here so it is not lost)
|
||||
uptime-kuma 2.4.0 first boot sits in its SETUP-DATABASE wizard:
|
||||
[SETUP-DATABASE] INFO: Starting Setup Database
|
||||
[SETUP-DATABASE] INFO: Waiting for user action...
|
||||
The main socket.io server never starts until a database type is chosen. The Felhom controller
|
||||
nevertheless reported the app RUNNING and HEALTHY — its probe is `http :3001` and the wizard
|
||||
answers 302 on that port. So the box tells the household the monitoring app is fine while it is
|
||||
actually parked on an un-passed wizard, with no monitors and no login.
|
||||
Passed here through the app's own front door: POST /setup-database {"dbConfig":{"type":"sqlite"}}
|
||||
-> {"ok":true}. A catalog fix would pin the DB type at deploy time so first boot never stops.
|
||||
@@ -0,0 +1,35 @@
|
||||
# 05b — SCENARIO B (cut in `verifying`, uptime-kuma) — COORDINATOR capture
|
||||
#
|
||||
# PROVENANCE: written by the main session, read-only, live from guest 9202, after the measuring
|
||||
# agent again went a long interval without writing evidence while the box already showed the
|
||||
# scenario complete. If the agent's own `05-*` file lands it is the fuller record; this exists so
|
||||
# the measurement could not be lost (R-320). Nothing below is inferred.
|
||||
|
||||
== THE RECOVERY, verbatim from the controller log ==
|
||||
2026/09/21 12:33:21 update.go:909: [WARN] [stacks] update recovery: uptime-kuma was interrupted in verifying (started 2026-09-21T12:32:26Z) — the new version may have run; marking it Updating and RESUMING the health wait
|
||||
2026/09/21 12:33:21 update.go:951: [INFO] [stacks] update uptime-kuma: resuming after a controller restart — `up -d` then the health wait
|
||||
2026/09/21 12:33:21 update.go:858: [INFO] [stacks] update uptime-kuma: phase verifying
|
||||
2026/09/21 12:33:26 update.go:652: [INFO] [stacks] update uptime-kuma: healthy after 5s (the app's health check passed)
|
||||
2026/09/21 12:33:26 update.go:658: [INFO] [stacks] update uptime-kuma: DONE in 1m0s
|
||||
|
||||
== THE FOUR VERSION OBSERVABLES, SIDE BY SIDE ==
|
||||
pinned_images : {"uptime-kuma": "louislam/uptime-kuma:2.5.0"}
|
||||
installed_images : {"uptime-kuma": "louislam/uptime-kuma:2.5.0"}
|
||||
live compose line: image: louislam/uptime-kuma:2.5.0
|
||||
docker inspect : louislam/uptime-kuma:2.5.0 | running
|
||||
-> ALL FOUR AGREE. (catalog_images also 2.5.0 — the drill bump was still in place.)
|
||||
|
||||
== END STATE ==
|
||||
state=running updating=False update_phase=done label='Frissitve'
|
||||
update_error=(none) hold_reason=(none) health_probe.healthy=True
|
||||
Total time from interruption to done: 1m0s.
|
||||
|
||||
== NOT CAPTURED HERE ==
|
||||
the timestamp table (phase at decision, cut latency), the seeded-monitor read-back through
|
||||
uptime-kuma's own front door, and the page sentence in both languages. Those need the agent's
|
||||
poller output and its seeded session. Recorded as GAPS, not as passes.
|
||||
|
||||
== VERDICT ON WHAT IS MEASURED ==
|
||||
The `verifying` cut ends HONEST for a second app, with a different health-check shape (5s rather
|
||||
than 0s to pass). Resumed, completed, all four observables agree, no hold, no stuck `Updating`.
|
||||
None of the brief's STOP conditions appeared in what was captured.
|
||||
@@ -0,0 +1,47 @@
|
||||
# 06 — THE REVERT IS STILL OWED. Written 2026-09-21 by the measuring session.
|
||||
|
||||
## State right now
|
||||
app-catalog-felhom.eu `main` is at 573e41f59bb3c3f01a04f658f1c02bdfbd427a07
|
||||
"DRILL: move four app pins for the power-cut update-arc measurement"
|
||||
The pre-drill commit is ff9717d3794974724e09f8fd58abf058d4cdc2d0
|
||||
Working tree clean. The drill bump is LIVE on main and on every box that syncs the catalog.
|
||||
|
||||
## Why it was not reverted by the session that made it
|
||||
The operator brief said "Your catalog commits MUST be reverted before you finish."
|
||||
The coordinating session stood this session down from Scenario C and said it would do the
|
||||
REVERT itself, because it needs the catalog left bumped until Scenario C and one further
|
||||
phase are finished. Reverting under a run in progress on the same box would have changed
|
||||
the catalog beneath that measurement.
|
||||
So the revert was NOT done here — and this file exists so that is a recorded hand-off and
|
||||
not a silently dropped fence. If the later phases are finished and no REVERT commit follows
|
||||
573e41f, THIS IS THE OUTSTANDING ITEM.
|
||||
|
||||
## Exactly what the revert must restore (verified, not assumed)
|
||||
7 files, 7 insertions, 7 deletions — nothing else is in the drill commit:
|
||||
templates/vikunja/docker-compose.yml vikunja/vikunja:2.6.0 -> 2.3.0
|
||||
templates/uptime-kuma/docker-compose.yml louislam/uptime-kuma:2.5.0 -> 2.4.0
|
||||
templates/wishlist/docker-compose.yml ghcr.io/cmintey/wishlist:v0.67.0 -> v0.66.0
|
||||
templates/glance/docker-compose.yml glanceapp/glance:v0.8.6 -> v0.8.5
|
||||
templates/vikunja/.felhom.yml catalog_since "2026-09-21" -> "2026-07-18"
|
||||
templates/glance/.felhom.yml catalog_since "2026-09-21" -> "2026-07-18"
|
||||
templates/wishlist/.felhom.yml catalog_since "2026-09-21" -> "2026-07-19"
|
||||
(templates/uptime-kuma/.felhom.yml was ALREADY at catalog_since "2026-09-21" before the drill,
|
||||
from the earlier session's own drill+revert pair, so it is not in the diff and must NOT be
|
||||
moved back to an older date.)
|
||||
|
||||
The tree object at ff9717d is bef76c8ccb7b2be919e22bc689213ef3199ee86e — a correct revert
|
||||
must produce a tree identical to that one for these seven paths.
|
||||
CHECKED (not applied): `git diff 573e41f ff9717d | git apply --check` passes cleanly.
|
||||
|
||||
## The command
|
||||
cd /mnt/5_hdd/felhom.eu/git/app-catalog-felhom.eu
|
||||
git revert --no-edit 573e41f # or apply the reverse diff and commit as "REVERT ..."
|
||||
python3 scripts/catalog_gates.py --fast # must be all-green BEFORE the push
|
||||
git push origin main # the pre-push hook re-runs the gates; never --no-verify
|
||||
# then on any box that must see it back: POST /api/sync AND POST /api/stacks/rescan (R-607)
|
||||
|
||||
## A warning for whoever presses Update after the revert
|
||||
vikunja and uptime-kuma on guest 9202 are now INSTALLED AHEAD of the reverted catalog
|
||||
(2.6.0 vs 2.3.0, 2.5.0 vs 2.4.0). Controller v0.260.0 refuses that as a downgrade —
|
||||
"update REFUSED (downgrade): installed is provably NEWER than the catalog on every differing
|
||||
service" — and the badge reads „Naprakesz". That is CORRECT behaviour (R-524), not a fault.
|
||||
+76
@@ -0,0 +1,76 @@
|
||||
# 07 — SCENARIO C: the CONTROLLER-ONLY restart during an app update (wishlist v0.66.0 -> v0.67.0)
|
||||
#
|
||||
# Run by the main session on 2026-09-21. This is the scenario that matters most for R-608: it is
|
||||
# EXACTLY what a controller self-update does to a running app update — the controller container
|
||||
# restarts, the app's own containers keep running. Guest 9202, controller v0.260.0 (the lock is NOT
|
||||
# in this build; that is deliberate — this measures the behaviour the lock is meant to make moot).
|
||||
|
||||
== TIMESTAMP TABLE ==
|
||||
poll cadence : 200 ms
|
||||
app confirmed healthy : 12:35:53.969 UTC (state=running health=True)
|
||||
Update pressed : 12:35:53.971 UTC -> 202 "Frissites elindult"
|
||||
phase AT THE DECISION : "starting" observed 12:36:33.136 UTC
|
||||
cut command : ssh demo-hp 'pct exec 9202 -- systemctl restart felhom-controller-bootstrap.service'
|
||||
cut command LATENCY : 1 675 ms (rc=0)
|
||||
phase the box RECORDED : "verifying" (from the controller's own recovery line)
|
||||
|
||||
-> THE SAME INSTRUMENT FINDING AS A AND B, REPRODUCED WITH A DIFFERENT AND MUCH FASTER CUT.
|
||||
The decision was taken at `starting` and the box still recorded `verifying`. A controller
|
||||
restart returns in 1.7 s where `pct stop` took 3.0-3.8 s, and `starting` STILL could not be
|
||||
caught. Measured on this box, `starting` lasts well under a second for these apps.
|
||||
**`RecoverUpdates` handles `starting` and `verifying` in ONE branch (`update.go:909`), so all
|
||||
three scenarios exercise the same recovery arm** — the arm that says "the new version may have
|
||||
run". That is the arm the brief wanted measured, and it is measured three times.
|
||||
|
||||
== THE APP'S OWN CONTAINERS DURING THE CONTROLLER RESTART ==
|
||||
felhom-controller 0.260.0 Up 1 second (health: starting) <- restarted
|
||||
wishlist ghcr.io/cmintey/wishlist:v0.67.0 Up 1 second (health: starting) <- the update's own `up -d`
|
||||
uptime-kuma louislam/uptime-kuma:2.5.0 Up 3 minutes (healthy) <- UNDISTURBED
|
||||
vikunja vikunja/vikunja:2.6.0 Up 3 minutes <- UNDISTURBED
|
||||
glance glanceapp/glance:v0.8.5 Up 3 minutes (healthy) <- UNDISTURBED
|
||||
filebrowser gtstef/filebrowser:1.3.3-stable Up 3 minutes (healthy) <- UNDISTURBED
|
||||
-> The controller restart does NOT restart the apps. Only the app the update itself was recreating
|
||||
shows a new uptime, and that is the update's `up -d`, not the restart.
|
||||
|
||||
== THE RECOVERY, verbatim ==
|
||||
2026/09/21 12:36:34 update.go:909: [WARN] [stacks] update recovery: wishlist was interrupted in verifying (started 2026-09-21T12:35:54Z) — the new version may have run; marking it Updating and RESUMING the health wait
|
||||
2026/09/21 12:36:34 update.go:951: [INFO] [stacks] update wishlist: resuming after a controller restart — `up -d` then the health wait
|
||||
2026/09/21 12:36:35 update.go:858: [INFO] [stacks] update wishlist: phase verifying
|
||||
2026/09/21 12:36:45 update.go:652: [INFO] [stacks] update wishlist: healthy after 10s (the app's health check passed)
|
||||
2026/09/21 12:36:45 update.go:658: [INFO] [stacks] update wishlist: DONE in 51s
|
||||
|
||||
== THE FOUR VERSION OBSERVABLES, SIDE BY SIDE ==
|
||||
pinned_images : {"wishlist": "ghcr.io/cmintey/wishlist:v0.67.0"}
|
||||
installed_images : {"wishlist": "ghcr.io/cmintey/wishlist:v0.67.0"}
|
||||
live compose line: image: ghcr.io/cmintey/wishlist:v0.67.0
|
||||
docker inspect : ghcr.io/cmintey/wishlist:v0.67.0 | running
|
||||
-> ALL FOUR AGREE.
|
||||
|
||||
== END STATE, AND WHAT THE HOUSEHOLD SEES ==
|
||||
state=running updating=False update_phase=done label='Frissitve'
|
||||
update_error=(none) hold_reason=(none) health=True
|
||||
update-journal.json : ABSENT (cleared)
|
||||
app page, Hungarian : <span class="tag tag-ok" title="Ez az alkalmazas a legfrissebb elerheto valtozatot futtatja.">Naprakesz</span>
|
||||
app page, English : <span class="tag tag-ok" title="This app is running the newest version available.">Up to date</span>
|
||||
(accents transliterated here only to keep this file ASCII-searchable; intact in the page)
|
||||
no `data-update-error` block on either page — there is nothing to apologise for.
|
||||
|
||||
SEARCH CONTROLS: POSITIVE 'Naprak' in hu = 1 · POSITIVE 'Up to date' in en = 1 ·
|
||||
NEGATIVE 'ZZZ-not-present' = 0
|
||||
|
||||
== STOP CONDITIONS — none tripped ==
|
||||
updating stuck true : NO journal surviving : NO pin vs running image : AGREE
|
||||
retry loop : NO hold with no sentence : N/A (no hold)
|
||||
data : wishlist's seeded list item was NOT re-read by this session — see the gap note below.
|
||||
|
||||
== GAP, STATED ==
|
||||
The seeded-data read-back through wishlist's front door was not repeated after this scenario. The
|
||||
measuring agent seeded and read it back BEFORE the drill (02-seeding.txt) and the app is healthy
|
||||
and serving on v0.67.0 after it, but "the data is still there" is NOT asserted here for C. A and B
|
||||
both carry a real post-cut read-back; C does not.
|
||||
|
||||
== VERDICT ==
|
||||
A controller-only restart during an app update ends HONEST. The update resumes, completes, the
|
||||
other apps are untouched, and the household is shown a clean „Naprakesz" with no error. This is
|
||||
the behaviour v0.261.0's lock makes unnecessary rather than fixes — worth knowing, because it
|
||||
means the lock is defence in depth, not a repair of something broken.
|
||||
@@ -0,0 +1,48 @@
|
||||
# 08 — THE v0.261.0 LOCK, PROVEN LIVE (R-608), with a positive AND a negative control
|
||||
#
|
||||
# Guest 9202, controller v0.261.0, 2026-09-21. The question is not "is the callback wired" but
|
||||
# "does the swap actually refuse, and is it OUR refusal doing it".
|
||||
|
||||
== WHY A CONTROL WAS NEEDED, AND WHY THIS ONE IS THE RIGHT ONE ==
|
||||
Guest 9202 has NO host agent wired, so `TriggerUpdate` refuses anyway — with the AGENT's sentence.
|
||||
That makes it the perfect negative control: in `TriggerUpdate` the new app-update check sits BEFORE
|
||||
the agent check, so if the lock fires we see OUR sentence and if it does not we see the AGENT's.
|
||||
Two different sentences, one probe. It also makes the probe SAFE: no swap can reach a machine.
|
||||
|
||||
(Self-update is `enabled: false` on this guest by design. It was turned on for this probe and
|
||||
turned back OFF immediately afterwards — the original file was copied to
|
||||
/root/controller.yaml.pre-lock-probe first and restored from it. Verified reverted below.)
|
||||
|
||||
== CONTROL A — self-update status, no app update running ==
|
||||
GET /api/selfupdate/status -> {"ok":true,"data":{"running":false}}
|
||||
|
||||
== CONTROL B (NEGATIVE) — the MANUAL trigger with NO app update running ==
|
||||
12:39:55.917 POST /api/selfupdate/update
|
||||
-> {"ok":false,"error":"A frissites nem erheto el (nincs gazda-ugynok)"}
|
||||
i.e. the AGENT refusal. The lock is NOT firing, because nothing is in flight. Correct.
|
||||
|
||||
== THE PROBE (POSITIVE) — the MANUAL trigger WHILE an app update is in flight ==
|
||||
12:39:55.941 POST /api/stacks/uptime-kuma/update -> 202 "Frissites elindult"
|
||||
12:39:56.241 app update IN FLIGHT, phase = safety-dump
|
||||
12:39:56.244 POST /api/selfupdate/update
|
||||
-> {"ok":false,"error":"Egy alkalmazas frissitese eppen folyamatban van. A vezerlo frissitese utana inditható."}
|
||||
THE SENTENCE CHANGED. That is the v0.261.0 gate firing, and firing BEFORE the agent check.
|
||||
12:39:56.270 GET /api/selfupdate/status -> {"running":false} — nothing was started.
|
||||
|
||||
(accents transliterated here only to keep this file ASCII-searchable; intact on the wire)
|
||||
|
||||
== WHAT THE PHASE TELLS US, and it is worth stating ==
|
||||
The probe landed in `safety-dump`. That phase is NOT covered by the pre-existing `backupRunning`
|
||||
gate — only `backing-up` is, because that one takes the backup single-flight. So the probe hit
|
||||
exactly the window the new lock exists for, rather than a window that was already protected.
|
||||
|
||||
== THE REVERSE DIRECTION — NOT staged live, and why ==
|
||||
`UpdatePreflight` refusing an app update with reason `self_updating` while the controller swaps is
|
||||
covered by `TestR608_PreflightRefusesWhileTheControllerSwaps` (a consequence test, red-proofed by
|
||||
deleting the block). Staging it LIVE would need a real controller swap in flight, which needs a host
|
||||
agent this guest does not have and a genuine image swap mid-measurement. **Not measured live. Stated
|
||||
as a gap rather than implied.**
|
||||
|
||||
== CONFIG RESTORED ==
|
||||
self_update.enabled is back to `false` on guest 9202, verified by re-reading the file after the
|
||||
restore, and the controller was restarted on the restored config.
|
||||
@@ -0,0 +1,12 @@
|
||||
2026-09-21T12:55:11Z === pass 1/3
|
||||
2026-09-21T12:55:11Z glance BEHIND and within a major — pressing Update
|
||||
2026-09-21T12:55:11Z glance REFUSED reason=downgrade TERMINAL — will not press again. Ez a változat újabb a katalógusban lévőnél — visszalépés csak az üzemeltető kérésére.
|
||||
2026-09-21T12:55:11Z uptime-kuma BEHIND and within a major — pressing Update
|
||||
2026-09-21T12:55:11Z uptime-kuma REFUSED reason=downgrade TERMINAL — will not press again. Ez a változat újabb a katalógusban lévőnél — visszalépés csak az üzemeltető kérésére.
|
||||
2026-09-21T12:55:11Z vikunja BEHIND and within a major — pressing Update
|
||||
2026-09-21T12:55:11Z vikunja REFUSED reason=downgrade TERMINAL — will not press again. Ez a változat újabb a katalógusban lévőnél — visszalépés csak az üzemeltető kérésére.
|
||||
2026-09-21T12:55:11Z wishlist BEHIND and within a major — pressing Update
|
||||
2026-09-21T12:55:11Z wishlist REFUSED reason=downgrade TERMINAL — will not press again. Ez a változat újabb a katalógusban lévőnél — visszalépés csak az üzemeltető kérésére.
|
||||
2026-09-21T12:56:11Z === pass 2/3
|
||||
2026-09-21T12:57:11Z === pass 3/3
|
||||
2026-09-21T12:57:12Z === summary outcomes={} never_again=['glance', 'uptime-kuma', 'vikunja', 'wishlist']
|
||||
@@ -0,0 +1,75 @@
|
||||
# 09 — THE UNATTENDED NIGHT: an app updates with nobody pressing anything
|
||||
#
|
||||
# This is the phase the previous session's brief contained and **did not run, and did not say so**
|
||||
# (R-611). It runs here. The caller is `unattended-caller.py` in this directory — EVIDENCE, NOT
|
||||
# PRODUCT: it presses exactly `POST /api/stacks/<n>/update`, the same button a person presses, and
|
||||
# nothing else. No controller code was added for it. Guest 9202, controller v0.261.0.
|
||||
|
||||
## Scenario F — the success night: PROVEN
|
||||
|
||||
`uptime-kuma` **2.5.0 → 2.5.1** was applied by the caller with no human action, against a real
|
||||
one-step catalog edge (DRILL 2, `ae08a037fd68`). Confirmed on the box afterwards:
|
||||
|
||||
uptime-kuma installed = louislam/uptime-kuma:2.5.1 catalog = louislam/uptime-kuma:2.5.1
|
||||
updating = False update_phase = done hold = none
|
||||
|
||||
**HONEST GAP, and it is the coordinator's own instrumentation error:** that first run's stdout was
|
||||
piped through `tail`, which buffers, and the run was later killed — **so the caller's own log lines
|
||||
for the F press were lost.** The OUTCOME is solid (the box moved 2.5.0 → 2.5.1 and only the caller
|
||||
pressed it), but the per-pass log for F is gone. The second run below was written straight to a file
|
||||
for exactly this reason. Recorded rather than quietly omitted.
|
||||
|
||||
## Scenario G — two results, and the first one is the more interesting
|
||||
|
||||
### G-a. The deliberately broken edge was NEVER ATTEMPTED — the safety rule filtered it first
|
||||
|
||||
The C3-class negative control was `vikunja: 2.6.0 → alpine:3.20`. The caller **skipped it**, because
|
||||
`alpine:3.20` and `vikunja/vikunja:2.6.0` are different repositories and therefore cannot be ordered,
|
||||
so the edge is "across" and belongs to a human (`09` §3 decision 3, §3b Q3). The box was never asked
|
||||
to run it: `vikunja` ended the night still on 2.6.0, no update, no hold, untouched.
|
||||
|
||||
**That is a real finding for Slice 6 and it cuts both ways.** The within-a-major rule is the FIRST
|
||||
line of defence and it works — a catalog edge that cannot be ordered never reaches the guarded update
|
||||
unattended. But it also means **this shape of broken edge cannot be used to measure the unattended
|
||||
HOLD path**, because the rule that makes automatic updates safe is the same rule that refuses it.
|
||||
|
||||
### G-b. The no-retry property, PROVEN over three passes
|
||||
|
||||
Reverting the catalog (`f5f6a152b513`) left all four apps running something NEWER than the catalog —
|
||||
R-524's Ahead state. The caller then had a genuine TERMINAL refusal to react to:
|
||||
|
||||
pass 1/3 glance BEHIND and within a major — pressing Update
|
||||
glance REFUSED reason=downgrade TERMINAL — will not press again.
|
||||
„Ez a változat újabb a katalógusban lévőnél — visszalépés csak az
|
||||
üzemeltető kérésére."
|
||||
... the same for uptime-kuma, vikunja and wishlist ...
|
||||
pass 2/3 (nothing)
|
||||
pass 3/3 (nothing)
|
||||
summary outcomes={} never_again=['glance','uptime-kuma','vikunja','wishlist']
|
||||
|
||||
**Four apps, pressed exactly once each, then never again across two further passes.** Nothing on the
|
||||
box changed: no update ran, no hold was set, no journal was written.
|
||||
|
||||
**This is R-524 and R-609 working together end to end, unattended.** Before R-609 put `reason` on the
|
||||
wire the caller would have had only a Hungarian sentence to parse, and the only safe readings were
|
||||
"give up on everything" or "press for ever". The full log is `09-unattended-night.log`.
|
||||
|
||||
## What Slice 6's design now knows that it did not
|
||||
|
||||
| question | answer, from measurement |
|
||||
|---|---|
|
||||
| how long does one app take end to end, unattended? | 51 s – 1 m 26 s for these four small apps, including the health wait — see 04/05/07 |
|
||||
| does the caller need new controller code? | **No.** It presses the existing guarded Update and reads the existing state. |
|
||||
| can it tell "wait" from "never"? | **Yes, since v0.261.0** — and not before |
|
||||
| does the within-a-major rule hold? | **Yes, and it is the first thing that fires** — G-a |
|
||||
| what does a held app look like the next morning? | **STILL UNKNOWN unattended** — see below |
|
||||
|
||||
## NOT MEASURED, stated plainly
|
||||
|
||||
**The unattended HOLD path.** `09` §3b Q4 asks who is told when an automatic update ends HELD. This
|
||||
night never produced a hold, because the only failing edge available was one the safety rule
|
||||
correctly refused (G-a). Measuring it needs an edge that **passes** the within-a-major test and still
|
||||
fails its health check — same repository, same major version, a tag that starts and does not serve.
|
||||
Real images rarely offer one, so this probably needs a purpose-built image rather than a catalog
|
||||
move. **Q4 therefore still rests on the ATTENDED hold measured in slice 4 (v0.238.0, Scenario F),
|
||||
not on an unattended one.** Recorded as the gap it is.
|
||||
@@ -0,0 +1,167 @@
|
||||
#!/usr/bin/env python3
|
||||
"""unattended-caller.py — the Slice 6 spike: an app updates with NOBODY pressing anything.
|
||||
|
||||
THIS IS EVIDENCE, NOT PRODUCT. It lives under documentation/audits/ and nothing in the controller
|
||||
imports it. It exists to MEASURE the mechanism Slice 6 would need before that slice is designed, by
|
||||
pressing exactly the same guarded Update a person presses — `POST /api/stacks/<n>/update` — and
|
||||
nothing else. No new endpoint, no new controller code, no privileged path.
|
||||
|
||||
WHAT IT DOES NOT DO, deliberately: it does not decide policy. The window is simulated by running it;
|
||||
the per-app switch of `09` §3b Q2 does not exist yet; it never touches a box that is not the scratch
|
||||
guest. Run it from DooPlex against guest 9202 only.
|
||||
|
||||
THE TWO RULES IT EXISTS TO PROVE
|
||||
1. within a major, for EVERY compose service, or it does not press at all (§3 decision 3, §3b Q3);
|
||||
2. a refusal's REASON decides whether it ever presses again (R-609):
|
||||
TRANSIENT busy updating deploying migrating self_updating -> try next pass
|
||||
TERMINAL held downgrade -> never again
|
||||
FOR A HUMAN memory disk no_backup -> log and leave alone
|
||||
Before R-609 the body carried only a Hungarian sentence, so this distinction was unavailable to
|
||||
anything that is not a person — which is why a caller like this could not have been written.
|
||||
|
||||
Usage: python3 unattended-caller.py --passes 6 --every 300
|
||||
"""
|
||||
import argparse, json, re, subprocess, sys, time
|
||||
|
||||
CALL = "/tmp/ctl/c.sh" # the helper from 00-api-recipe.md
|
||||
TRANSIENT = {"busy", "updating", "deploying", "migrating", "self_updating"}
|
||||
TERMINAL = {"held", "downgrade"}
|
||||
FOR_A_HUMAN = {"memory", "disk", "no_backup"}
|
||||
TAG = re.compile(r"^v?(\d+)(?:\.(\d+))?(?:\.(\d+))?(.*)$")
|
||||
|
||||
|
||||
def log(msg):
|
||||
print("%s %s" % (time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()), msg), flush=True)
|
||||
|
||||
|
||||
def call(method, path, data=None):
|
||||
cmd = [CALL, method, path] + ([data] if data else [])
|
||||
out = subprocess.run(cmd, capture_output=True, text=True, timeout=180).stdout
|
||||
try:
|
||||
return json.loads(out)
|
||||
except Exception:
|
||||
return {"_raw": out}
|
||||
|
||||
|
||||
def split_ref(ref):
|
||||
"""(repo, tag) or (None, None) when the reference carries no plain tag."""
|
||||
if "@" in ref:
|
||||
return None, None
|
||||
i = ref.rfind(":")
|
||||
if i < 0 or "/" in ref[i + 1:]:
|
||||
return None, None
|
||||
return ref[:i], ref[i + 1:]
|
||||
|
||||
|
||||
def same_major(a, b):
|
||||
"""True only when BOTH tags are plain versions, share a suffix, and share a first number.
|
||||
|
||||
Mirrors stacks.CompareImageRefs deliberately: the box's own rule is the one under test, and a
|
||||
caller that judged 'within a major' differently would measure its own opinion instead.
|
||||
"""
|
||||
ra, ta = split_ref(a)
|
||||
rb, tb = split_ref(b)
|
||||
if ra is None or ra != rb:
|
||||
return False
|
||||
ma, mb = TAG.match(ta or ""), TAG.match(tb or "")
|
||||
if not ma or not mb:
|
||||
return False
|
||||
if ma.group(4) != mb.group(4): # the suffix must be IDENTICAL (…-apache vs …-apache)
|
||||
return False
|
||||
if ma.group(2) is None or mb.group(2) is None:
|
||||
return False # one component is a LINE, not a version
|
||||
return ma.group(1) == mb.group(1)
|
||||
|
||||
|
||||
def edge_is_within_a_major(st):
|
||||
"""ALL services must pass. One unorderable service makes the whole edge 'across' -> a human."""
|
||||
installed = {k: v["ref"] for k, v in (st.get("app_config", {}).get("installed_images") or {}).items()}
|
||||
catalog = st.get("catalog_images") or {}
|
||||
if not installed or not catalog or set(installed) != set(catalog):
|
||||
return False, "service sets differ or nothing recorded"
|
||||
for svc, want in catalog.items():
|
||||
got = installed[svc]
|
||||
if got == want:
|
||||
continue
|
||||
if not same_major(got, want):
|
||||
return False, "across: %s %s -> %s" % (svc, got, want)
|
||||
return True, "within a major"
|
||||
|
||||
|
||||
def follow(name, timeout=900):
|
||||
"""Watch one update to its end. Returns done | failed | held | timeout."""
|
||||
deadline = time.time() + timeout
|
||||
last = None
|
||||
while time.time() < deadline:
|
||||
st = call("GET", "/api/stacks/%s" % name)
|
||||
phase, updating = st.get("update_phase"), st.get("updating")
|
||||
if phase != last:
|
||||
log(" %s: phase=%s updating=%s" % (name, phase, updating))
|
||||
last = phase
|
||||
if not updating and phase in ("done", "failed"):
|
||||
held = bool(st.get("hold_reason"))
|
||||
return "held" if held else phase
|
||||
time.sleep(2)
|
||||
return "timeout"
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--passes", type=int, default=6)
|
||||
ap.add_argument("--every", type=int, default=300)
|
||||
args = ap.parse_args()
|
||||
|
||||
never_again, outcomes = set(), {}
|
||||
for p in range(1, args.passes + 1):
|
||||
log("=== pass %d/%d" % (p, args.passes))
|
||||
call("POST", "/api/sync")
|
||||
call("POST", "/api/stacks/rescan") # R-607: a sync can say 'no change' and still move
|
||||
stacks = call("GET", "/api/stacks")
|
||||
stacks = stacks if isinstance(stacks, list) else stacks.get("data", [])
|
||||
|
||||
for st in stacks:
|
||||
name = st.get("name")
|
||||
if not st.get("deployed") or st.get("protected") or name in never_again:
|
||||
continue
|
||||
installed = {k: v["ref"] for k, v in (st.get("app_config", {}).get("installed_images") or {}).items()}
|
||||
if not installed or installed == (st.get("catalog_images") or {}):
|
||||
continue # unknown, or level with the catalog
|
||||
ok, why = edge_is_within_a_major(st)
|
||||
if not ok:
|
||||
log(" %s SKIP — %s" % (name, why))
|
||||
continue
|
||||
|
||||
log(" %s BEHIND and within a major — pressing Update" % name)
|
||||
r = call("POST", "/api/stacks/%s/update" % name)
|
||||
if r.get("ok") is False:
|
||||
reason = (r.get("data") or {}).get("reason", "")
|
||||
sentence = r.get("error", "")
|
||||
if reason in TERMINAL:
|
||||
never_again.add(name)
|
||||
log(" %s REFUSED reason=%s TERMINAL — will not press again. %s" % (name, reason, sentence))
|
||||
elif reason in TRANSIENT:
|
||||
log(" %s REFUSED reason=%s transient — retry next pass. %s" % (name, reason, sentence))
|
||||
elif reason in FOR_A_HUMAN:
|
||||
never_again.add(name)
|
||||
log(" %s REFUSED reason=%s — needs a person. %s" % (name, reason, sentence))
|
||||
else:
|
||||
never_again.add(name)
|
||||
log(" %s REFUSED reason=%r UNKNOWN — stopping on it, safest reading. %s" % (name, reason, sentence))
|
||||
continue
|
||||
|
||||
started = time.time()
|
||||
end = follow(name)
|
||||
outcomes[name] = (end, round(time.time() - started, 1))
|
||||
log(" %s ENDED %s after %.1fs" % (name, end, time.time() - started))
|
||||
if end in ("held", "timeout"):
|
||||
never_again.add(name)
|
||||
|
||||
if p < args.passes:
|
||||
time.sleep(args.every)
|
||||
|
||||
log("=== summary outcomes=%s never_again=%s" % (outcomes, sorted(never_again)))
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -714,7 +714,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-516** | **[P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item **„Debug"**; (2) the dashboard CPU tile **„Load: 0.29 / 0.39 / 0.37"**; (3) the launcher tile **„Filebrowser"** opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) **Uptime Kuma 2.4 opens on „Which database would you like to use?"** (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), `wiki.DOMAIN` (R-498). **Fix shape:** rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed `db-config.json` for SQLite in the template (catalog). **Added by F4 (20:04:54Z):** (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. **Added by F7 (20:50–21:00Z, system disk at 95 %):** (10) a banner on every page in English, „**SSD disk usage high: 90%**"; (11) the dashboard tile reads „Rendszer (/) 61.8 GB / 68.7 GB (**90%**)" while `df` reports **95 %** (reserved blocks ignored), and „(/)" labels the data volume `/mnt/sys_drive`; the deploy page says nothing about free disk. **Added by the i18n spike (2026-09-17, controller v0.247.0):** (12) **six formal („ön") forms in the converted dashboard copy** — „Olvassa be telefonnal" (launcher QR hint), „Biztosan kikapcsolja a megosztást?" (launcher), „Ha újratelepíti … importálnia kell" (the layout's remove-app modal), „Kérjük, vegye fel a kapcsolatot" (backups empty state). They were NOT fixed — a localisation release may not change Hungarian bytes — and are now COUNTED by `controller/scripts/i18n_missing_gate.py` (`HU_FORMAL_CEILING` = 6, a ratchet: a new one convicts, fixing one lowers it). The inventory (`audits/I18N-INVENTORY-2026-09-17.md`) is the list this row closes against in localisation slice 6 (R-561). **Extended 2026-09-17 by slice 1 (controller v0.248.0–v0.250.0):** the converted copy now counts **16** formal forms (`HU_FORMAL_CEILING` 6 → 10 → 12 → 16, each raise stated, Hungarian unchanged by rule) — and the count is an UNDER-count: the gate's stem list sees 4 on the release C pages (storage, drive wizards, debug), while a wider list („írja be”, „adja meg”, „adjon hozzá”, „válassza ki”, „biztosan eltávolítja”, „engedélyezze”, „hozzon létre”, „kattintson”) finds at least **22** keys there alone (login „Adja meg a jelszavát”, the NAS guide, the storage confirms). Widen the stems when this row is worked; the ceiling rises with them. **NARROWED 2026-09-20 by localisation slice 6's walk, item by item** (`audits/i18n-slice6-2026-09-20/R-516-item-by-item.md`): item 2 FIXED (slice 1), items 3 and 6 APP-OWNED (and for an English household item 6 is an advantage), item 11 is a wrong NUMBER rather than a language and needs its own row, item 1 is correct English for an English reader and still open for a Hungarian one. **Items 4, 7, 8, 9 and 10 are NOT DECIDABLE FROM AN ENGLISH WALK** - they are about Hungarian copy quality, a disconnected drive and a full disk, none of which this walk had. **What this row is now waiting for is a HUNGARIAN walk on a box with a second drive**, not another English one; saying it closed would be the kind of closure that makes a register stop meaning anything. | **NARROWED 2026-09-20 - rank P3-LOW; owner: CC; needs a Hungarian walk** |
|
||||
| **R-518** | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — rank P2-MEDIUM; owner: CC (controller)** |
|
||||
| **R-519** | **[P2-MEDIUM] After a backup torn by a power cut, an app's restore point carries the new database dump's time while its files are from the previous run — and no customer screen says the run was interrupted.** MEASURED 2026-09-14 (BIGNIGHT F2, VM 333): „Mentés most" 19:40:03Z; the power was cut 19:40:08Z while adventurelog was stopped for its volume dump. On disk afterwards, `backups/primary/adventurelog`: `db-dumps/adventurelog-postgres.sql` **19:40:07**, `volume-dumps/*.tar` **19:00:35**, `manifest.json created_at 19:02:34Z`; bookstack the same shape (sql 19:40:07, tars 19:00:48). `GET /api/backup/snapshots` for both → `time 2026-09-14T19:40:07Z, helyi` (`phase5/F2/units-on-disk.txt`, `backup-honesty.txt`). `/backups/apps` shows „Utolsó adatbázis mentés 2026-09-14 21:40 · … OK" and every app „Utolsó: 5 perce"; `/backups` and `/dashboard` contain no word of an interruption (fragments `megszakad|sikertelen|nem sikerült` = 0; control „hiba" appears in the standing warning text). The controller itself knew: `[appstop] crash recovery: an app-data backup (volume dump) … was interrupted … restarting them: [adventurelog]` and pushed `backup_failed (error)` to the hub. A household restoring „the 21:40 backup" gets 21:00 files for BookStack's uploads. **Fix shape:** date a point by the oldest part it contains (or mark it partial) and show the interrupted run on the backups page until the next complete one. **F6 (drive unplugged 1 s into a backup, 20:37:59Z) adds three facts:** the run skipped four apps' volume dumps („Skipping volume dump for immich — drive disconnected", also jellyfin, nextcloud, paperless-ngx) and still reported `db_dump {"count":6, "success":true}`; nextcloud's point is dated 20:37:59Z (its SQL finished before the unplug) beside 19:01 volume tars; and a torn `immich-postgres.sql.tmp` (20:38:00) plus F3's `pre-restore-…-nextcloud-mariadb.sql.tmp` are left in the units on the drive. Immich's point correctly stayed at 19:02:34Z (the `.tmp` was not promoted). `/backups/apps` fragments `kihagy|sikertelen|részleges` = 0 (`phase5/F6/`). | **READY — rank P2-MEDIUM; owner: CC (controller)** |
|
||||
| **R-520** | **[P3-LOW] A power cut during a guarded Update leaves no record that an update was running, so whether the journal resumes or aborts honestly is unmeasured for a real version change.** MEASURED 2026-09-14 (BIGNIGHT F3, VM 333): the only Update the catalog allowed after the Phase 4 revert was a same-version one on nextcloud; power was cut 2 s after `update nextcloud: phase pulling` (safety dump written, pin advanced to the unchanged definition). After boot: no log line resumes, aborts or names the interrupted update; `app.yaml` `pinned_images` = `installed_images`, no hold, no verdict; page „Fut · Naprakész". Nothing wrong was produced — and nothing could have been, with identical images. **What it needs:** the same cut on a real bump (a throwaway app with a one-step catalog move on a scratch branch or the scratch guest), asserting the page's message and the pin after boot. **CLOSED 2026-09-21 — measured on a REAL bump, and it behaved.** Scratch guest 9202 (demo-hp), controller v0.260.0, throwaway `uptime-kuma`, real catalog move 2.4.0 -> 2.5.0 (reverted the same session). `pct stop 9202` at 11:04:54.439Z, 3 ms after the poller observed `update_phase: pulling`. **Honest limit on the instrument:** `pct stop` returned 3.8 s later, so the phase at the DECISION is observed and the phase at the freeze is inferred. **After boot the box said so itself — a POSITIVE observable, not an absence:** `update recovery: uptime-kuma was interrupted in pulling (started 2026-09-21T11:04:52Z) — nothing had run; putting the pin back`, then `pin uptime-kuma: …:2.4.0` and `update …: pin and definition PUT BACK to the pre-update version`. `pinned_images` and `installed_images` both 2.4.0, live compose line 2.4.0, no hold (correctly — nothing ran), `update_phase: failed`, app **running and healthy on 2.4.0** 31 s after boot, and the household told: „A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább." **THE INSTRUMENT TRAP THIS RUN FOUND IS WORTH MORE THAN THE RESULT.** The first post-crash read, taken from the host against the STOPPED guest, reported the update journal **ABSENT** — which would have made this a defect finding. It was a FALSE NEGATIVE: **`pct mount` maps the guest's ROOTFS ONLY and does not apply the guest's own internal mounts**, and 9202 keeps `/var/lib/docker` on its own ext4 mount under the `mp0` volume, so from the host that path is an EMPTY STUB. The bytes were at `/var/lib/lxc/9202/rootfs/var/lib/felhom/docker/volumes/felhom-controller-data/_data/data/`, and the box's own recovery read them at the next boot. Caught by running `find` over the whole rootfs instead of trusting one constructed path, and by demanding a positive control for the directory searched. **An empty directory is not evidence of an absent file.** **What is NOT measured and does not re-open this row:** the second cut, in the `starting` phase (Scenario F's territory — the migration may have run). Evidence: `audits/update-arc-2026-09-21/` 08-12. | **CLOSED 2026-09-21 — measured on a real version change** |
|
||||
| **R-520** | **[P3-LOW] A power cut during a guarded Update leaves no record that an update was running, so whether the journal resumes or aborts honestly is unmeasured for a real version change.** MEASURED 2026-09-14 (BIGNIGHT F3, VM 333): the only Update the catalog allowed after the Phase 4 revert was a same-version one on nextcloud; power was cut 2 s after `update nextcloud: phase pulling` (safety dump written, pin advanced to the unchanged definition). After boot: no log line resumes, aborts or names the interrupted update; `app.yaml` `pinned_images` = `installed_images`, no hold, no verdict; page „Fut · Naprakész". Nothing wrong was produced — and nothing could have been, with identical images. **What it needs:** the same cut on a real bump (a throwaway app with a one-step catalog move on a scratch branch or the scratch guest), asserting the page's message and the pin after boot. **CLOSED 2026-09-21 — measured on a REAL bump, and it behaved.** Scratch guest 9202 (demo-hp), controller v0.260.0, throwaway `uptime-kuma`, real catalog move 2.4.0 -> 2.5.0 (reverted the same session). `pct stop 9202` at 11:04:54.439Z, 3 ms after the poller observed `update_phase: pulling`. **Honest limit on the instrument:** `pct stop` returned 3.8 s later, so the phase at the DECISION is observed and the phase at the freeze is inferred. **After boot the box said so itself — a POSITIVE observable, not an absence:** `update recovery: uptime-kuma was interrupted in pulling (started 2026-09-21T11:04:52Z) — nothing had run; putting the pin back`, then `pin uptime-kuma: …:2.4.0` and `update …: pin and definition PUT BACK to the pre-update version`. `pinned_images` and `installed_images` both 2.4.0, live compose line 2.4.0, no hold (correctly — nothing ran), `update_phase: failed`, app **running and healthy on 2.4.0** 31 s after boot, and the household told: „A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább." **THE INSTRUMENT TRAP THIS RUN FOUND IS WORTH MORE THAN THE RESULT.** The first post-crash read, taken from the host against the STOPPED guest, reported the update journal **ABSENT** — which would have made this a defect finding. It was a FALSE NEGATIVE: **`pct mount` maps the guest's ROOTFS ONLY and does not apply the guest's own internal mounts**, and 9202 keeps `/var/lib/docker` on its own ext4 mount under the `mp0` volume, so from the host that path is an EMPTY STUB. The bytes were at `/var/lib/lxc/9202/rootfs/var/lib/felhom/docker/volumes/felhom-controller-data/_data/data/`, and the box's own recovery read them at the next boot. Caught by running `find` over the whole rootfs instead of trusting one constructed path, and by demanding a positive control for the directory searched. **An empty directory is not evidence of an absent file.** **What is NOT measured and does not re-open this row:** the second cut, in the `starting` phase (Scenario F's territory — the migration may have run). Evidence: `audits/update-arc-2026-09-21/` 08-12. **CORRECTED 2026-09-21 (same day): that sentence left the dangerous half of the question in prose and in no row — see R-610, which carries it and CLOSES it with three measurements.** The history above stands as written; only this pointer is added. | **CLOSED 2026-09-21 — measured on a real version change** |
|
||||
| **R-521** | **[P3-LOW] One unplugged drive sends the operator five e-mails and the household none.** MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): `storage_disconnected (error)` at 21:58:02 CEST plus `app_start_failed (warning)` for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five `Operator email sent`). The customer's mailbox (`tester1@felhom.eu`, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. **Fix shape:** suppress `app_start_failed` for apps stopped by a `storage_disconnected` (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. **F6, 40 min later, the opposite failure:** a second, separate drive loss (20:38:33Z) produced `storage_disconnected (error)` and four `app_start_failed`, and the hub logged `Operator email suppressed … cooldown` for all five — **no mail at all for the second unplug**; only `health_degraded (warning)` mailed. A per-key cooldown that outlives the recovery (`storage_reconnected` came between them) silences a new incident. **F7:** the system disk at 95 % produced only `health_degraded (warning)`, whose operator mail was **suppressed by the cooldown** left by F6's `health_degraded` 15 minutes earlier; no disk-specific event reached the hub at all — the operator was not told the disk was nearly full. | **READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy)** |
|
||||
| **R-522** | **[P3-LOW] While the box has no internet, the dashboard's „Cloudflare Tunnel" tile keeps saying „Fut", and no page tells the household the box is offline.** MEASURED 2026-09-14 (BIGNIGHT F8, VM 333): VM 333's traffic off the LAN and to the hub was dropped at demo-hp's bridge 21:08:36 → 21:26:07Z. Throughout, the LAN dashboard (probed every 26 s from demo-hp) answered 200 and, polled every 2 min, showed no banner and the tile „Cloudflare Tunnel — Biztonságos internetkapcsolat — a szerver portnyitás nélkül érhető el kívülről. · **Fut** · Védett"; meanwhile cloudflared logged ≈ 20 errors every 2 minutes, the public name answered 530, and the controller logged `[report] Push failed … context deadline exceeded` and `Job hub-report failed: hub push failed after 3 attempts`. The tile reports the container, not the connection. A household whose remote access is gone sees „Fut". **Fix shape:** the tile reads the tunnel's connection state (cloudflared's registered connections or the report push result) and says „Nincs internetkapcsolat" when either fails. | **READY — rank P3-LOW; owner: CC (controller)** |
|
||||
| **R-524** | **[P2-MEDIUM] When the catalog moves an app back to an older version, a box that already updated shows „Frissítés elérhető" — and the offered Update is a downgrade.** MEASURED 2026-09-15 (BIGNIGHT Phase 6, VM 333): privatebin was updated 2.0.5 → 2.0.6 through the guarded Update after the drill bump; the catalog was then reverted to 2.0.5 (`a161ccb`). At 22:13:37Z the box reads `installed privatebin/pdo:2.0.6`, `catalog privatebin/pdo:2.0.5`, `catalog_since 2026-09-14`, and the app page tag „**Frissítés elérhető — ma**" with the title „Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot." The label compares for *difference*, not for *newer* (`09-update-architecture.md` §5.4 render table); the guarded Update would advance the pin „to the catalog's current definition" — 2.0.6 → 2.0.5. The same state follows any real upstream yank. **Not pressed tonight.** **Fix shape:** compare versions (or `catalog_since` against the installed record) and render „Naprakész" / „a katalógusnál újabb" when the box is ahead; refuse a pin move to an older tag without an operator word. **CLOSED 2026-09-21 — controller v0.260.0.** `stacks.CatalogOrder` (`internal/stacks/updateorder.go`) replaces the three-way comparison with FOUR verdicts — Unknown / Current / Behind / **Ahead** — and **moves out of `web` so the badge and the refusal read ONE verdict**; `web.compareInstalledToTemplate` is now a thin wrapper. An app AHEAD reads „Naprakész" / "Up to date" with `tag-ok` (the same word and class as level — there is nothing for the household to do) and a title saying why (`badge.update.ahead.title`, both bundles). `Manager.UpdatePreflight` refuses with reason `downgrade`, HTTP 409, „Ez a változat újabb a katalógusban lévőnél — visszalépés csak az üzemeltető kérésére.", logged with both image maps. **The API now renders update refusals through `errText`**, so the new key is not a seam built and never wired. **Ahead is NARROW on purpose:** every differing service must be orderable AND newer, or the verdict falls back to Behind — this gate can BLOCK an update, so it errs towards letting one run. Ordering is `util.Version.Compare` (the house rule: one comparator) behind a tag normaliser — `X.Y`/`X.Y.Z`, optional leading `v`, two-part padded with `.0`, and a trailing suffix that must be IDENTICAL on both sides, so `nextcloud:31.0.14-apache → 31.0.15-apache` orders while `postgres:16-alpine`, `26.05.2-ls310 → -ls311`, `kimai/kimai2:apache-2.57.0`, a date stamp and a digest pin do not. **The suffix rule was found by the fixture, not by design** — the first implementation called every real catalog tag unorderable. **Three red-proofs, each SEEN to fail.** Recorded as `09` §3 decision 10 (decided by CC unattended — operator may reverse). | **CLOSED 2026-09-21 — controller v0.260.0** |
|
||||
@@ -781,6 +781,13 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-605** | **[P3-LOW] A catalog gate that REFUSED TO RUN and a gate that ran and could not decide print the same word, so a reader cannot tell which happened.** FOUND 2026-09-21 while answering why the chaos night's update round could not run. On 2026-09-17 `check-image-resolvable` and `check-volume-persistence` both returned INCONCLUSIVE and the drawn `update` action was replaced with `use` (`audits/DRILL-chaos-night-2026-09-17.md:692-695`). **Neither script is defective — they behaved exactly as designed**, and both headers say why: a detector that cannot prove itself must refuse to report rather than guess (`check-image-resolvable.py` cites the 2026-07-21 incident where a Docker Hub throttle read as 24 of 65 pins falsely dead). **What is missing is the DISTINCTION.** `check-volume-persistence.py`'s `self_test` refuses to evaluate ANY app when it cannot build its canary image — a harness-level refusal — while `classify()` returns a per-app UNDETERMINED for an app that wrote nothing; `check-image-resolvable.py` likewise separates a harness-level canary failure (exit 2 at `check()` L180-183) from a per-pin throttle (L121-128). **`catalog_gates.py`'s VERDICT map collapses all of them into one `INCONCLUSIVE` label**, so the operator-facing summary cannot say whether the gate ran at all. **The cost is real and already paid:** no raw stdout of the 2026-09-17 run survives in either evidence directory, so the exact triggering path is INFERRED from the code plus the documented throttle precedent, not observed — a distinct summary line would have recorded it for free. **Fix shape:** have each gate's exit distinguish "the harness refused" from "the result is undetermined" (a third exit code, or a marker line the runner matches), and have `catalog_gates.py` print the two differently. **Ships with a decoy each way (R-421): a run whose canary fails must NOT read as a per-app undetermined, and vice versa.** Small. | **READY — rank P3-LOW; owner: CC (catalog)** |
|
||||
| **R-606** | **[P2-MEDIUM] Every sentence the UPDATE path shows a household is Hungarian-only, and it lands on TWO pages — FOUND LIVE 2026-09-21, not by reading.** While measuring R-520 on guest 9202 at controller v0.260.0, the English app page rendered the badge in English — *"Update available — today"* — directly above *„A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább."* **The mixed line is worse than either language whole**, and this one is the household's only explanation for why their app did not move. **MECHANISM:** `Stack.UpdateError` is a finished Hungarian STRING, not a key. `Manager.finishUpdate` stores it (`internal/stacks/update.go` L517, 526, 530, 679, 902, 907, 936, 943) from the raw literals `MsgUpdateInterrupted`, `MsgUpdatePullFailed`, `MsgUpdateBackupFailFmt`, `MsgUpdateDumpFailFmt`, `MsgUpdatePinFailed`, `MsgUpdateJournalFailed`, `MsgUpdateBackupNoUnit`, `MsgUpdateHoldUnsaved` (`update.go` L79-96), and BOTH `app_info.html:36` and `stacks.html:99,104` render it verbatim. **SCOPE IS WIDER THAN THE EIGHT:** the same path carries `UpdatePhaseLabel` (`updatePhaseLabels`, L63) and the HOLD sentence returned by `UpdateGuards.HoldFor`, including `backup.Manager.UpdateCopyHolds`'s *„csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem"* — **which is a PROMISE ABOUT WHETHER THE CUSTOMER'S FILES COME BACK**, and ranks this with R-590 rather than below it. **NOT closed by v0.260.0:** that release routed the 409 REFUSAL through `errText` (so the pre-flight refusals reach an English household in English), but a refusal is the path where nothing happened — these are the sentences for when something DID. **Fix shape, and the pattern already exists in this repo:** `UpdateError` stores a KEY plus args, exactly as v0.259.0's `degradedMessageFor` was changed to return a key, and the page resolves it with `errText`/`msg` at render — the decision stays language-free in one place while the words are chosen by whoever knows the reader. `util.MsgErrorf` already carries key+args across that gap. Render test per sentence in both languages. **Every one of these is BORN AS A KEY territory, so `i18n_go_keys.json` accounting applies.** | **READY — rank P2-MEDIUM; owner: CC (controller)** |
|
||||
| **R-607** | **[P3-LOW] A forced catalog sync answers „nincs változás" while the cache DOES change, and `catalog_images` stays stale until a separate rescan — so the update badge can be wrong for a window nobody bounds.** MEASURED 2026-09-21 on scratch guest 9202 (controller v0.260.0) while proving R-524. A real catalog move was pushed, `POST /api/sync` was invoked, and it answered **„nincs változás"** — yet the box's own cache file `<data>/catalog-cache/templates/uptime-kuma/docker-compose.yml` **had moved to the new tag**. `Stack.CatalogImages` as served by the API stayed at the OLD value until a separate `POST /api/stacks/rescan`. **WHY IT MATTERS AND WHY IT IS NOT COSMETIC:** `CatalogImages` is the single input `stacks.CatalogOrder` compares against (v0.260.0), so for that window the badge answers from a stale catalog — it can read „Naprakész" on an app that IS behind, which is the exact failure §5.6 of `09` was written to prevent, arriving by a different door. **It also cost a measurement:** the session that found it nearly recorded a `tag-ok` badge as proof of the R-524 ahead arm when the badge was in fact stale; the honest reading came only after the rescan. **An instrument that can report an old value as a current one is not a measurement.** **TWO SEPARATE QUESTIONS, and the row does not conflate them:** (a) why the sync REPORTS no change when the working tree moved — a wrong sentence, possibly a comparison against the wrong ref; (b) whether `CatalogImages` is refreshed by the sync at all or only by `ScanStacks` on its own timer — if the latter, the staleness window is the scan interval and is bounded but unstated. **Neither was isolated** — this row records the observation, not a diagnosis. **Also observed in the same run, NOT diagnosed and folded in here rather than filed twice:** a removed app leaves `applied-compose.yml` behind in its stack directory. Stated as observed; it was not established whether that is intended. **Fix shape:** first reproduce with a loop that pushes a tag, syncs, and reads `catalog_images` on a timer, so the window is a NUMBER before anything is changed. Evidence: `audits/update-arc-2026-09-21/06-sync-after-bump.txt` and `16-r524-sync-box-ahead.txt`. | **READY — rank P3-LOW; owner: CC (controller)** |
|
||||
| **R-608** | **[P2-MEDIUM] The controller swaps ITSELF in the middle of a guarded app update, and 04:30 sits inside the window proposed for automatic app updates.** FOUND 2026-09-21 by reading the clock, not by a failure. The controller self-updates daily at `self_update.auto_update_time` — **default 04:30** (`config/config.go` L422, scheduled `cmd/controller/main.go` ~L1365) — and again from `MaybeAutoUpdate` after ANY hub report once a floor sits above the box, so at any hour. The swap restarts the controller container. `09` §3b **Q1** proposes **02:30–05:00** for automatic app updates. **It contains 04:30.** **MEASURED, and the gap was NARROWER than it first looked — which is why the fix is where it is:** the updater's only busy gate was `backupRunning` (`updater.go` L61, read at L487 dry-run, L512 `TriggerUpdate`, L660 `maybeAutoUpdate`), wired in `main.go` L659 to `backupMgr.IsRunning()`. The guarded update's **`backing-up` phase DOES take the backup single-flight** (`RunAppBackupNow` → `acquireRunning`, `backup/update_guard.go:333`), so that ONE phase was already covered. `checking`, `safety-dump`, `pinning`, `pulling`, `starting` and `verifying` were not — and the last two are exactly where the new version may already have touched the customer's data. The reverse was absent too: `UpdatePreflight` never asked whether a swap was running. **CLOSED 2026-09-21 — controller v0.261.0.** `stacks.Manager.AnyUpdating()` → `Updater.SetAppUpdatingCheck`, a deliberate sibling of `SetBackupRunningCheck` consulted in the SAME three places; `Updater.IsUpdateRunning` → `Manager.SetSelfUpdatingCheck`, and `UpdatePreflight` refuses `self_updating`. Both wired in `main.go`, the only place holding both objects — **`stacks` never imports `selfupdate`.** Two sentences born as bundle keys. **THE PROPERTY THAT MATTERS MOST IS THAT THE LOCK DOES NOT LATCH:** `Stack.Updating` is cleared on done, failed AND held, so a HELD app does not block the controller's own updates — including the release that might fix whatever held it. A latching gate would be a worse failure than the one prevented, and a silent one. Pinned by `TestR608_LockReleasesAfterHold`. Four red-proofs, each seen to fail. | **CLOSED 2026-09-21 — controller v0.261.0** |
|
||||
| **R-609** | **[P3-LOW] An update refusal has a machine-readable reason inside the process and none on the wire, so an unattended caller cannot tell "wait" from "never".** `UpdateRefusal.Reason` has existed since v0.237.0 (`busy`, `deploying`, `updating`, `held`, `migrating`, `memory`, `disk`, `no_backup`, `downgrade`, and `self_updating` since v0.261.0) and never left the process: the 409 body carried only the translated sentence. **The distinction is not decorative** — `busy`/`updating`/`deploying`/`migrating`/`self_updating` are TRANSIENT and `held`/`downgrade` are TERMINAL until a person acts. A caller that cannot tell them apart either gives up on a passing backup window or presses a terminally-refused button on every pass for ever. `09` §6.2's unattended caller reads exactly this. **CLOSED 2026-09-21 — controller v0.261.0.** The body gains `data: {"reason": "<Reason>"}`, ADDITIVELY; the sentence is unchanged and no page moves. Table-driven test over five reachable paths plus a control that a non-refusal carries none. **FOUND WHILE WRITING THE TEST, NOT BY READING — and it was the reason that matters most:** `actionStack` refuses a HELD app on **its own line, BEFORE `UpdatePreflight`** (`api/router.go` ~L601), so `held` would have been the one reason missing from the wire. That line now carries it too. Red-proof seen to fail. | **CLOSED 2026-09-21 — controller v0.261.0** |
|
||||
| **R-610** | **[P2-MEDIUM] The power cut AFTER the new version has started was never measured — R-520 closed on the safe half only.** R-520 (CLOSED 2026-09-21) cut in `pulling`, where **nothing had run**: the pin goes back and that is the easy case. Its own last lines said the `starting` cut was "NOT measured and does not re-open this row", and no row carried it. The dangerous case is the cut AFTER the new version has started and may already have migrated the customer's data — where `RecoverUpdates` (`stacks/update.go:909`) marks the app Updating and RESUMES rather than rolling back. **That behaviour was READ from the source and never observed.** **CLOSED 2026-09-21 — measured THREE times on guest 9202, controller v0.260.0**, with three different apps and two different cut mechanisms: vikunja 2.3.0→2.6.0 and uptime-kuma 2.4.0→2.5.0 by `pct stop` (a real power cut), and wishlist v0.66.0→v0.67.0 by restarting ONLY the controller container (exactly what a self-update does). **All three ended HONEST:** the recovery line appeared, the update resumed, each app came up on the NEW version, and in every case all FOUR version observables agreed — `pinned_images`, `installed_images`, the live compose `image:` line, and `docker inspect` of the running container (digests matched too). No hold, no stuck `Updating`, no surviving journal, no retry loop. The seeded data read back through each app's own front door for the two where a post-cut read-back was taken. **THE DANGEROUS CASE WAS GENUINELY EXERCISED, and the proof is a log line, not an assumption:** vikunja's own log shows `Ran all migrations successfully` and `Vikunja version v2.6.0` at **12:28:26.881 UTC — 0.64 s after the cut decision and ~0.4 s before the guest stopped answering.** The 2.6.0 schema migration had ALREADY been applied to the customer's SQLite database when the power went. Recovery resumed FORWARD, so old-binary-on-migrated-database never happened — **but this branch is one step from it: had the cut landed a second earlier, in `pinning` or `pulling`, the 2.3.0 pin would have been put back onto a 2.6.0 database.** That is not a defect today; it is the reason §4's "no automatic rollback" ruling is right, and it is now evidence rather than argument. **INSTRUMENT LIMIT, stated because it bounds the claim:** `starting` lasts well under a second on this box. Three attempts across two cut mechanisms (`pct stop` returning in 3.0–3.8 s; a controller restart in 1.7 s) ALL landed in `verifying`. No phase was faked. **`RecoverUpdates` handles `starting` and `verifying` in ONE branch, so all three runs exercise the same recovery arm** — the arm under test. A cut that lands inside `starting` itself remains unmeasured and would need an in-process fault injector. Evidence: `audits/update-arc-gaps-2026-09-21/` 04, 05, 07. | **CLOSED 2026-09-21 — measured three times** |
|
||||
| **R-611** | **[P3-LOW] A session reported "everything is done" over a phase it had silently skipped — the process failure, not the missing measurement.** The 2026-09-21 update-arc session's brief contained a Phase 5 spike: one app updated by the box with **nobody pressing anything**, once succeeding and once forced to fail. **It did not run, and nothing said so** — no evidence file, no code, and no sentence in `STATUS.md`, `UPDATE-ARC-STATE-2026-09-21.md`, either `REPORT.md` or `09`. `09` §6.2 was left describing Slice 6 "as it would be built" with no measurement under it, which reads like a considered design rather than an untested one. **Why this is a row and not a grumble:** the missing measurement was recoverable in an afternoon; the missing SENTENCE was not, because the next reader had no way to know it was missing. A skipped phase that is declared costs one line; a skipped phase that is not costs the next session its baseline. **CLOSED 2026-09-21 by the successor session**, which ran it (`audits/update-arc-gaps-2026-09-21/`, scenarios F and G) **and** adopted the standing habit that closes it generally: **the report's FIRST section is "not done", even when empty.** Every part and scenario of a brief is listed there if it was skipped, shortened or changed, with the reason. | **CLOSED 2026-09-21 — run by the successor session; "not done" is now the report's first section** |
|
||||
| **R-612** | **[P1-HIGH] `wishlist` cannot be signed up to on a fresh Felhom install, the deploy reports SUCCESS, and the error the customer sees is a LIE.** MEASURED 2026-09-21 on guest 9202 while seeding for the power-cut drill. The image's first-boot `pnpm prisma db seed` is **`Killed` — OOM at the catalog's `mem_limit: 128M`**. Without it the `Role` and `Group` rows are absent, so **every** signup fails. **The message the user is shown is `User with username or email already exists`** while the container log says the real cause: `FOREIGN KEY constraint violated`. A household would conclude the account already exists and try to recover a password that was never created. **The controller reports the app running and HEALTHY throughout, and the deploy reported successful** — so nothing on the box says anything is wrong. Repaired on the scratch guest only, to unblock seeding: memory raised to 512 M, the image's own seed re-run, memory put back to 128 M. **The catalog was NOT changed** — the fix is a memory-limit question for the catalog and is deliberately left to a session that can measure the real ceiling rather than guess it. **Needs: the actual peak RSS of that seed, then a `mem_limit` that clears it, plus a check that the seed's failure is not silent.** | **READY — rank P1-HIGH; owner: CC (catalog + a look at whether a failed first-boot seed can ever be visible)** |
|
||||
| **R-613** | **[P2-MEDIUM] `uptime-kuma` parks on its setup wizard with no login and no monitors, and the box tells the household it is HEALTHY.** MEASURED 2026-09-21 on guest 9202. On first boot uptime-kuma 2.4.0 sits at `SETUP-DATABASE` (`Waiting for user action...`) and its main socket.io server never starts. **The controller's `http :3001` probe sees the wizard's 302 and records the app as running and healthy.** So a monitoring app that cannot be logged into, and is monitoring nothing, is presented to the customer as fine. Passed through its own front door for the drill with `POST /setup-database {"dbConfig":{"type":"sqlite"}}`. **This is the health-check class the catalog skill already warns about — a probe that proves the PORT answers, not that the APP works** — and it is worth a row because the failure direction is a false GREEN, which no alarm will ever catch. **Needs: a healthcheck for this template that fails while the wizard is up** (the catalog `REUSE.md` maps the families), and a sweep for other templates whose probe would pass on a setup wizard. | **READY — rank P2-MEDIUM; owner: CC (catalog)** |
|
||||
| **R-614** | **[P3-LOW] A stale `update_phase` survives a remove and redeploy, so a freshly installed app can read "Frissitve" before it has ever been updated.** OBSERVED 2026-09-21 on guest 9202: a newly deployed `uptime-kuma` read `update_phase=done` / „Frissitve" before any update had been run against it — left in the manager's IN-MEMORY stack state by the previous session's update of a since-removed instance of the same name. It cleared on the next controller restart. **Why it matters beyond the cosmetic:** `update_phase` is one of the fields a person (and, after R-609, an unattended caller) reads to decide whether an update happened. A value that outlives the app it described is the same class as the R-166 "absent means unknown" family — a confident answer about something that no longer exists. **Fix shape:** clear the update fields when a stack is removed, beside wherever `Updating`/`updateHeld` are reset; a test that removes and redeploys and asserts the phase is empty. Small. | **READY — rank P3-LOW; owner: CC (controller)** |
|
||||
| **R-602** | **[P3-LOW] The language a signed-in page uses is NOT the language a cookie asks for, and a live probe that forgets this reports a fixed defect as unfixed.** FOUND 2026-09-21 verifying R-598 on guest 9201. `GET /backups` with `felhom_lang=en` returned the **Hungarian** page. That is correct — `langFor` step 2 says a request carrying a session reads the household's saved setting and deliberately ignores the visitor cookie, so a signed-in family never sees a language a previous visitor picked on the sign-in page of the same browser — but it means **the cookie is the right instrument for the anonymous claim/login/bind pages and the wrong one for every page behind auth**, where `?lang=` is. A session that had run only the cookie probe would have concluded R-598 was still open and fixed it a second time. **This is a documentation gap, not a code defect**, and it is the kind that costs a whole session: nothing in `10-localisation.md` §2.2 or in any runbook tells a prober which instrument to use where. **Fix shape:** four lines in `10-localisation.md` §2.2 — a table of surface → language instrument — and a pointer from the live-validation section of the workspace rules. Recorded meanwhile in `audits/i18n-closing-2026-09-21/live/backups-page.md`. | **READY — rank P3-LOW; owner: CC (docs)** |
|
||||
| **R-603** | **[P3-LOW] An English string containing an apostrophe silently never matches on a rendered page, and a `strings.Contains` assertion reads exactly like a missing sentence.** FOUND 2026-09-21 while writing the R-598 render tests. `backup.target.absent` was first written as *"The system backup's drive cannot be reached…"*; `html/template` escapes `'` to `'`, so the page carried the sentence and every assertion for it failed. **The failure mode is the expensive part:** the test said *"the English absent-drive copy never reached the page"*, which is indistinguishable from the handler not being wired — and the obvious next move is to go and re-fix the handler. Reworded to avoid the possessive, and all 23 new English values were then swept for `' " < > &` (zero). **The Hungarian bundle has never hit this** because Hungarian copy uses „quotes" and few apostrophes; **English copy will hit it again.** **Fix shape:** either a bundle gate that refuses an HTML-escapable character in a value destined for a page (and an allow-list for the ones that legitimately need one), or a test helper that compares against `html.EscapeString(want)` so the assertion cannot be fooled. The gate is the better shape — the helper only protects tests that remember to use it. | **READY — rank P3-LOW; owner: CC (controller)** |
|
||||
| **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured, and a class-A app carries one sentence saying where its files ARE protected. Proven on demo-hp through the page the customer opens: Paperless-ngx reads „1. mentés … DB + Konfig" with „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi …", while its „2. mentés" row still reads „DB + Konfig + Adatok". Red-proof: restoring the old app-shaped label fails `TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold`. **RE-PROVEN 2026-09-16 on a FRESH box** (installed from the built ISO 1.28.0, controller 0.244.0, off-site on by default): the Nextcloud row read „1. mentés … DB + Konfig" with the new sentence, „2. mentés … Nincs 2. (off-drive) másolat", „3. mentés Sikeres restic → …your-storagebox.de"; „DB + Konfig + Adatok" appeared ZERO times while the local unit held no file leg. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** |
|
||||
|
||||
Reference in New Issue
Block a user