Written BEFORE the soak so the comparison cannot be fitted afterwards. Sources are the box's own [scheduler] Registered lines and the agent's backup tier armed lines, read at 23:56. MUST happen: db-dump 02:30, tier2-backup + fill-watch 03:30, metrics-prune 04:00, offbox-backup 04:15 (snapshot_count 1 -> 2), a backup_run_digest event, reports every ~15m. MUST NOT happen: host_stale/host_recovered, offsite_selfheal_restaged (this is also R-218's negative control — a box WITH a target must not declare), offbox_repo_orphaned, offsite_repo_key_changed, escrow blob SERVED, any controller self-update. Genuinely uncertain and recorded as such: the restore-test has a 24h settle and every archive here is hours old, so NOT running is correct, not a gap.
37 KiB
CAMPAIGN 11 — Phases 2 and 4 journal
Every observable, in the order taken. Unattended session, started 2026-08-05 22:38 CEST.
Clocks, stated once because three are in play and they differ:
demo-hp and the appliance = CEST · the guest (LXC 9201) = UTC · the hub = CEST (the hub's own log
lines read 22:12:06 at a wall-clock of 22:48, so the hub logs local time, not UTC as the
project memory's "hub=UTC" note says — see the hygiene section).
Pre-flight — baselines re-read on arrival
| What | Value | How read |
|---|---|---|
felhom-controller main |
v0.201.0 @ 05cf352a2f17 |
git rev-parse HEAD == origin/main, tree clean |
felhom-agent main |
v0.125.0 @ 0404f60e6a7b |
same |
felhom.eu main |
@ 3a539ea5306a |
same |
| hub, LIVE | gitea.dooplex.hu/admin/felhom-hub:0.97.1 |
kubectl -n felhom-system get deploy hub -o jsonpath=… |
| Day-0 manifest | agent 0.125.0 (f7d8339b…7016b3) · golden 0.201.0 (e730d7ca…f007654) · min_agent 0.125.0 |
hub /configuration, the selected options |
| Highest register ID | R-223 | fresh grep -rhoE 'R-[0-9]{1,3}' --include=*.md over all four repos |
All four cited commits match the brief exactly. The manifest matches the journal's rebake record.
Venue, re-confirmed alive
| Host | demo-hp (HP t740), Tier 0. Reached by SSH key, first try — R-129 stands |
| VM | 321 c11-appliance, running, uptime of host 13 d |
| Appliance | 192.168.0.105, hostname c11, felhom-agent 0.125.0, guest 9201 running |
| Guest | 192.168.0.106 (DHCP — was .207 in Phase 0 and .227 in Phase 1; it moves) |
| Containers | felhom-controller:0.201.0 (healthy) · filebrowser (healthy) · traefik |
| Hub | c11-36d660 ONLINE, agent 0.125.0, 1/1 guests |
| Floor | c11 v0.200.0 (override); every other customer v0.156.0 — blast radius still zero |
| Untouched | drill-r50 (VM 300, stopped), guest 9201 on both demo boxes, every other hub customer |
calibre-web is NOT deployed. Phase 3 destroyed the guest and never redeployed it. The venue's
current shape is therefore a rebuilt box that has completed a second escrow ceremony, which is
exactly what Phase 2's F-series needs.
Harness — the access path, and two credential traps hit on the way in
Transport: DooPlex → (Tailscale) → demo-hp → 192.168.0.105 (appliance) → pct exec 9201, and
for HTTP, curl run on demo-hp, which shares 192.168.0.0/24 with the guest. No server logic is
skipped. Harness, not product. (The journal's Phase-0/1 SOCKS proxy was not needed; running the
client on demo-hp is the same transport with fewer moving parts.)
Two traps, both recorded because each cost a wrong reading first:
- Never
source ~/.config/credentials. It holds keys with hyphens (R_DEMO-HP=…) that bash cannot assign, and the resultingcommand not founderror echoes the secret value into the transcript. Two demo-box recovery codes were printed this way before the helper was rewritten togrepthe single key it needs. Nothing was written to a file and the values are not reproduced here — but the mechanism is a live I4 hazard for any session that sources that file. - Values in that file are quoted, so a bare
cut -d=keeps the quotes and yields the wrong secret — the hub returned302until the quotes were stripped. Already in project memory (credentials-file-values-are-quoted); re-confirmed.
§4 — the two positives that were owed
§4.1 — is the floor actually SERVED? MEASURED. YES.
The previous session recorded "no HELD line and no held reason" — two absences — and correctly refused to call that a measurement. Here is the positive.
The box's own view, rendered from GetFloor() (/settings, the „Rendszer" subpage):
Verzió és frissítés — Jelenlegi verzió 0.201.0 · Legújabb verzió 0.201.0 — naprakész · Utolsó ellenőrzés
2026-08-05T15:57:36Z· Automatikus frissítés – Minimális verzió (üzemeltető)0.200.0— a rendszer automatikusan erre a verzióra frissít, ha régebbi · Utolsó frissítés Sikeres (0.192.0 → 0.200.0)
Why that value is proof and not decoration, established from source rather than assumed:
data["ControllerFloor"] = s.updater.GetFloor()(internal/web/handlers.go:1454).u.floorhas exactly one writer —SetFloor(selfupdate/updater.go:556); the only non-test caller is the report-ACK handlerupdater.SetFloor(resp.MinControllerVersion)(cmd/controller/main.go:696).- Both hold branches of
ResolveManagedFloorsetd.Floor = ""(store.go:2086and:2095), andmanaged_floor_test.go:94pins it (!fd.Held || fd.Floor != "").
So a non-empty floor on the box can only have come from an ACK that carried one. A held floor is served as the empty string and would render as none.
The temporal half, which rules out a stale pre-hold value:
| time (CEST) | observable | source |
|---|---|---|
| 17:57:05 | controller starts | guest log 15:57:05Z |
| 17:57:11 | managed floor HELD for c11 … (controller floor withheld) |
hub log |
| 17:58:35 | [offsite-apply] settle-gate: GO — floor still unknown after 1m30s |
the box's own log |
| 17:57 → 22:12:06, every 15 min | managed floor HELD for c11 ×18 |
hub log |
| after 22:12:06 | no HELD line | hub log |
| 22:48 | box renders floor 0.200.0 |
/settings |
At start the box's floor was empty by its own account, and the hub withheld it for the next four
hours. The 0.200.0 it holds now can therefore only have arrived on an ACK after the vouch
released the hold. That is a measurement, not arithmetic.
The absence is backed by a liveness positive (rule 6): the HELD lines stop at 22:12:06 while the
hub demonstrably keeps logging — 22:44:41 host-report from demo-hp-bb76ea, 22:44:57 wgsync: pushed 5 peers. Two further report cycles (22:27, 22:42) passed with no HELD line.
And a correction to the brief's own plan for §4.1
The brief says to take this positive during F8's controller restart, "at DEBUG if necessary".
A restart would not have produced it. SetFloor's log line is u.dbg(...)
(updater.go:554), and u.dbg is gated on a private u.debug flag set from
cfg.Logging.Level == "debug" (main.go:478) — it writes to the logger, never to the logx debug
ring, so it cannot appear in /api/debug/logs at any level.
Verified live: the debug ring was pulled (4 000 entries, 17:03:08Z → 20:46:46Z — a window that
covers the 20:12→20:27 UTC release), and it contains no SetFloor line. Positive control
on the instrument first: the same ring carries 1196 DEBUG / 2802 INFO / 2 WARN entries, so the
ring works and the level is not filtered — the absence is structural, not evidential.
This is the same class the brief warns about, caught on the instrument rather than the subject: had the ring been read without the level census, "no
SetFloorline after a restart" would have looked exactly like "the floor is still held".
§4.1 verdict: the floor IS being served, measured from the box's own state. Recorded against R-216's row.
§4.2 — R-218's live half: deferred, with the reason
R-218's fix is present and readable in backup/offbox.go:1257:
func (m *Manager) needsOffsiteCredential(t *settings.OffboxTarget) bool {
if t != nil { return false } // a TARGET exists — not a rebuild
if m.settings == nil || !m.settings.GetHubEscrowIdentityPresent() { return false }
return true
}
The short-circuit on OffboxRepoPasswordHash() that caused the deadlock is gone — the declaration
now retires on the target, not on the key.
The live half cannot be taken in the venue's current state, and the reason is the fix itself: c11
has a target (host/user/port/repo_path all set, applied_marker present since 14:57), so
t != nil and the box correctly does not declare. Declaring here would be the bug.
Reaching the state that exercises it needs a box with the hub's identity blob, a placed key and no target — which the venue last held during Phase 1 and does not hold now. Recorded as still not measured, deliberately, rather than reported as measured from the unit test. F7's set-aside was examined as a candidate route and is reported in F7 below.
Phase 2 — the faults
Order: reversible first. Every fault carries a positive control proving the fault was real, and every control that failed is reported as a failed control rather than a result.
The four messages, as identified in internal/web/recovery_handlers.go (v0.201.0)
| trigger | first words | |
|---|---|---|
| M1 | the unseal failed and the hub holds no superseded package | „A megadott helyreállítási kódot nem fogadtuk el. Ellenőrizd, hogy mind a tíz szót…" |
| M2 | the capability gate says the agent cannot do it (R-216) | „Ez a gép még nem tudja megnyitni a mentéseidet… A kódoddal semmi baj…" |
| M3 | unlocked, inventory unreadable (R-217) | „A kulcs visszakerült, de a mentések listáját most nem sikerült beolvasni…" |
| M4 | the unseal failed and a superseded package is retained (R-222) | „Ez a kód nem nyitja meg azt a csomagot, amit most őrzünk ehhez a géphez…" |
The venue holds a superseded package (hub_escrow_superseded_present: true, …_at 2026-08-05 15:03:14), so the M4 branch — which is tested before M1 on the failure path — is live for every
failed unlock. That single fact shapes F1, F3 and F4 below.
F1 — a wrong recovery code, three times
Injected: a ten-word hyphenated code that is neither real code, POSTed three times.
| attempt | HTTP | elapsed | message |
|---|---|---|---|
| 1 | 200 | 1.194 s | M4 |
| 2 | 200 | 1.004 s | M4 |
| 3 | 200 | 1.014 s | M4 |
- Refused each time ✅ · no lockout ✅ (three identical responses; the handler documents
NO LOCKOUT (§8.4)by design) · HTTP 200, no error page ✅ - ~1.0–1.2 s is the positive tell that a REAL unseal was attempted —
age's scrypt cost, visible. Phase 1's 0.134 s is what "no attempt was made" looks like. This discriminator is used throughout. - Nothing was written ✅ — all four
/data/offboxfiles byte-identical with mtimes unchanged at 14:57:35:
846ac6e78a71e868 2026-08-05 14:57:35 applied_marker
75dbbc9be6909912 2026-08-05 14:57:35 known_hosts
e11a6c542b73477a 2026-08-05 14:57:35 repo_password ← the Phase-3 key, unchanged
85579eb0033832a0 2026-08-05 14:57:35 ssh_key
- I4 held ✅ — a sweep for three fragments of the entered code across the controller's data dir,
/var/log,/tmpanddocker logsreturned product hits: 0. Both controls passed first: a planted canary was found by the same sweep, anddocker logs … | grep -c recoveryreturned 6. The only hits were in the harness's own script (/tmp/.cc.sh) — a self-match, cleaned up.
The deviation from the brief's expected answer: message 4, not message 1. F1's stated right answer is "message 1 only". On this box M1 is unreachable, because the M4 branch precedes it and fires whenever a superseded package exists.
Judged on the message rather than the outcome: M4 is hedged („Ha egy korábbi kódot adtál meg"), it does not accuse the customer of mistyping, and it states two true facts. But a customer who simply fumbled their current code is given no hint to re-check their ten words and is instead routed to customer support about older backups. M1 — the only message that tells a customer to check their typing — cannot be reached on any box that has re-escrowed, which is exactly the population most likely to be typing a recovery code. Recorded as a finding.
F2 — the customer has no code at all — PASS
Two halves, both taken.
- The screen states it plainly, unprompted:
„A helyreállítási kódot senki nem tudja pótolni — sem a Felhom, sem az ügyfélszolgálat, sem az üzemeltető. Ez szándékos… Ha a kód elveszett, a korábbi mentések nem nyithatók meg többé."
- It offers the set-aside path rather than a dead end — „…nem kérem vissza a korábbi adatokat"
is rendered, because
CanSetAside(=OffboxOrphaned()) is true here. ✅ - An empty submission returns „Add meg a helyreállítási kódot." in 0.027 s — no unseal attempted, and no blame attached to a code that was never given. ✅
F3 — the hub unreachable when the code is entered — FAIL
Injected on the appliance: iptables -I OUTPUT 1 -d 37.191.56.193 -j REJECT.
Control: hub http=302 before → http=000, exit=7 after. The fault is real.
Then the correct current recovery code was entered, so the only possible cause of failure is the hub.
HTTP 200 · elapsed 0.0556 s ← far below the ~1.0 s scrypt cost: NO unseal was attempted
M4 — „Ez a kód nem nyitja meg azt a csomagot, amit most őrzünk ehhez a géphez. Ha egy korábbi kódot adtál meg: a géped azóta új mentési kulcsot kapott…"
The code was perfect and current. The hub was down. The customer is told their code does not open their package and is sent to customer support about older backups. The hub is never named. I6 breach — an unreachable service reported as a fact about the code.
And the machine knew. The agent's own journal distinguishes the two cases exactly:
F1 (wrong code): err="escrow: the recovery code did not unwrap the identity escrow
(wrong recovery code, or a corrupt blob): exit status 1"
F3 (hub down): err="escrow: fetching the sealed bundle: hub: transport error:
Get \"https://hub.felhom.eu/api/v1/hosts/c11-36d660/escrow\":
dial tcp 37.191.56.193:443: connect: no route to host"
The discriminator exists in the agent's err field and is discarded at the HTTP boundary: both
return HTTP 400 with one merged sentence — "the recovery code did not open the sealed bundle,
or the bundle could not be fetched" — and the agent's own msg= collapses them too
("FAILED (wrong recovery code, or the blob could not be fetched)"). The controller then has no
branch for "could not reach", so rerr != nil falls into the code/package messages.
Reverted: rule deleted, chain empty, hub http=302 again.
F4 — the agent stopped when the code is entered — FAIL
Injected: systemctl stop felhom-agent. Control: is-active = active → inactive, and no
listener on :8443. Real.
HTTP 200 · elapsed 0.0299 s ← again no unseal
M4 — the same sentence as F1 and F3.
The stated right answer is M2 — "names the machine's inability to ask, the R-216 fix under pressure". M2 did not fire. The controller log names the truth the customer never sees:
recovery: unlock failed: agentapi: POST /escrow/recover-offsite-password:
Post "https://169.254.253.1:8443/escrow/recover-offsite-password":
dial tcp 169.254.253.1:8443: connect: connection refused
Why the R-216 gate did not catch it, measured rather than reasoned — the box's own debug ring:
[web] recovery capability gate: offsite_key_recovery=yes (source=version)
source=version. recoverySupport asks netFeatures.SupportsWithSource(...), which answers from
the known agent version without probe traffic — by design, and correct for the question it was
built to answer. But the question it is being used for here is "can this machine ask right now?",
and a version number cannot answer that. The gate discriminates age, not reachability; a dead
agent of the right version sails through it, and the failure that follows is attributed to the code —
the precise outcome R-216's comment says "must never" happen ("An attempt that cannot succeed must
never be made, because its failure is attributed to the code").
F3 and F4 are one defect with two faces: on the unlock failure path there is no branch for the machine could not ask or could not reach, so every non-code cause is rendered as a statement about the code or the package. This is Phase 1's headline finding, relocated from the version channel to the transport.
Reverted: agent active, :8443 listening, felhom-agent 0.125.0.
F5 — the store unreachable after a successful unlock — PASS
The first attempt's control failed and is reported as a failed control: the probe used nc, which
the controller container does not have, so blocked and unblocked both printed the same fallback. A
second attempt used the IPv6 address getent hosts returned — and the container has no IPv6 route
at all, so that too could not discriminate. The container resolves the store to IPv4
91.98.242.176, and only a /dev/tcp probe on that address separated the states:
BEFORE: TCP-OPEN INJECT: iptables REJECT (OUTPUT + FORWARD) AFTER: TCP-CLOSED
Only then was the fault real. With the correct current code:
HTTP 200 · elapsed 1.198 s ← a REAL unseal
„A kulcs visszakerült, de a mentések listáját most nem sikerült beolvasni. A mentéseid nincsenek veszélyben — nézd meg a Biztonsági mentés oldalt néhány perc múlva."
- The unlock still counts ✅ · M3, the "could not be read" shape ✅
- R-217's false claim is genuinely absent ✅ — verified in Python on the local file with UTF-8 matching, after the accented-substring trap was accounted for (the project's own documented false- zero hazard):
MUST BE ABSENT "A tároló megnyílt" 0 · "van benne tartalom" 0 · "nem tudtuk alkalmazásokhoz rendelni" 0
CONTROLS PRESENT "A kulcs visszakerült" 1 · "nézd meg a Biztonsági mentés oldalt" 1
R-217's fix holds under exactly the fault it was written for.
Reverted: rules deleted, TCP-OPEN again.
F6 — „Most nem", then return — PASS
Baseline first (the interception is on /launcher and /dashboard, not /):
BEFORE GET / -> 302 /launcher
GET /launcher -> 302 /recovery ← the full-page interruption
GET /dashboard -> 302 /recovery
GET /recovery -> 200
POST /recovery/postpone → 302 → /launcher, and:
AFTER GET /launcher -> 200 ← no longer interrupted
GET /dashboard -> 200
GET /recovery -> 200 ← still reachable, permanently
The backups-area entry point SURVIVES ✅ — /backups/remote still carries exactly one /recovery
link and the sealed-package sentence, before and after:
„A korábbi, házon kívüli mentéseid visszaszerezhetők… Adatok visszaszerzése"
That asymmetry — full page silenced, entry point permanent — is what the handler documents, and it holds. And the unlock still works after postponing (measured: HTTP 200, 3.214 s, a real unseal).
F8 — the controller restarted mid-unlock — state PASS, message FAIL
Injected: an unlock started, then docker restart felhom-controller at T+0.7 s, inside the
~1.1 s unseal window. Control: the container's StartedAt moved 2026-08-05T15:57:04Z →
2026-08-05T21:06:05Z. The restart really landed mid-request.
What the customer got:
HTTP 502 · elapsed 1.914 s · body: "Bad Gateway"
- No half-state ✅ — after the restart the four
/data/offboxfiles are byte-identical with mtimes still 14:57:35, andsettings.jsondiffers from the pre-F8 snapshot in exactly one field,postponed: True, which F6 set. The controller returned healthy in 40 s. - „the screen says which" — it does not ❌. The customer is shown
Bad Gateway: a raw upstream error, in English, from traefik. It names no reason, offers no action, and says nothing about whether the key was installed. I3 breach — "every refusal names a reason a person can act on, in Hungarian, with no raw error."
What F8 does NOT establish, stated so the pass is not overclaimed. On this venue the unlock resolves to
RecoverUnchanged(the recovered password is byte-identical to the one already on disk), soInjectOffboxPassword— the only writing branch — was never reached. F8 therefore shows that a mid-request restart leaves state coherent and returns a raw 502; it does not establish that the password write is atomic. Testing that needs shape (a), which the venue does not hold. Recorded as a gap, not as a pass.
§4.1 again — and this time the box says it out loud
The brief expected the floor positive here, and F8's restart delivered a better one than /settings:
a cold-started process stating the floor it was served, ten seconds in. The same log line, same
box, same code, before and after the vouch:
15:58:35 [offsite-apply] settle-gate: GO — floor still unknown after 1m30s ← hold IN FORCE
...
21:06:06 [offsite-apply] settle-gate: awaiting floor knowledge (first report ACK)
21:06:16 [offsite-apply] settle-gate: GO — at/above floor 0.200.0 (we are 0.201.0),
no managed update running ← hold RELEASED
§4.1 is therefore measured twice, independently: the rendered GetFloor() value, and this
box-side line from a fresh process that had no prior in-memory state to inherit.
The finding F6 turned up on the way past — the store reports 0 when it cannot read
Not one of the eleven; found while checking F6's "the listing is coherent" clause on
/backups/remote. It is reported because a suspicion checked and confirmed is a result.
The page renders, on one screen:
Tároló méret · 0 pillanatkép Tárhelykeret: 0 / 50 GB (0%)
and, immediately below:
„A távoli tároló másik kulccsal készült mentéseket tartalmaz. A távoli tárhelyen lévő mentések egy korábbi, már nem elérhető kulccsal készültek… A meglévő mentések nem sérültek…"
Ground truth, measured directly against the Storage Box over SFTP — a read-only listing, no decryption, using the box's own transport credential:
/home/felhom-repo/snapshots:
-rw------- 253 Aug 5 13:14 f3d9cd67d539c00359f0454ea7a782f45beac8406c55275bee5a6886afa8d791
/home/felhom-repo: du -s → 12535 (KB)
/home/felhom-repo/keys: exactly ONE key
That is snapshot f3d9cd67 — the Phase 0 snapshot holding all three sentinels, byte-size and
timestamp matching the journal's repo_size_bytes 12 611 522 and its 13:13:32Z capture.
One snapshot and ~12.5 MB of the customer's only surviving copy, rendered as
0 pillanatképand0 / 50 GB, on the page the customer opens to find out whether their backups still exist — directly above a paragraph telling them the store contains backups.
Mechanism, from the box's own state rather than inferred. After the Phase 3 rebuild the
controller's settings.json offbox block carries no snapshot_count and no repo_size_bytes
key at all (the full block is quoted in the pre-flight section above). The values are unknown,
and unknown renders as the zero value.
This is R-217's defect class in a second location: a field whose zero value is indistinguishable
from a real measurement, defaulted past on a failure/unknown path. OffsiteInventory.Empty exists
precisely because "len(Apps)==0 is also what a failed read looks like" — the same reasoning applies
to a snapshot count the box has never successfully taken. I6: an absence reported as a fact.
I5 checked and NOT breached — the hub's /offsite page shows Campaign 11 · 0.0 GB · 50 GB, but
that is 12.5 MB rounded to one decimal of a GB, and the pool total (Used 3.8 GB) is consistent. The
hub and the box do not disagree; both understate, for different reasons. Only the box's
snapshot count — an integer, not a rounding — is false.
F11 — the box offline for a whole reporting window — PASS
Injected: qm shutdown 321 at 23:07:45, restored 23:29:12 (≈21½ min). Control: qm status →
stopped, and the last report before it was 23:06:12.
Both directions fired, and each sent an operator mail:
23:28:11 Host staleness: c11-36d660 ok → stale (host_stale) + Operator email SENT
23:29:59 Received report from c11 (2514 bytes) ← the box returns, unaided
23:30:11 Host staleness: c11-36d660 stale → ok (host_recovered) + Operator email SENT
23:30:11 Recovery host_recovered for c11: customer mail skipped — no unanswered customer down
mail (pairing miss)
- The dead-man's switch fires and clears itself.
ONLINE → STALE → ONLINE, hub-side, with no operator action. - The customer mail was correctly withheld — pairing-gated, exactly as the power-outage audit records. The gate held under a real outage.
- I5 holds: after the return the box's
offboxblock and the hub's view agree (nothing changed across the outage, and both say so). Judged after a full report cycle rather than from a single read — the Phase 0 lesson.
What was NOT reached, stated rather than implied: the box never went to DOWN (>1 h) — the
outage was ended once STALE and the recovery mail had both fired, because Phase 4 needed the venue
back. The STALE → DOWN escalation is untested.
F7 — „I do not want the old data", then a change of mind — the set-aside PASSES, the afterwards does NOT
Pre-state, measured on the Storage Box (read-only SFTP listing, no decryption):
/home felhom-repo (mtime Aug 5 13:13) ← exactly one repo
/felhom-repo/snapshots f3d9cd67d539c003…afa8d791 ← the Phase 0 history, 3 sentinels
/felhom-repo du -s → 12535 KB
Confirmation 1 (/recovery?setaside=1) states all four consequences before anything happens:
„a korábbi mentéseket félretesszük — nem töröljük; a helyreállítási kód nélkül többé nem lesznek megnyithatók; a gép új, üres mentési tárolót kezd; ez az oldal többé nem jelenik meg."
Confirmation 2 (POST /backup/offbox/reset confirm=1) → 302, flash
„Új távoli mentés indítása folyamatban — a régi előzmény félretéve (nem törölve)."
[WARN] [offbox] resetting orphaned repo (operator-confirmed (claimed)):
move-aside /home/felhom-repo -> /home/felhom-repo.orphaned-20260805, then re-init
NOTHING WAS DELETED — the assertion, verified against the remote:
/home felhom-repo (mtime 21:32) ← fresh, re-inited
felhom-repo.orphaned-20260805 (mtime 13:13) ← UNCHANGED mtime
/felhom-repo.orphaned-20260805/snapshots f3d9cd67… ← the Phase 0 history, still there
/felhom-repo.orphaned-20260805 du -s → 12535 KB ← byte-for-byte the pre-state
/felhom-repo/snapshots (empty)
Set aside, never deleted. PASS, and the two confirmations are honest about what they will do.
The change of mind — and this is where it fails
The customer returns wanting the old data. What they find:
GET /recovery -> 302 /backups/remote ← the screen is GONE
GET /launcher -> 200 ← no interruption
POST /recovery/unlock -> 302, 0.028 s ← the handler's gate; NO unseal, NO message
The brief's expected answer is "message 4, not a typing accusation". There is no message at all — which is not a typing accusation, but is also not an explanation.
And the set-aside history is invisible. The box knows exactly where it is:
"orphaned_renamed_to": "/home/felhom-repo.orphaned-20260805"
/backups/remote after the set-aside contains zero occurrences of orphaned-20260805,
„félretéve", „régi előzmény" or any equivalent (instrument controls: felhom-repo → 2,
„letétbe helyezve" → 1, so the page and the matcher both work). And a census of the whole web layer:
grep -rn "OrphanedRenamedTo" internal/web/templates/ internal/web/*.go → (no hits)
The field is written and read by nobody. 12.5 MB of the customer's retained data sits on the Storage Box at a path the box records and never shows. The only mention it ever gets is a flash message on the redirect, which is gone on the next click.
This is the project's own recurring "seam built but never wired" pattern, and it lands on the one promise the set-aside screen makes — we set it aside, we do not delete it. → R-228
(Correct and recorded as a pass: the recovery screen retiring itself is exactly what confirmation 1 promised — „ez az oldal többé nem jelenik meg" — so its disappearance is declared, not sprung.)
F9 — a box that never had off-site backups — the gate PROVEN LIVE, on a narrower shape
F9 as literally specified needs a box that never had off-site backups. The venue cannot be that box without being rebuilt, which the brief forbids before Phase 4 — so the literal precondition was not staged.
But the assertion that FAILED in Phase 1 was tested and it now passes. Phase 1's failure was not
about the precondition; it was that GET /recovery never asked the predicate and rendered its
story regardless (R-215). F7's set-aside drives recoveryOffer() false (a repository password exists
and the tier is no longer orphaned), and:
GET /recovery -> 302 /backups/remote ← the page consults the predicate
POST /recovery/unlock -> 302, no unseal ← the POST sibling was already gated
R-215's fix is proven live. What is NOT covered: the GetHubEscrowIdentityPresent()==false arm
(a box the hub holds nothing for). Recorded as partial, not claimed as full.
F10 — an app with a mandatory data path missing — COULD NOT BE INJECTED. Three attempts, each with its control.
Setup, through the customer's own endpoints:
- R-220 reproduced live, a third time —
GET /api/disks/candidatesreturnedinitialize: [], attach: []while both drives sat mounted at both/mnt/felhom-drives/<name>and the raw/mnt/<name>. Its red-proof reproduced too: unmounting only the raw mounts flipped the scan to both drives withfstype: ext4, data_bearing: true. (Deviation: an operator action on the appliance host, unavoidable while R-220 is open, and it doubles as the red-proof.) - Both drives re-enrolled via
POST /api/storage/attach→{"registered":true,"where":"/mnt/felhom-drives/adatok"}. calibre-webdeployed viaPOST /api/stacks/calibre-web/deploy(202, the same 409→202 flip Phase 1 recorded), healthy in 39 s, bound at/mnt/felhom-drives/adatok/userdata/media/books— its declared-mandatory path.- Baseline off-site run:
last_status: ok, 1 snapshot, logbacked up calibre-web (…, 1 mandatory path(s)).
A clean confirmation of R-225's mechanism, taken here: after this first successful run
settings.jsongained"snapshot_count": 1and"repo_size_bytes": 27371. Those keys were absent before it — which is exactly why the page rendered0 pillanatkép · 0 GB.
The three injection attempts:
| # | method | control | why it failed |
|---|---|---|---|
| 1 | mv the directory aside, app running |
ABSENT confirmed | the app recreated it (mtime 21:45) before capture |
| 2 | docker stop calibre-web, then rm -rf |
ABSENT confirmed, held 25 s | the controller's monitor auto-restarted the stack; the app recreated it |
| 3 | stop the stack through the controller API, then rm -rf |
ABSENT confirmed, held 25 s | the directory reappeared at 21:51:38.65 — coincident with the run's own start (21:51:38) — and the stack was found Up again a minute later |
All three runs reported ok with 1 mandatory path(s), and the R-203 stat-gap WARN
(mandatory data path missing on disk, skipped from offsite) never fired — correctly, because by
the time os.Stat ran the path existed again.
Verdict: harness, not product. F10 is recorded as NOT INJECTED, because the state it describes is not reachable on a deployed app of this kind: the app (or the monitor that keeps it up) recreates its own data directory within a second, every time.
But one observation is worth keeping, with its evidence. At capture time the directory held:
metadata.db metadata.db-shm metadata.db-wal ← recreated, 21:51
(no F10-SENTINEL.txt — the customer's file, written at 21:39)
The run reported ok. The verdict is about the presence of a declared path, not its content —
which is correct as designed and could hardly be otherwise — but it means an ok off-site run can
immediately follow the loss of everything that path contained. Not filed as a defect; recorded
because "the run said ok" is exactly the reassurance R-203 was filed to make trustworthy, and this is
its edge.
(The customer's sentinel was restored into the live library afterwards and the harness copy removed.)
Venue state at the end of Phase 2 — WORKING
| Containers | calibre-web healthy · filebrowser healthy · felhom-controller:0.201.0 healthy · traefik up |
| Drives | both enrolled; backup target {"degraded":false,"known":true,"label":"mentes","target":"felhom-backup"} — the I5/I6 baseline restored |
| Off-site | fresh repo, last_status: ok, last_success 2026-08-05T21:52:43Z, 1 snapshot, 26.4 KB, escrow_state: escrowed |
| Set aside | /home/felhom-repo.orphaned-20260805 — the Phase 0 history, 12 535 KB, untouched |
| Recovery screen | retired (correct — the set-aside declared it would be) |
| Deviations left in place | the raw /mnt/adatok and /mnt/mentes mounts remain unmounted (R-220's workaround); the stable /mnt/felhom-drives/* mounts are what everything uses |
Phase 4 — leave it alone and watch
Window: 2026-08-05 23:56 CEST → 2026-08-06 ~04:45 CEST. The venue is not touched: no endpoint is driven, no service restarted, nothing configured. Only logs and the hub are read.
Pre-registered expectations — written BEFORE the window, so the comparison is not fitted afterwards
A soak that reports only "nothing broke" has measured nothing. Both directions are named here in
advance. Sources: the box's own [scheduler] Registered … lines and the agent's own backup tier armed lines, read at 23:56.
The box's own schedule (controller, restarted 23:29:53 CEST after F11)
| when | job |
|---|---|
| 02:30 | db-dump |
| 03:30 | tier2-backup · fill-watch |
| 04:00 | metrics-prune |
| 04:15 | offbox-backup ← the one this campaign cares about |
| every 15 m | hub-report |
| every 1 m / 5 m / 2 m | agent-channel-health · system-health · stack-scan |
| every 10 s / 30 s | status-refresh, health-probes · deadapp-check, ring-spill |
| every 6 h | selfupdate-check · disk-health-check |
The agent's schedule (host domain)
backup tier armed target=felhom-backup cadence=24h keep_last=3 primary=true
backup tier armed target=felhom-pbs cadence=168h keep_last=0 primary=false
backup: restore-test scheduler starting (per-archive due-check) eval_interval=6h settle=24h
MUST happen (an absence here is a finding)
- Reports keep arriving every ~15 min; the host stays ONLINE all night.
db-dumpat 02:30,tier2-backup+fill-watchat 03:30,metrics-pruneat 04:00.offbox-backupat 04:15 runs unprompted and succeeds —snapshot_count1 → 2,last_status: ok,last_successstamped in the 04:15 window.- A
backup_run_digestevent reaches the hub after a backup run (this is the "operator digest"; it is event-driven per run, not a nightly cron — checked in the hub source rather than assumed). - The agent's
felhom-backuptier (24 h cadence) takes a whole-guest archive when it comes due.
MUST NOT happen (firing here is equally a finding)
- No
host_stale/host_recovered— the box is up and staying up. - No
offsite_selfheal_restaged— the box has a target, soneedsOffsiteCredentialreturns false and it must not declare. (This is also §4.2's negative control: if the box DOES declare tonight with a target present, R-218's fix is wrong in the other direction.) - No
offbox_repo_orphaned— the repo is fresh and written under the key the box holds. - No
offsite_repo_key_changed— nothing re-escrows tonight. - No
escrow blob SERVED— nobody is driving a recovery. - No controller self-update —
auto-update: false, and the box is at 0.201.0 with floor 0.200.0.
Genuinely uncertain, recorded as open questions rather than predictions
- The restore-test evaluates every 6 h with a 24 h settle, and every archive on this box is hours old. Expected NOT to run — and if it does not, that is correct behaviour, not a gap. The distinction matters because "no restore-test fired" would otherwise read as a failure.
- Whether the agent's local tier fires tonight depends on when its 24 h cadence last elapsed on a box built today.