Files
felhom.eu/REPORT.md
T
admin f5a4fceeeb
gates / gates (push) Successful in 16s
DRILL 2026-08-21: the off-site restore never replays named volumes (R-354..R-365)
Diagnostic only — no code changed, no version bumped, nothing deployed.

The verdict is a mixture. The unit and the off-site snapshot HOLD the data, proven
by identity in both storage classes including two Hungarian accented filenames. The
loss is in the last leg: ReconstituteFromOffsite skips every isUnit placement and the
volume tars live inside the unit, so the off-site full restore has no named-volume
leg at all — while the local restore-from-unit does, and returned the same tar
byte-identical minutes later.

Twelve rows opened, ceiling R-353 -> R-365. Three HIGH:
  R-354 off-site restore never replays volume dumps
  R-355 paperless-ngx's Postgres is dumped under a non-existent stack, so its unit
        has no DB dump, no safety dump is taken, and the customer is told it has none
  R-356 the off-site restore refuses for all 40 no-drive apps saying the running app
        "is not installed", with a remedy those apps make impossible

R-353's instruction (2) is satisfied and annotated: the 40-class DOES reach the
off-site tier. Its instruction (1) stands and is now larger. R-329 confirmed still
live and now the only bad-severity emit fleet-wide.

Evidence: documentation/audits/DRILL-backup-truth-2026-08-21/evidence/
2026-08-21 23:30:27 +02:00

30 KiB
Raw Blame History

REPORT — DRILL: does the backup hold the data, and does the restore tell the truth? (2026-08-21 night)

Unattended diagnostic drill on demo-hp. No code changed in any repository. No version bumped, no image built, nothing deployed. Evidence: documentation/audits/DRILL-backup-truth-2026-08-21/evidence/.


1. THE VERDICT

It is a mixture, and the drill's three options are all present — but they belong to different faults, and only one of them explains the thing you actually saw.

What explains YOUR observation (OpenGist, 2026-08-21 afternoon): THE BACKUP IS EMPTY, and then THE MESSAGE LIED

Reproduced independently tonight, and it agrees with what R-353 already recorded:

  1. The off-site restore for OpenGist never ran. It refused, because OpenGist declares no data drive, and the refusal says „a(z) opengist nincs telepítve" — "OpenGist is not installed" — about an app that was installed, deployed, running and healthy.
  2. The person therefore used the local restore-from-unit. The local unit on the freshly rebuilt box was genuinely empty of data — no dump cycle had run yet on a one-hour-old machine — so it held compose/ and nothing else.
  3. The restore returned that configuration and reported a bare completion.

So for that specific event the data was not in the thing that was restored. R-353 called this correctly and I did not find an error in it. I nearly filed a correction against it and was wrong to think so; its text is more careful than the CHANGELOG's summary of it.

What the drill found that nobody had seen: THE RESTORE LOSES IT

This is new, it is worse, and it is not the same fault:

When the off-site snapshot DOES hold the data, the off-site full restore still does not return it. The off-site restore has a files leg and a database leg. It has no named-volume leg at all.

Proven live on calibre-web at 22:23:36 with planted files:

leg in the unit in the off-site snapshot in the checking folder returned by the off-site restore
declared user files (media/books) n/a yes, 5/5 byte-identical yes, 5/5 byte-identical yes, 5/5 byte-identical
named volume calibre_web_config yes, 1 422 848 B yes yes NO — silently skipped

and the customer was told:

„A(z) calibre-web: 5 fájl visszaállítva (mentés: 2026-08-21 22:16) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa."

Five files came back. A 1.4 MB tar of the app's own configuration volume did not, and the sentence does not mention it. For calibre-web the lost leg is the app's settings. For the 40 catalogue apps that declare no data drive, that leg is the entire dataset.

Why: ReconstituteFromOffsite skips every placement flagged isUnit (controller/internal/backup/offbox_reconstitute.go:341-346), and the volume tars live inside the unit. grep for a volume-restore call in the whole off-site path returns nothing. The local restore does have one (restore.go:99 restoreDockerVolumes) — proven tonight by restoring PrivateBin's planted 1 MB from its volume tar, byte-identical. Two code paths, the same tar, one of them replays it.

And a third, separate: THE BACKUP IS EMPTY — really empty — for paperless-ngx's database

paperless-ngx runs a 72-table PostgreSQL. Its recovery unit records db_dumps: null. It always has. The dump is taken — 284 617 bytes, valid, 72 tables — and written to /mnt/sys_drive/felhom-data/backups/primary/**paperless**/db-dumps/, a directory named after a stack that does not exist, on the wrong drive. Nothing off-sites it. Nothing restores it. And because the safety-dump code filters on the same wrong name, the destructive restore takes no undo at all and then says:

„A(z) paperless-ngx: 0 fájl visszaállítva … Ennek az alkalmazásnak nincs adatbázisa."

The controller had dumped that database five minutes earlier.

One app in 53 is affected (catalogue-wide sweep in §6). It is the document archive.


2. THE FULL/EMPTY CONTRAST

Apps chosen, and why. From the two storage classes: calibre-web declares a data drive (needs_hdd: true, backup.userdata: media/books class: mandatory) and privatebin / opengist declare none (the 40-class; data lives in a Docker named volume). I ran three cases rather than two, deliberately: FULL and EMPTY alone cannot separate "it was empty" from "the class is broken", so privatebin was run FULL as the disambiguator.

The fixture — 5 files, two with Hungarian accented names, recorded as raw bytes:

SENTINEL.txt                 0c23c8531214fe20cbc1ed177da22f51d65570c5aaedb39e380aa8d4d3e62991
binary-1mb.bin               725763bbe679b22d5c231a14083ff13155f475c028c22ade4623955b50a2a84a
nested/őszibarack.md         a39ad6f623da67ac72f2e62a24245eef46c722c279ae89cd6b0c7c45afa6a4c4
plain.txt                    07e91a985809fc96752f97cddcf6523b211bc607c81a5808befabe35573926f1
árvíztűrő-tükörfúrógép.txt   0d6a22ec56acf61b8a80e73eb12cb223d607acaa3684acfcfa7c3e34ecdfabea

name bytes (UTF-8 NFC):
  árvíztűrő-tükörfúrógép.txt = c3a1 72 76 c3ad 7a 74 c5b1 72 c591 2d 74 c3bc 6b c3b6 72 66 c3ba 72 c3b3 67 c3a9 70 2e 74 78 74
  őszibarack.md              = c591 73 7a 69 62 61 72 61 63 6b 2e 6d 64

The comparator was proved able to convict before it was trusted. One byte flipped at offset 500 000 of binary-1mb.bin (af → 00): sha256sum -c reported binary-1mb.bin: FAILED, rc=1, while the other four passed; the unmodified set passed rc=0. The mutant was discarded.

The results

case class data leg in unit in off-site snapshot checking folder off-site restore local restore
calibre-web FULL drive user files n/a 5/5 identical 5/5 identical 5/5 identical —
calibre-web FULL drive named volume 1.42 MB yes yes yes NOT restored —
privatebin FULL no drive named volume 1.06 MB yes 5/5 identical 5/5 identical REFUSED — false reason 5/5 identical
opengist EMPTY no drive named volume 181 KB skeleton yes yes yes REFUSED — false reason —

The contrast decides it. The unit and the off-site snapshot hold the data, verified by identity, in both classes, accented filenames included. So the capture is sound. The failure is entirely in the last leg.

Both messages, verbatim:

  • FULL, drive class → ok=true, „A(z) calibre-web: 5 fájl visszaállítva (mentés: 2026-08-21 22:16) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa."
  • FULL and EMPTY, no-drive class → ok=false, „a(z) privatebin nincs telepítve, ezért nincs hová visszaállítani az adatait. A mentése szerint az adatai itt voltak: /mnt/sys_drive. Telepítsd újra az alkalmazást (Alkalmazások) ugyanerre a helyre…"

The second is the important one. FULL and EMPTY got the identical sentence, so the message cannot distinguish them — but the sentence is worse than uninformative, it is false twice over: the app is installed, and the instruction ("reinstall it to the same place") cannot be followed, because a 40-class app is offered no storage field at deploy time (that is R-352's own measurement).

This closes R-353's second instruction. It asked for proof that a named-volume app reaches the off-site tier rather than inference from gate order. It does. Manifests read tonight:

privatebin     volume_dumps = ['privatebin_privatebin_data.tar']       db_dumps = None
opengist       volume_dumps = ['opengist_opengist_data.tar']           db_dumps = None
kimai          volume_dumps = ['kimai_kimai_db_data.tar', 'kimai_kimai_var.tar']
                                                                       db_dumps = ['kimai-mariadb.sql']
calibre-web    volume_dumps = ['calibre-web_calibre_web_config.tar']   db_dumps = None
paperless-ngx  volume_dumps = [3 tars]                                 db_dumps = None   ← the defect

3. PART 0 — did the floor move the machine?

Yes, unaided, in 17 seconds.

moved from → to 0.216.0 → 0.217.0
who initiated the hub — the operator saved the floor at 19:48:37Z; a poke reached the agent from 10.77.0.1:58093 at 19:48:32Z for the artifact-manifest save 6 s earlier. No customer action, no agent-side decision.
how long floor saved 19:48:37Z → "controller-swap: new controller healthy" 19:48:54Z = 17 s. Swap requested 19:48:41Z → healthy = 13 s.
the container's own tag on the box gitea.dooplex.hu/admin/felhom-controller:0.217.0, created 2026-08-21 19:48:45 UTC — read from docker ps in guest 9201, not from the hub.

The hub's own host page for demo-hp shows the guest's Controller column as „—" — the hub does not know which controller version the box runs, while the controller_updated event it received says exactly that. Two hub surfaces, one blind.

demo-felhom also runs 0.217.0, but it started it at 19:31:04Z — 17 minutes before the floor was saved — and emitted controller_started with no controller_updated. It was moved by hand during the golden bake, not by the floor.


4. PART 2 — the four answers, from source

1. What puts a dump into a backup unit, and when? Which apps qualify?

  • Database dumps — runDBDumps (backup/backup.go:444-560). Qualification is a running container whose image matches a database image, mapped to a stack by deriveStackName (appbackup/dbdump.go:770). Written to <nsRoot>/backups/primary/<stack>/db-dumps/<stack>-<type>.sql.
  • Volume dumps — runVolumeDumps (backup/backup.go:607). Qualification is a deployed, non-protected stack with at least one Docker named volume (GetDockerVolumes). An app with zero named volumes is skipped silently and is never stopped.
  • When — one cycle, both legs, scheduled db-dump daily at 02:30 CEST; and again as the coherence pre-phase of every off-site run (offbox.go:938-951), which is what makes a snapshot an internally coherent {DB@T, files@T} pair.
  • CaptureRecoveryUnit only enumerates what is already on disk (recovery_unit.go:131-132). It writes no dump itself.
  • The gap this leaves: an app whose data is a bind mount and which has no database container produces neither leg. Its unit is configuration only — and nothing anywhere says so.

2. What should a correct backup contain?

  • Declares a data drive (13 of 53): the recovery unit (compose + app.yaml carrying the portable secrets + whatever dumps exist) plus the paths its .felhom.yml backup: block marks mandatory, appended to the restic snapshot as extra paths (offbox_capture.go:32).
  • Declares none (40 of 53): unit only. offboxCaptureSet returns (nil, nil) when the app has no backup block, so the off-site snapshot is the unit and nothing else. That is correct given that their data is inside the unit's volume tars — and tonight confirmed it is.

3. What does the restore report, and on what evidence? — THE ANSWER IS: FROM THE ABSENCE OF AN ERROR.

ReconstituteFromOffsite returns res and nil. Nothing in it ever asks whether anything was restored. The handler then calls EndRestoreOp(**true**, reconstituteOutcomeMsg(...)) (web/offbox_handlers.go:455). reconstituteOutcomeMsg (web/offbox_handlers.go:463-475) formats counters:

if res.DBsReplayed == 0 {
    return fmt.Sprintf("A(z) %s: %d fájl visszaállítva%s — az alkalmazás újraindult. "+
        "Ennek az alkalmazásnak nincs adatbázisa.", app, res.FilesPlaced, when)
}

So with FilesPlaced == 0 and DBsReplayed == 0 the customer reads "0 files restored — the application restarted" under a success. This is a finding on its own and is recorded whichever way the rest goes. Two aggravations found on top of it:

  • "This application has no database" is asserted from a counter, not from a fact. reimportDBDumpsFrom returns (0, nil) when the dump directory is absent or holds no .sql (restore_db.go:30-46), so an app that certainly has a database is told it has none. Proven live on paperless-ngx.
  • The file count is rsync's transfer count, not a restore count. A correct restore of unchanged data reports "0 fájl visszaállítva" — indistinguishable from a restore that did nothing. Observed: the same app reported 5, then 2, then 0 across three runs.

4. What did the 9 August off-site snapshot contain, and is it still readable?

Readable, and it contained the data. Three snapshots at 2026-08-09 08:30:

id app contents
41c830db calibre-web unit + userdata/media/books incl. the 9-Aug rehearsal fixture (árvíztűrő-tükörfúrógép.txt, őszibarack.md, binary-3mb.bin, plain.txt)
9e38b84c opengist unit + volume-dumps/opengist_opengist_data.tar, 182 272 B — manifest records volume_dumps: ["opengist_opengist_data.tar"]
78b93f04 privatebin unit + volume-dumps/privatebin_privatebin_data.tar, 2 560 B

All still readable tonight; restic check over the whole repository reports "no errors were found". So the 9 August off-site copy of OpenGist's data exists and is intact — it simply was not the thing the afternoon's restore read, and the off-site route that would have read it refuses for that class of app.

The repository had also been dead since 9 August and nothing said so: no snapshot between 2026-08-09 08:30 and tonight. The cause is visible in the hub event at 16:01 — the off-box target was lost by the guest rebuild (R-193's shape) — and, after tonight's self-heal restored the target, every per-app off-site toggle was still off, so the first run I triggered reported:

[offbox] backup run started (0 app(s) toggled)
[offbox] backup OK: 0 app(s) backed up, 18 snapshot(s), 14s

Credit where it is due: the card does not lie about this. It reads „Aktív — nincs kijelölt alkalmazás" and „Sikeres — nincs mentésre jelölt alkalmazás" beside the green tick. The tick still leads, and the log line alone says only „backup OK".


5. THE TWO CYCLES COMPARED

(filled in after the scheduled run — see §11)


6. PART 4 — everything attempted, and the message judged

# test outcome message judged
1a destructive restore: nothing is ever deleted PASS both ways. A file created after the snapshot survived; a file mutated after the snapshot was overwritten back to the snapshot's content. count is rsync transfers, not files restored — see §4.3
1b safety dump taken and verified before the stop PASS. 21:02:47Z dump written → 21:02:47Z StopStack romm. correct: „0 fájl és az adatbázis visszaállítva"
1b …and the whole operation refuses if it cannot be PASS. Made impossible by putting a regular file at the db-dumps path. Refused; plain.txt kept its post-snapshot mutation; the container's StartedAt was unchanged — the app was never stopped. honest and names the path, but leaks a raw Go mkdir … not a directory into a customer surface
1b the hole in it FAIL. For paperless-ngx the discovery resolves the database to the wrong stack, so hasDB is false: no safety dump is taken and the refusal cannot fire. The undo is absent, not refused. „Ennek az alkalmazásnak nincs adatbázisa" — false
2 end of the abandonment countdown state created and overdue; fires at 05:10 CEST — see §11 card states a past date in the future tense while overdue
3 damaged store MIXED — see below honest at the point of failure, then forgotten
4 drive pulled mid-restore restore failed, nothing written to the wrong place, the agent re-bound the drive within 5 s (23:15:29 pulled → 23:15:34 re-bound) WRONG DIAGNOSIS. „restore dir: mkdir …: permission denied" for a drive that had vanished. A person reads that and goes looking at permissions.
5 full disk, non-destructive path PASS — refuses before it starts. exemplary: „Nincs elég szabad hely a visszaállításhoz (183.6 MB szükséges, 99.2 MB szabad)." Both numbers named.
5 full disk, destructive path FAIL — no gate at all. offbox_reconstitute.go contains zero references to offboxFree; the three headroom gates are all on non-destructive paths (offbox_restore.go:231,297,423). It stopped the app, failed halfway on ENOSPC, left DRILL-2026-08-21/ holding 2 of 5 entries, and restarted the app. honest (No space left on device (28)) but raw rsync output
6 controller killed mid-restore PASS. SIGKILL inside the stop→restore→start window. On restart: "[appstop] crash recovery: an off-site restore (op "offbox-reconstitute:paperless-ngx") was interrupted and left 1 app(s) stopped — restarting them". App restarted, marker cleared, and the hub was told (event 3009). good — names the operation, not just the app
7 the 40-class under pressure the backup noticed; the alarm did not. Filled the 69 GB filesystem that holds the Docker data-root, the system namespace and all 40-class data, to 99% / 1.2 GiB free. All 12 containers stayed healthy. The reserve refused per app: „App backup REFUSED for kimai (size) … reserve: 97% used or 1.0 GiB free", and the hub got recovery_unit_capture_failed (error) naming the filesystem and its numbers. The fill watcher said nothing — it runs once a day at 03:30 plus once at startup (cmd/controller/main.go:1092). the refusal messages are good; the silence is the problem

Where the clock stopped me: nothing in Part 4 was skipped for time. Item 2's terminal deletion is scheduled rather than forced, because the sweep has no on-demand entry point — it is a daily job only.

4.3 in detail — the damaged store

One byte flipped inside pack 967853d2… at offset 5 000 000, over the repository's own SFTP transport. Pack files are named by their content hash, so this is genuine corruption.

  • restic check detects it — "ciphertext verification failed", "Fatal: repository contains errors".
  • But nothing in the product ever runs it. The only restic verbs in the entire controller are restore, snapshots, backup, unlock, stats, init, forget, prune, cat. The agent's restore-test is PBS-tier only. The off-site store is never verified by any layer, at any time. Corruption is discovered at restore time — the worst possible moment.
  • A restore that touches the damage fails honestly: ok=false, naming the file and "ciphertext verification failed".
  • But the failure is not remembered. It left a partial checking folder — 78 MB, 54 files, 15 of 16 originals. OffboxFullScratchReady (offbox_restore.go:305) asks only "does the directory exist and is it non-empty", so the wizard then offered all three actions including „Teljes visszaállítás indítása".
  • Pressing it ran the destructive restore from that known-incomplete copy and reported SUCCESS.

The repository was repaired from byte-identical originals; restic check now reports "no errors were found".


7. PART 5 — the two rows that were observed and never filed

5.1 — the delete guard on verification copies is blind, and its own comment says otherwise. offboxVerifyCopyDeleteHandler (web/offbox_handlers.go:502) gates on backupMgr.IsRunning() — the concurrency flag — while its comment states "It refuses while a backup/restore op is running: the copy being deleted could be the one currently being written." R-351b moved all seven restore handlers onto restoreOpBlocked() (which reads both flags); this handler was left behind, and one other site (:239) reads the bare flag correctly and documents why.

Reachability is not a race — it is the whole operation. RestoreOffboxScratch (offbox_restore.go:211) never calls acquireRunning at all, so IsRunning() is false for the entire duration of an off-site verification restore. Demonstrated live, flags read immediately before and after the delete:

--- BEFORE delete 22:35:25   display= True offbox-restore kimai   concurrency= False
--- DELETE calibre-web verification copy:  „Az ellenőrző másolat törölve…"
--- AFTER  delete 22:35:25   display= True offbox-restore kimai   concurrency= False

The copy was removed. The guard is app-agnostic, so the same call naming the restoring app hits the directory the restore is writing into. Rank: MEDIUM — it needs a customer to press delete during a restore, but both controls live on the same page, the window is the whole restore, and the target is the restore's own source. (Reconstitute and place do hold the flag, so the exposure is the verification-restore window only.)

5.2 — accented-text search is an instrument that fails silently, and it nearly did again tonight. Filed as an instrument defect. Occurrences I can evidence:

  1. 2026-07-20 — an accented grep through ssh → pct exec → bash -c nearly produced a wrong "banner cleared" claim. Recorded in felhom-controller/.claude/rules/ui-hungarian.md:19-22.
  2. 2026-08-13 — kubectl exec … sh -c "grep '<accented>'" returned 0 for three strings that were present, one step from being reported as a failed hub v0.105.0 deploy.
  3. 2026-08-21, tonight, 22:07 — tar -tf rendered őszibarack.md as \305\221szibarack.md (octal escaping). Recording the accented filenames' raw bytes from that listing would have been wrong. Caught by extracting the archive and reading the names with xxd.

Correction to the task's premise: that is two inside two weeks, plus the founding case a month earlier. I looked for a third inside the two-week window and did not find one on record.

The smallest guard I would propose — and did NOT build: the problem is not grep, it is that every one of these tools silently transforms the bytes. So the guard is not "use ASCII fragments" (a discipline, which is what failed three times) but a negative control that the harness cannot skip: any search whose pattern contains a byte ≥ 0x80 must be run twice — once for the target and once for a string that MUST be absent — and a zero result from the first is only reportable when the second also returns zero and a third probe for a known-present ASCII anchor returns non-zero. Three probes, one helper, no judgement required at the call site. Everything else has been tried and is what "nearly" means in all three cases.


8. RANKED REGISTER ROWS OPENED

Ceiling was R-353; it moved to R-365.

id rank what
R-354 HIGH The off-site full restore has no named-volume leg. The tar is in the unit, in the snapshot and in the checking folder, and is never replayed; the outcome reports success. For the 40-class this is the entire dataset. offbox_reconstitute.go:341-346.
R-355 HIGH paperless-ngx's PostgreSQL is dumped to a directory for a non-existent stack (…/primary/paperless/), so its unit records db_dumps: null, nothing off-sites it, no safety dump is taken on a destructive restore, and the customer is told the app has no database. appbackup/dbdump.go:770-798. One app in 53.
R-356 HIGH The off-site restore refuses for all 40 no-drive apps with „nincs telepítve" about an installed, running app, and instructs the customer to reinstall it "to the same place" — an instruction those apps' deploy page makes impossible. offbox_reconstitute.go:208-227.
R-357 MEDIUM The destructive restore has no headroom gate (the three that exist are all on non-destructive paths). It stops the app, fails halfway on ENOSPC and leaves a partially-restored data directory.
R-358 MEDIUM A failed scratch restore leaves a partial copy that OffboxFullScratchReady reports as ready; the destructive restore then runs from it and reports success.
R-359 MEDIUM The off-site restic store is never verified by any layer — restic check is not among the verbs the controller runs, and the agent's restore-test is PBS-only.
R-360 MEDIUM Verification-copy delete gates on IsRunning(), which RestoreOffboxScratch never holds — deletable throughout a restore. Its comment asserts the opposite. (Part 5.1)
R-361 MEDIUM The safety dump overwrites the unit's own DB dump: DumpOne writes the canonical <stack>-<type>.sql and only then renames it away. The comment at offbox_reconstitute.go:147-148 states it "can never overwrite the app's real dump". Proven: romm's romm-mariadb.sql was present before and absent after.
R-362 MEDIUM A data drive detached mid-restore is reported as „permission denied". The restore path never consults drive state.
R-363 MEDIUM The fill watcher runs once a day (03:30) plus at startup. A filesystem that fills at 03:31 is unannounced for ~24 h — while the backup reserve is already refusing apps.
R-364 LOW Accented-text search is a silently-transforming instrument; discipline has failed at least three times. Guard proposed, not built. (Part 5.2)
R-365 LOW An overdue abandonment countdown renders its past due-date in the future tense („…2026-08-20 napján véglegesen töröljük" shown on 2026-08-21).

Confirmed still live, not re-filed: R-329 — app_start_failed emits severity "warn" (internal/notify/notifier.go:546), outside the hub's vocabulary, so it coerces to info and emails nobody. Observed tonight as event 3006, severity info. A sweep of every emit( call site shows this is now the only remaining instance fleet-wide.

R-353: its instruction (2) is satisfied — see §2. Its instruction (1) stands and is now strictly larger than when it was written, because R-354 shows the bare completion can also be reported over a unit that did have a data leg.


9. WALL CLOCKS, AND EVERY STEP OFF THE CUSTOMER'S PATH

time (CEST) what
21:57 start; baselines
22:00 break-glass into demo-hp
22:09 fixture planted, comparator control passed
22:12 manual local backup
22:13 manual off-site run — 0 apps toggled
22:17 off-site run with 3 apps enabled
22:19–22:20 checking-folder restores
22:21 reconstitute privatebin → refused
22:23 reconstitute calibre-web → the conviction
22:25 local restore-from-unit privatebin → data returned
22:27 Part 4.1a — no-delete invariant
22:39–22:45 paperless-ngx → R-355
22:51–22:57 Part 4.3 damaged store; repo repaired
23:02–23:04 Part 4.1b safety dump, both directions
23:10–23:12 Part 4.5 full disk; Part 4.6 controller killed
23:15 Part 4.4 drive pulled
23:17–00:16 Part 4.7 filesystem filled and freed
23:35 abandonment countdown created (fires 05:10)

Steps off the customer's path, named:

  1. Break-glass root access to demo-hp via the hub-vaulted host_recovery credential — the box had lost DooPlex's SSH key (its authorized_keys held only its own root@demo-hp RSA key) and its tailnet address was unreachable. DooPlex's public key was re-added to /root/.ssh/authorized_keys and an ssh alias hp → 192.168.0.104 was added to ~/.ssh/config on DooPlex.
  2. A hub DB snapshot (hub.db + -wal + -shm) was streamed to the scratchpad to read host_recovery and the events table. It holds every host's secret; it is in the session scratchpad only and is not in any committed file.
  3. settings.json was edited directly (controller stopped, backup at /root/settings.json.drill-backup) to create the overdue abandonment state. There is no product path to shorten a countdown, and the real orphan→reset path would have destroyed demo-hp's entire off-site history.
  4. The set-aside store the sweep will delete was created by hand at u629488-sub3:/home/felhom-repo-superseded-drill-20260821, for the same reason.
  5. Two restic pack files were deliberately corrupted and then restored from byte-identical copies.
  6. fallocate fillers were used to fill two filesystems and were removed.
  7. Apps deployed for the drill: opengist (re-deployed empty), kimai, paperless-ngx, romm.

10. THE FENCES

  • ep0 / the off-site endpoint — untouched outside this machine's own path, and here is how I know. demo-hp authenticates as the Hetzner Storage Box sub-account u629488-sub3, which is chrooted to its own home: ls / returns Permission denied, and /home contains exactly .ssh and felhom-repo. Every write, the two corruptions, the set-aside store and the sweep's target are inside that home. demo-felhom is a different sub-account (u629488-sub1) and a real customer's copy is a different sub-account again — none reachable with this key. restic forget/prune were never invoked by me; the nightly retention that ran as part of the customer-path off-site button kept every 9-August snapshot (verified by listing all 24, not by a count).
  • demo-felhom's two fixtures — confirmed intact, not assumed. (a) the unopenable set-aside store u629488-sub1:/home/felhom-repo.orphaned-20260810 — listed tonight, config/data (258 shards)/index/keys/locks/snapshots, mtimes still 18 Jul and 3 Aug; (b) the retained-key case — hub host_escrow_superseded rows 11 and 12 for demo-felhom-8363b5, each with a 572-byte identity_blob, dated 2026-08-12. demo-felhom is healthy on 0.217.0, its own off-site ran at 02:15 with last_status: ok, and --abandon-status there reports „no abandonment countdown is running on this box".
  • peti-felhom — not contacted. It does not appear in the hub host list at all; no command in this session named it.

11. SCHEDULED CYCLE AND THE COUNTDOWN

(filled in as they fire — 02:30 local backup, 04:15 off-site, 05:10 abandonment sweep, all CEST)


12. MACHINE STATES AT THE END

(final state recorded in §13 after the scheduled runs)