Files
felhom.eu/documentation/audits/D5-drive-alone-restore-2026-07-30.md
T
admin 0a9bd3829d D5 SHIPPED: Tier-1/2 restore no longer depends on the whole-guest tier
Records controller v0.188.0 across the four coupled artifacts.

07-backup-architecture.md is the owning doc:
- new 7.4 = the recovery chain AFTER D5 (7.1 leg 1 superseded; leg 2,
  the living-app dependency, explicitly unchanged so this is not read
  as more than it is)
- 7.3 collapsed to history, with the correction that the target as
  written (data_key-only) was tested in Part 0 and rejected
- 3 records that the two-lane split is now real, not just intended
- matrix rows 3 / 3c (new) / 13; 10.1 D5 itself shipped

Also: new capability-map row, D5 collapsed in ROADMAP + OPEN-ITEMS,
and R-127 filed in both (data_key flag unreliable; O4 can regenerate a
DB password that no longer matches the restored data directory).

The audit is named D5-drive-alone-restore rather than "...secrets..."
because .gitignore blocks *secret* -- a guard worth respecting, not
forcing past.
2026-07-30 16:58:06 +02:00

16 KiB

D5 — an app restore works from the drive alone (controller v0.188.0, 2026-07-30)

Verdict: SHIPPED + PROVEN-LIVE. Tier-1/Tier-2 no longer depend on the whole-guest tier. A customer needs the drive and nothing else to restore an app.

Owning design doc: architecture/07-backup-architecture.md §7.4 (new), with §3, §7.1, §7.3, §8 rows 3/3c/13 and §10.1 corrected. Capability map: new row in §D.


1. Part 0 — the rulings, and why the brief's own recommendation was rejected

The task proposed that only data_key-flagged secrets travel, and asked for that to be tested rather than obeyed. It did not survive.

1.1 data_key: true is not a trustworthy classification → R-127

Only 5 fields across 4 apps set it (adventurelog/SECRET_KEY, homebox/HBOX_AUTH_API_KEY_PEPPER, papra/AUTH_SECRET, sparkyfitness/{API_ENCRYPTION_KEY,BETTER_AUTH_SECRET}). But the catalog's own Hungarian labels contradict the flag elsewhere:

Field Catalog's own label data_key
n8n/N8N_ENCRYPTION_KEY „Titkosítási kulcs" false
wanderer/POCKETBASE_ENCRYPTION_KEY „Adatbázis titkosítási kulcs" false
calcom/CALENDSO_ENCRYPTION_KEY „Titkosítási kulcs" false
bookstack/APP_KEY „Alkalmazás kulcs" false
adventurelog/SECRET_KEY „Titkosítási kulcs" true

Same label, opposite flag. Travelling "only data keys" would have omitted real data keys, and the fail-closed gate (restore_unit.go) would not fire for them — a restore that succeeds onto unreadable data.

1.2 A DB password is not resettable in practice — proven, not argued

DumpAppVolumes (internal/backup/backup.go:549) dumps every compose named volume with no DB exclusion, so a postgres data volume is captured and restored. Probe on a throwaway postgres:16-alpine (seed with pw A → drop container, keep volume → redeploy with regenerated pw B, which is exactly the O4 path):

Leg Result
replaydocker exec psql -U app -d app, no password (verbatim appbackup.ImportDump) SUCCEEDS — the image's local socket is trust
app path — pw B over the compose network FATAL: password authentication failed for user "app"
app path — pw A (the hash still in the restored PGDATA) SUCCEEDS

POSTGRES_PASSWORD is ignored once PGDATA is non-empty (initdb skipped). So the restore reports success, the dump replays, the rows are there, and the application cannot reach them — the §4 "looks like a pass" failure. 18 DB/root-password fields affected.

A first attempt tested -h 127.0.0.1 and got a false pass: the postgres image ships a trust line for localhost. Only the compose-network path is the app's real path. Recorded because it is the same class of false observable this task exists to catch.

MariaDB fails louder: getMariaDBPassword (appbackup/dbdump.go:737-754) reads the regenerated value from container env against a datadir holding the old hash → Access denied, so the replay itself fails (nextcloud, romm).

1.3 The rulings (operator, 2026-07-30)

Classifying all 53 templates mechanically produced a boundary the brief's framing missed:

  • TRAVELS — every type: secret field (45): 5 declared data keys + 18 DB/root passwords + 22 internal signing/encryption secrets. Each decrypts data on the same drive or authenticates to a container on an internal compose network with no external listener — so possessing it adds nothing to possessing the drive, which is exactly D2's argument for plaintext DATA.
  • WITHHELD — every type: password field (7) plus the nonPortableSecrets register (vaultwarden/ADMIN_TOKEN, which gates the app's public /admin panel). These authenticate against published services, so their blast radius is not bounded by the drive.
  • Plaintext, as the data already is.

Excluding the withheld class is what licenses the plaintext ruling. The two are coupled and must not be relaxed independently.

What a customer must possess to complete a Tier-1/2 restore after this change: the drive.


2. Implementation

One place per side, no parallel path.

Concern Where
The boundary stacks.PortableSecretEnvVars + the nonPortableSecrets register (internal/stacks/deploy.go) — code, not a catalog flag: a boundary a catalog push can silently move is not a boundary (R-97a)
Decrypt for capture GetStackRecoveryInfo (cmd/controller/main.go), via the same LoadAppConfigDecrypted the restore side uses
Write the split buildUnitAppYaml (was buildStrippedAppYaml) — 0600, and it names the withheld class in the header so an operator sees WHY a credential is absent rather than suspecting a capture bug
Read it back readUnitEnv — split driven by the manifest's portable names, never guessed from key names
Reconcile reconcileRestoreSecrets stays a pure function; the new source arrives as an argument
Manifest schema 2 + portable_secret_env_vars (NAMES only — the manifest is 0644)

Precedence: the UNIT WINS. Not "newest wins". The unit's secrets are captured in the same run as the dumps beside them (runVolumeDumpscaptureAllRecoveryUnits), so the unit's value matches the data about to be restored; the guest's is merely the most recent. A rotated data key does not decrypt data encrypted with the old one; a rotated DB password does not match the hash in the restored data directory. An undefined precedence between two sources of a decryption key is a data-loss bug waiting for its first disagreement.

The fail-closed gate is unchanged. A data key in neither source still refuses outright. D5 makes it normally present; "normally" is not a reason to soften a gate.

Backward compatible. A schema-1 unit carries no secrets and still restores from the guest; the next capture rewrites it (the app.yaml checksum changes).

Three comments that asserted invariants D5 makes false were corrected rather than left to read as settled (CaptureRecoveryUnit "NEVER writes a secret value", RestoreFromRecoveryUnit "no secret is read from the unit", appbackup/paths.go + appdata.go "secret-free"), and the O4 WARN that claimed "stored data is unaffected" for every non-data-key secret was scoped to what §1.2 shows is true.


3. Tests

Suite green, rc=0, read on its own (never combined with a commit).

Test Asserts
TestRestoreFromRecoveryUnitWithGuestAbsent D5's claim as a test: the restore SUCCEEDS with the guest absent
TestRestoreFromRecoveryUnitGuestAbsentStillFailsClosed data key in neither source ⇒ refuse, nothing mutated, app not even stopped
TestReconcileRestoreSecretsPrecedence unit-wins; guest fills in when the unit is silent or empty; a secret never shadowed by stray plain config
TestCaptureRecoveryUnitCarriesPortableSecretsOnly portable present + 0600 + schema 2; the withheld value appears NOWHERE in the unit (wrong-outcome check)
TestReadUnitEnvSplitsByManifest manifest-driven split; schema-1 degrades without losing entries
TestPortableSecretEnvVars the boundary itself, both directions, incl. the register being slug-scoped
…Orchestration/schema-1 unit still restores no regression for pre-D5 units

Seam (R-125): Manager.stackProvider only — the docker/compose ops and the guest app.yaml decrypt. RecoverStackSecrets returning nil is the guest being gone: it is what the real adapter returns when the stack or its app.yaml cannot be read. Everything else is production code, and fixtures come from a unit written by the real CaptureRecoveryUnit, so capture and restore meet at real bytes rather than a test agreeing with itself.

Four red-proofs, each mutation verified present in the working file before the run:

# Mutation Result
1 drop the PortableSecrets merge in buildUnitAppYaml consequence + capture tests FAIL (data-encrypting key did NOT travel)
2 if false && the fail-closed gate 4 failures across orchestration + pure-function tests
3 flip precedence to guest-first both sources disagree — the UNIT wins FAILS (got "guest-key")
4 widen the class to include type: password the boundary test FAILS (admin password becomes portable)

4. Live proof — the data, not the exit code

Venue: a scratch drill LXC (9301) on felhom-pve, destroyed after the run. runbooks/target-selection.md makes this the correct class: D5's claim is about restore code behaviour, not the install path / golden / first-boot state, so a fixture-class venue applies and a fresh day-0 is not required. Guest 9201 (both hosts), drill-r50 (VM 300) and demo-hp's local-lvm were all fenced by the task and untouched. demo-hp could not host the guest at all: its local storage has no rootdir content type and local-lvm was forbidden.

Image identity confirmed: the guest pulled digest sha256:f297fbf0… — byte-identical to what build.sh 0.188.0 --push produced.

Path exercised through the real endpoints the UI invokes: POST /api/stacks/{app}/deployPOST /api/backup/runPOST /backup/restore.

4.1 The portable half — AdventureLog (SECRET_KEY data_key + DB_PASSWORD)

[backup] Recovery unit captured for adventurelog → …/backups/primary/adventurelog
         (images=3, secrets-referenced=2, data_keys=1, portable-carried=2/2, withheld=0)

Unit compose/app.yaml: mode 0600, both secrets present in plaintext, DB_PASSWORD fingerprint identical to the password the running postgres was built with. The guest's own app.yaml holds the same names ENC:-encrypted. Manifest: schema_version: 2, portable_secret_env_vars: [SECRET_KEY, DB_PASSWORD].

The guest's app.yaml was then moved aside, and the restore run:

[backup] Restoring adventurelog from recovery unit: images=3, secrets recovered=2/2, data_keys=1
[backup] Restore-from-unit completed: adventurelog          (27.6 s)

2/2 with the guest gone. Pre-D5 this was 0/2 and the data-key gate refused.

The positive observable — the application reading its own data:

APP-READ-OK rows= ['D5-CUSTOMER-DATA-2026-07-30']
connected_as= adventurelog  over_TCP= True  row_count= 1

Read through Django's own connection, as the app, over TCP — not the localhost trust socket that produced §1.2's false pass. Plus manage.py check --database default"System check identified no issues", and 51 public tables intact.

The discriminator held, which is what makes this a restore and not a no-op: the row seeded before the backup came back, and a row inserted after the backup was gone. And the unit contained no .sql dump (db_dumps: None) — the DB came back from the volume tar, i.e. exactly the case where a regenerated password would have failed silently.

Guest app.yaml was recreated from the unit with both secrets re-encrypted at rest.

4.2 The withheld half — Grafana (GF_SECURITY_ADMIN_PASSWORD, type: password)

[backup] Recovery unit captured for grafana → … (images=1, secrets-referenced=1, data_keys=0,
         portable-carried=0/0, withheld=1)

A distinctive sentinel was supplied as the admin password. It was live in the container (1 env match) and ENC: in the guest app.yaml — so the absence below is not vacuous — and it appeared in:

  • 0 files in grafana's recovery unit;
  • 0 files anywhere under the whole backup namespace.

The unit's app.yaml header names it explicitly: # WITHHELD (internet-reachable logins — stay in the guest, regenerated on restore): GF_SECURITY_ADMIN_PASSWORD

4.3 Not exercised live

  • O4 regeneration of a withheld secret on restore (unit-tested only).
  • Tier-2's cross-drive copy of a secret-bearing unit (unit-tested only; no code changed).
  • The offsite leg. Recorded consequence, not a change: the unit already travels to Tier-2 (another customer drive, plaintext, same reasoning) and offsite via restic (encrypted at rest under the customer-owned repo password). No tier code was touched — the secrets simply travel inside the unit those tiers already carried.

5. R-120's gate does not apply to this task

The gate lives in hub/internal/web/configs.go handleSetArtifacts — the golden vouch form, the sole UI path to SetArtifactManifest. It compares a submitted golden against store.NewestReportedControllerVersion(). It therefore never runs on a controller image deploy, and did not block anything here.

Re-baking the golden (currently 0.186.0) is a follow-on so that FRESH installs get D5 — not a prerequisite. It is its own ceremony (bake fixture, upload, vouch, live day-0 proof) and bundling it into D5 would have put two unrelated risky ceremonies in one session.


6. Teardown — all three layers

Layer Disposition
Guest LXC 9301 destroyed (pct destroy --purge); vm-9301-disk-0 removed. pct list shows only 9201, untouched
Host (felhom-pve) no leftover drill files in /tmp; thin pool back to 27.88 %, identical to pre-drill, only 9201's volumes remain
Hub nothing to dispose — the drill ran hub.enabled: false and never registered. Verified positively, not by assuming: /hosts and /configs both loaded (9880 / 10160 bytes) and returned 0 occurrences of d5drill

Secret-bearing files shredded (shred -u -z) inside the guest before the volume was destroyed: both units' compose/app.yaml, both stack app.yamls, encryption.key, the hidden guest app.yaml, and every transported helper script — plus the host-side tarballs and the local scratchpad scripts.

One self-inflicted slip, recorded rather than smoothed over: an inspection script's redaction regex assumed 2-space YAML indentation while the unit's app.yaml uses 4, so the drill's SECRET_KEY and DB_PASSWORD values printed to the session transcript once. They were throwaway values in a guest that no longer exists, and they reached no committed file — but the guard was mine to get right, and the lesson is that a redaction pattern is itself a claim that needs checking against the real format.


7. Findings

  • R-127 (NEW, READY): data_key: true is unreliable (leg a) and O4 can regenerate a DB password that no longer matches the restored data directory (leg b). Filed, not fixed — leg (a) is a catalog change plus a flag-vs-label gate; leg (b) wants either fail-closed treatment for DB passwords or an ALTER USER after the volume restore. Both are outside D5's ruling and neither is what D5 was asked to do.

D5's own defect surface — the false "stored data is unaffected" WARN and the three stale invariant-asserting comments — was corrected in place, since that is the defect this task was inside.

8. Arc status

The arc's definition of done is MET. Its four named components — R-106, R-109 (both agent v0.118.1 / hub v0.83.0), R-108 (controller v0.187.0) and D5 (controller v0.188.0) — are all SHIPPED + PROVEN-LIVE.

Still open, none of them part of that definition: R-105 (three hub-held DR records empty fleet-wide), R-121 (installed-vs-vouched agent drift), R-123 (PARTLY CLOSED — registry hygiene), R-124 (PBS root namespace spelled "root" vs the API's ""), R-126 (.fab export onto a NAS), R-117g (no guided recovery for an aborted fs), R-117h (StablePathForRaw hardcodes the live parent, so the stale-bind repair is not test-seamable), and R-127 (new, above).

R-119 does not exist — it appears in neither OPEN-ITEMS.md nor ROADMAP.md. Reported as absent rather than given an invented status.