Files
felhom.eu/documentation/audits/night-2026-09-26/E/E1-README.md
T

4.0 KiB
Raw Blame History

E1 — kept data: the spike on 9202 (controller 0.272.0, live catalog), 2026-09-25 12:36–13:04 CEST

Written before any build. nextcloud installed through the product with HDD_PATH=/mnt/felhom-drives/scratch_hdd (the drive root, as households have it — demo-hp's apps read HDD_PATH: /mnt/felhom-drives/hdd_1), so its files live at <drive>/appdata/nextcloud. Seeded through occ user:add; before-backup.txt written through WebDAV as that user; a unit taken by the product's own nightly db-dump leg (window moved to 12:44 by POST /backups/window, then back to 02:30 — E1-01, E1-02, E1-03); after-backup.txt written through WebDAV AFTER the unit.

Q1 — remove „keep my data" (+ backups kept) → the removed-app row → restore: does nextcloud come back whole?

FALSE on 0.272.0 — it comes back BROKEN. (E1-04, E1-05)

  • The row is there (R-487: Nextcloud in the /backups/restore picker), but the picker's snapshot list for it is EMPTY (snapshots() offered 0) — so a household cannot even choose a copy. The form post with snapshot_id=helyi (what the R-487 row sends) was accepted.
  • The restore then logged No readable recovery unit for nextcloud at /mnt/sys_drive/felhom-data/backups/primary/ nextcloud — falling back to volume-only restore — the unit sat on the DATA drive, .../scratch_hdd/backups/ primary/nextcloud, fully readable. compose up ran WITHOUT the app's env: The "HDD_PATH" variable is not set, the app container mounted /appdata/nextcloud on the guest's ROOT disk, nextcloud-db crash-looped (Database is uninitialized and password option is not specified), volumes 0/0, dbs 0/0, no app.yaml.
  • Root cause (read from source): backup.primaryUnitDirFor and backup.ListRestorePoints decide "is this app removed?" with stackProvider.GetStackComposePath(name), which in production answers true for every catalog app (it asks whether the STACK exists — every template is a stack). The R-487 test's fake answers it for deployed apps only, so TestR487_PrimaryUnitDirForNamesTheRemovedUnitWhereItSits passes while the box fails — a seam that lies. Neither R-487 branch has ever run on a box for a catalog app on a data drive. Filed R-690, fixed in this build (isStackDeployed, the list's own predicate); /appdata/paperless on 9202's root (2026-09-15) looks like the same accident, earlier.

Q2 — a file written after the last backup: visible, or does nextcloud need occ files:scan?

TRUE — it needs the scan. (E1-07, BY HAND — the product route was broken, so the unit's OWN volume tars were poured into fresh volumes and the stack started with the unit's HDD_PATH; the database is what the unit holds.) After the start: the seeded account reads back; before-backup.txt visible; after-backup.txt NOT visible though it is on the drive. occ files:scan --all (0.6 s, 130 files, 4 updated) → after-backup.txt visible. So nextcloud's template gets after_load: = php occ files:scan --all as www-data in nextcloud.

Q3 — remove with the backups ALSO deleted → files only

TRUE. (E1-08) remove_backups deleted the unit (928 MB); appdata/nextcloud (126 MB, both files) stayed; the removed-app row is gone. So "use my kept data" must be OFF here (no database copy) — the brief's use-off sentence.

And R-657's loop — measured again as a baseline

A fresh install over the kept folder on 0.272.0 did not loop this time: occ status → installed: true after 60 s. The old account (drilld112bd) and its files are invisible to the new install — its database knows nothing of them. So the harm is data that silently stops being reachable, and the loop is one way it can show. Either way a silent install over kept files is what E2 replaces with the choice.

Teardown: nextcloud removed through the product (keep data, backups deleted); the hand-made stack down -v; appdata/nextcloud (126 MB) kept on purpose — it is E5's kept-data fixture. docmost untouched. Controller on 9202 was 0.272.0 throughout — not interrupted.