Files
felhom.eu/documentation/audits/DRILL-the-28-2026-09-22-HEAD.md
T
admin 1de6aaf904
gates / gates (push) Successful in 25s
corrections: calibre-web restore was REFUSED, calcom was INCONCLUSIVE
Both from their own records, not re-run.

calibre-web: the flash_error refusal is in its log; that walk predated the refusal-capture code.
The report's own prose already said so while the table disagreed.

calcom: the restore was accepted and the app read running; the classifier then saw "starting", a
settling state it did not list beside running/unhealthy, and fell through to failed. The brief
supposed a read-back artefact - that is wrong, and restore_state_seen says so.

Restores correctly refused: 2 -> 3.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-22 20:18:37 +02:00

6.0 KiB

THE TWENTY-EIGHT — every app no update drill had ever touched, 2026-09-22

Evidence: the-28-2026-09-22/. What came before: DRILL-update-night-2026-09-21.md (21 edges, 19 apps) and PROBE-FIX-2026-09-22.md (the probe fix and the fifteen moves).


Not done, or changed from the brief

Read this first.

Claims in the brief that turned out wrong

  1. There is no requires: key in .felhom.yml. The brief said "A .felhom.yml requires: (HDD, x86) that 9202 cannot meet → recorded, not forced." No template has such a key. The constraints live under resources: as needs_hdd and pi_compatible. And none of them excluded anything: 9202 is x86_64 with /mnt/felhom-drives/scratch_hdd mounted, so every one of the 28 was installable on those grounds. Nothing was skipped for a resource reason.

  2. The brief's file-leg list is wrong, and it told me where to check. It grouped "immich, jellyfin, plex, emby, komga, calibre-web, gokapi, homebox, gramps-web" as file-leg apps. Read from 07-backup-architecture.md §6.2 — which the brief itself says to read rather than re-derive — only four of the 28 are class A (at least one readable file leg): calibre-web, immich, komga, paperless-ngx. jellyfin, plex and emby are class B precisely because their only bind is a :ro media mount, which ClassifyBinds excludes; gokapi, homebox and gramps-web keep everything in named volumes. The table below uses 07 §6.2's classes, not the brief's.

  3. The brief's database list is incomplete. It named calcom, claper, outline, paperless-ngx, rallly and sparkyfitness for PostgreSQL and kimai for MariaDB. immich also carries PostgreSQL (a vectorchord variant) and redis, and wanderer carries meilisearch — two apps the brief put in the "file-leg" group actually run their own datastore.

  4. One of the 28 cannot be installed at all, by design. plant-it is lifecycle: "abandoned", and the product's lifecycle gate refused the deploy with 409 „Ez az alkalmazás jelenleg nem telepíthető." That is correct behaviour, measured live for the first time. It is the only lifecycle-gated template in the whole catalog of 53.

Claims in the brief that were verified true, by looking

  • The drill repo's Actions are off (R-629). Read from the API: has_actions: false, private: true. And measured rather than trusted: 47 CI jobs before the reset push, 47 after — the push produced no run and no mail.
  • repoint_drill.py still works after the catalog moved. 9202's cache now reads origin …/app-catalog-drill.git at 1ad1f34.
  • 9202 has the capacity. 25.9 GB RAM (23.5 free), 28 GB free on /, 842 GB on the scratch drive. The root disk is the binding constraint, so each app's images are removed by name after its verdict — never prune (rule 3). Disk held at 1.9 GB used throughout.

Changed method, named

  1. A restore that is REFUSED is recorded as refused-with-a-sentence, not as a failed restore. The first version of the harness collapsed the two and mislabelled calibre-web. The product had in fact done the right thing — see the finding below — and a harness that calls a correct refusal a failure would have buried it.

  2. The harness now waits for a restore to settle before removing. It did not at first, and that race produced R-633, a real defect. The race was left in the record for gokapi and fenced out afterwards, so the remaining apps measure the product rather than the harness.

  3. Three bugs in tonight's own harness, each named with what it cost. A missing import re in the restore step killed the restore half for wanderer, claper, sparkyfitness and calcom. A variable named m shadowed the app's metadata and broke paperless-ngx's teardown. An empty phase list crashed on [-1] when the Update was refused before any phase existed, which cost rallly its whole walk. All four apps, and five more, were re-walked serially afterwards — and that re-walk is what corrected R-634 and produced ghost's proof. An instrument that can drop results silently is not a measurement; these dropped them loudly and were re-run.

  4. Concurrency is part of the method and it changed two results. Three walks ran at once to fit 28 apps in one night. POST /api/backup/run is box-wide, so a second caller gets 409 „Mentés már folyamatban", and the Update refuses while a backup or restore is in flight (409 busy). Both refusals are the product being right and both are quoted below. But they cost ghost and rallly their edge on the first pass, and they are implicated in two of R-634's three instances. The nine re-walks were serial for exactly this reason.

Corrected 2026-09-22 (evening), from the records rather than by re-running

  1. calibre-web's restore was refused-with-a-sentence, not failed. The refusal is in this app's own log.txt line 12 as a flash_error on the redirect, and its state stayed running throughout. It was recorded failed because that walk ran before the refusal-capture code was added later the same night — the document's own section "A restore that is REFUSED" already said so while the table and the record disagreed with it.

  2. calcom's restore was inconclusive, not failed — and that was my harness, not the product. The restore was accepted (flash=flash.restore.started, no flash_error), the app read running at +45 s with no hold and no phase, and the classifier looked once more and saw starting — a settling state it did not list beside running/unhealthy, so it fell through to failed. The task brief supposed a different cause — "a read-back of data that was never seeded" — and that is wrong: the seed half is recorded separately and was already no route. The record's own restore_state_seen: "starting" is the evidence.

    Totals move with it: restores correctly REFUSED go from 2 to 3.