Files
felhom.eu/documentation/audits/SPIKE-restic-restore-test-2026-08-31.md
T
admin 130f7a6eba
gates / gates (push) Failing after 17s
R-87 SPIKE: measured, do not build it as written (R-407..R-409 filed)
Spike. NO production code. No version bump, no build, no deploy, no golden.
felhom-controller and felhom-agent were READ ONLY. The fleet stays on v0.230.0.

Q1 restic is 0.14.0 (go1.19.8, bookworm 12.15) - the four source comments asserting
it are CONFIRMED, not corrected.

Q2 --verify DOES exist and is NOT a content check. Red-proof: one byte changed in a
restored 160 MB tar with size and mtime preserved passed clean, rc=0. Verify took
131 ms on a 213 MB / 7-file tree, which cannot be hashing. A size or mtime mismatch
causes a silent re-download, not a failure. Controls: --target 1 hit, four post-0.14
flags and a nonsense string 0 hits each. Neither --verify nor --no-lock appears
anywhere in the controller source.

Q3 no reference for "correct" exists. restic ls --json carries no content hash in
0.14.0, and the unit manifest hashes 4918 B of a 213231242 B unit - 0.0023 percent,
the config files and not the dumps or the tars. R-409.

Q4 it is CHEAP. All 8 apps / 774378123 B logical restored back to back in 25 s, against
40257 ms for the weekly 100 percent check beside it. Individual restores 2253-3978 ms
regardless of size: cost is per-snapshot round-trip plus ~1 s per 200 MB. Peak scratch
is the app's full logical size. The 1.1 MB restic cache is index only and hides nothing
(--no-cache 5423 ms vs cached 3198 ms, trees byte-identical).

Q5 skip-if-busy stays right at 25 s against a 2m52s nightly backup. But
RestoreOffboxScratch takes NO acquireRunning, while offbox_integrity.go:28 asserts
every off-site operation does. R-408.

Q6 observed with a positively-controlled lock sampler: restic restore takes NO lock;
restic check DOES (locks 0 -> 1 for nine samples -> 0 across the check, zero across two
restores). The product writes anyway - unlockStale runs `restic unlock`, a delete verb,
before every restore (offbox_restore.go:289). The task's lead was right in direction and
wrong in mechanism. R-95's constraint IS satisfiable: --no-lock plus skipping unlockStale
writes nothing, and both mechanisms exist unused. Neither was fixed - the task forbids it.
offbox_integrity.go:255's "It NEVER writes to the repository" is R-407.

Q7 THE DECIDING ONE: of R-353/354/356/358/403 an unattended scratch-restore would have
caught ONE (R-356). The value is elsewhere, and the weekly check structurally cannot
reach it: `check` proves the stored bytes are the stored bytes, never that we stored the
RIGHT thing. A hollow unit backs up, checks at 100 percent and restores cleanly and
recovers nothing - R-403, measured in bytes on 31 August.

RECOMMENDATION: option C, the NARROW test - one app a night, restored to scratch, checked
against its own manifest.json through the existing unitCarriesData, scratch deleted, the
SNAPSHOT recorded as the proof. Options A (do not build) and B (scheduled attended drill)
considered explicitly; B is weakest because it is what already happens. R-87 should be
RE-SCOPED, not built as written, and that is Viktor's call - the row stays open carrying
the verdict and STATUS.md item 4 asks it in plain words.

Also corrected in 07-backup-architecture.md: matrix rows 4 and 10 both said "the depth
that ships ON does not re-read pack contents (R-399)". R-399 CLOSED in v0.228.0 and the
depth is 100 percent. Two stale cells, fixed, and the spike verdict added beside them.
Row 4's verdict is UNCHANGED by the spike and now says so.

Teardown: all three layers, none of them "nothing was created" - 6 files on the PVE host,
9 in the guest, 5 plus 2 run-flags in the container, all removed and verified empty. The
four scratch directories this session's restores created were removed; three that
pre-date the session were left alone. Two state changes recorded rather than hidden: the
control integrity run recorded its verdict (depth structure -> 100%, due-ness +7 days),
and four restores appear in the controller log. Nothing was written to the off-site
repository by hand.

Evidence: documentation/audits/evidence-spike-restic-restore-2026-08-31/ - 31 files,
every one pulled off the box BEFORE teardown (R-320).

golden-currency is RED at this commit and was already red at dddcc80. Pre-existing, not
this session's debt. Second --no-verify push of the day for that reason; R-404's count
goes six -> seven and its row says so.

Ceiling R-406 -> R-409.
2026-08-31 15:59:09 +02:00

26 KiB
Raw Blame History

SPIKE — can the off-site (restic) copy be restore-tested without a person? (R-87)

Date: 2026-08-31 · Venue: demo-hp (Tier 0, disposable), guest 9201, controller v0.230.0 Class: spike. No production code was written. No version bump, no build, no deploy, no golden. Evidence: documentation/audits/evidence-spike-restic-restore-2026-08-31/ — 31 files, all pulled off the box before teardown (R-320).

Baselines re-confirmed at the start of the session: felhom-controller main @ 2d802d75e88616d86cbade8a0e16965c2b85771c, v0.230.0; felhom.eu @ dddcc808be95d1c89b276b4d791491bad3c96bba (clean tree, HEAD == origin/main); felhom-agent 058b945, v0.130.0.


The one-paragraph answer

restic 0.14.0 cannot tell us a restore produced correct files, and neither can anything else the box holds today — --verify exists but checks size and modification time, not content, and the recovery unit's own manifest records a hash for 4 918 bytes of a 213 231 242-byte unit. But an unattended restore-test is far cheaper than expected: restoring every app on the box — 8 snapshots, 774 MB logical — took 25 seconds, against 40.3 seconds for the weekly integrity check sitting beside it. And the question worth asking is not the one R-87 was filed for. Of the five restore-path defects human drills found between 2026-08-26 and 2026-08-31, an unattended scratch-restore would have caught one. What it would catch, and what the weekly check structurally cannot, is a snapshot that is perfectly intact and contains nothing recoverable — the R-403 shape, measured live nine days ago. Recommendation: build the narrow version (option C below), and do not build the thing R-87 asks for.


Part 1 — the register correction (done, pushed as 6e550ae)

When and where it went wrong. R-87 was moved into CLOSED-ITEMS.md by commit ef6ac6f, 2026-08-22, "One register, enforced by a gate; closed work compressed into siblings (R-376..R-378)". Established from git log -S, not inferred: that commit is the only one that ever added an R-87 row to CLOSED-ITEMS.md and the only one that removed it from OPEN-ITEMS.md.

It is a survivor of R-378, not a separate incident. R-378 records that same sweep moving six still-open rows — R-123, R-190, R-214, R-264, R-295, R-352 — and restoring them verbatim in the same session. R-87 is a seventh it missed. Its state cell read READY — RE-RANKED UP 2026-08-03 (R-86 closed): the leading verdict is READY, and the word closed later in the same cell describes a different row. Nine days in the wrong file, while OPEN-ITEMS.md's ranking paragraph ranked it fourth and pointed at nothing.

The count, reproduced independently — and the predicate decides the answer.

predicate rows convicted in CLOSED-ITEMS.md (of 151) which
open word anywhere in the row 144 meaningless
open word anywhere in the state cell 3 R-87, R-224, R-260
open word in the leading verdict 1 R-87

R-224 and R-260 are genuinely closed; their long prose verdicts merely contain the words "open" and "OPEN". The task author's count of one is right, and it is right only under the leading-verdict predicate — which is R-378's own lesson, restated by measurement.

The gate: scripts/closed_register_gate.py, two rules — no open state word leading a CLOSED-ITEMS.md row's verdict, and no R- id with a row in both registers. Red-proofed on both rules; negative-controlled against the files as they were pushed, where it convicts R-87 by name and exits 1. Registered as the 12th gate in repo_gates.py after it was green. Four residual holes are named in its docstring.

One thing the second rule turned up: R-398 also had a row in both registers — a deliberate cross-reference stub. It is now prose beneath the table, not a table row.

A duplicate this session did NOT fix: OPEN-ITEMS.md carries two unrelated findings both numbered R-133 (:267 hub customer_id uniqueness; :273 plaintext break-glass credential). Filed as R-406; deliberately not gated, because a within-register duplicate rule would fail on a pre-existing row and a registered-but-failing gate refuses every push.


Q1 — what restic is actually running?

Answer: restic 0.14.0, and every source comment asserting that is correct.

From the running container on demo-hp, not from the Dockerfile:

restic 0.14.0 compiled with go1.19.8 on linux/amd64
/usr/bin/restic
/etc/debian_version → 12.15   (Debian bookworm, as the Dockerfile says)

Method: pct exec 9201 -- docker exec felhom-controller restic version, rc=0. Evidence 02-q1-restic-version.txt. The comments at offbox_capture.go:15, offbox_restore.go:21, offbox.go:726, offbox_progress.go:48,66 are confirmed, not corrected.


Q2 — what can that version verify about a RESTORE?

Answer: --verify exists, and it does NOT verify content. It cannot tell us a restore produced correct files.

--verify is real. restic restore --help on the running container lists --verify verify restored files content. Controls, because a grep -c 0 must be earned:

control result
positive — --target present 1 hit
--verify present 1 hit, quoted verbatim above
negative — --delete, --dry-run, --overwrite, --sparse (all post-0.14) 0 hits each
negative — ZZZ-NOT-A-FLAG 0 hits

Evidence 03-q2-restore-help.txt, 04-q2-controls.txt. --verify and --no-lock both exist in 0.14.0 and neither appears anywhere in the controller source — grep -rn over felhom-controller/controller/ returns rc=1 for both, with --json/--target as the positive control (05-q2-codebase-verify-grep.txt).

What --verify actually checks — measured, with a red-proof and a negative control.

test what was done result
cost virgin restore of kimai (213 231 242 B, 7 files) with --verify verify itself 131 ms; restore+verify 3 438 ms
red-proof corrupt one byte in a restored 160 MB tar, size and mtime preserved, re-run restore --verify PASSED clean, rc=0 — the corruption was not detected
control (size) truncate the same file by 1 byte, re-run restic silently re-downloaded it; verify reported "7 files, 133 ms", rc=0
control (mtime) corrupt content and bump mtime, re-run restic silently re-downloaded it; rc=0
negative control same command against the untouched copy identical output — so "passed" carries no information

131 ms cannot hash 213 MB. Combined with the red-proof, --verify in 0.14.0 is a size-and-mtime reconciliation that re-fetches anything that disagrees. It is useful — it makes a restore self-repairing — and it is not a content check. Evidence 20-, 21-, 22-.

This is the point at which the spike's shape changed, per §9's instruction to stop and reconsider after Q1/Q2: the tool cannot supply the reference, so Q3 became the hard question.


Q3 — if not restic, then what is the reference for "correct"?

Answer: there is none available to an unattended test today. The one hash record that exists covers 0.002 % of a unit's bytes.

candidate verdict why
restic --verify rejected size + mtime only — Q2's red-proof
the snapshot's own metadata rejected restic ls --json file nodes in 0.14.0 carry name, size, mode, uid/gid and three timestamps — no content hash (24-q3-hash-coverage.txt)
a hash the controller already records rejected as a content reference, kept as a completeness one see below
restic check --read-data already shipped, different question it proves the STORE's packs, never the restored files
a planted sentinel rejected a drill technique. An unattended test may not write data into a customer's app to have something to look for
the live data rejected it drifts by design; the snapshot is 12 h old by the time a check runs
restore twice and compare rejected proves determinism, not correctness

The hash record that exists, and its exact coverage. The recovery unit's manifest.json carries a checksums object. For kimai, measured on the restored unit:

checksums: .felhom.yml (2 235 B), app.yaml (488 B), docker-compose.yml (2 195 B)   = 4 918 B
files in the unit:  + kimai-mariadb.sql 48 217 B
                    + kimai_kimai_db_data.tar 160 331 776 B
                    + kimai_kimai_var.tar      52 845 056 B
                    + manifest.json 1 275 B                  total 213 231 242 B

4 918 of 213 231 242 bytes — 0.0023 %. The three config files are hashed; the database dump and the two volume tars, which are the recoverable data, are not. Filed as R-409.

What the manifest CAN answer is a different and better question. It declares db_dumps and volume_dumps by name, and unitCarriesData (r403_hollow.go:40) already asks it. A restored unit can therefore be checked for completeness — does every file the manifest declares exist — with no new metadata, no new reference, and no content hash. That is the whole of the recommendation in §Recommendation.


Q4 — what does one restore-test cost?

Answer: about 4 seconds per app and 25 seconds for the whole box — cheaper than the weekly integrity check it would sit beside. The cost is dominated by per-snapshot round-trip, not by data volume.

Through the product's own path (POST /backup/offbox/restore → RestoreOffboxScratch), timed from the POST to the completion log line:

app mode logical size wall clock
docmost unit 118 207 270 B 9 s (13:40:10 → 13:40:19)
kimai full 213 231 242 B 11 s (13:42:54 → 13:43:05)
opengist unit 185 664 B ~8 s

Raw restic, all eight snapshots back to back (26-q4-all-apps-cost.txt):

app logical restored to disk ms
privatebin 2 123 006 2 159 870 2 628
opengist 185 664 222 528 2 253
calibre-web 5 808 703 5 878 335 3 642
paperless-ngx 83 473 654 83 547 382 3 062
bookstack 166 641 768 166 682 728 2 745
docmost 118 207 270 118 248 230 3 676
romm 184 706 816 184 747 776 3 978
kimai 213 231 242 213 272 202 3 202
all eight 774 378 123 774 759 051 25 s

185 KB takes 2.25 s and 213 MB takes 3.20 s. Nearly all of it is fixed per-snapshot cost — opening the repo, loading the index, one SFTP session. Data adds roughly 1 s per 200 MB (≈ 67 MB/s on this link).

The cache is not hiding the cost. /root/.cache/restic is 1.1 MB — index and metadata only. Cached 3 198 ms vs --no-cache 5 423 ms for kimai; the trees are byte-identical (diff -r → YES). Pack data always crosses the wire. Evidence 25-q4-cache-effect.txt.

Peak scratch: the restore writes the full logical size — 213 272 202 B for the largest app. Sequential-with-cleanup needs only the largest app; all-at-once needs 774 MB.

Against R-359's numbers. R-359 measured 35.0 s at structure depth and 39.2 s at 100 % read-data on a 140 829 678 B store. Re-measured today through the product's own debug button on a 141 959 062 B store: 40 257 ms at 100 % (15-q5-integrity-run.txt). So:

restore-testing every app on the box (25 s) costs LESS than one weekly integrity check (40 s). Same order of magnitude, and on the cheaper side of it.

Extrapolation — labelled as extrapolation. R-401 records that the check's cost tracks the index and a read-data run's tracks the data. A restore tracks the data too, plus a fixed per-snapshot cost. With 8 apps and the measured ≈ 67 MB/s:

store fixed (8 × ~2 s) data total
today, 774 MB logical 16 s ~9 s 25 s (measured)
10× — 7.7 GB 16 s ~115 s ~2.2 min
100× — 77 GB 16 s ~19 min ~19 min

Why it may not hold. One store, one link, one afternoon. This is DooPlex→Hetzner; a customer's domestic line is the real variable, and restore is the download direction, usually the faster one on such a line. Scratch space becomes the binding constraint long before time does: at 100× the largest app would need ~21 GB of free scratch, and RestoreOffboxScratch's headroom gate would start refusing. Unknown, and what would settle it: a measurement on a store above 10 GB. None exists. This is the same single-data-point hole R-401 already owns.


Q5 — contention

Answer: skip-if-busy stays right, because the operation is seconds and not minutes — but the restore path takes NO single-writer flag at all today, which is a bigger problem than contention.

The configured window on this box, read from the scheduler's own registrations rather than assumed (28-q5-schedule.txt, times are CEST — the guest scheduler runs in local time, not UTC):

job time measured duration
db-dump 02:30 —
tier2-backup 03:30 —
offbox-backup 04:15 2m52s (last run, last_duration)
offsite-abandon-sweep 05:10 —
offsite-integrity 06:00 40.3 s at 100 % depth

A restore-test of all eight apps holds anything it holds for 25 s — one seventh of the nightly off-site backup, and shorter than the check beside it. There is an empty gap from ~04:18 to 06:00. Skip-if-busy remains the right policy, and it is not the "minutes rather than seconds" case the task worried about. The integrityCheckTimeout reasoning transfers unchanged.

The real finding here. offbox_integrity.go:28 states the invariant:

"Every off-site operation takes acquireRunning for exactly that reason."

It does not. grep -rn 'acquireRunning()' finds nine non-test callers; RestoreOffboxScratch (offbox_restore.go:234) is not among them. restore_wizard.go:174 says the same thing independently — "RestoreOffboxScratch never acquires it at all" — and the UI works around it with a separate display flag. So a scheduled, unattended caller built on RestoreOffboxScratch today would run with no single-writer flag, which is precisely the hazard the integrity file's header is shaped around. A comment asserting an invariant with no test pinning it. Filed as R-408.


Q6 — what does a restore-test WRITE to the repository?

Answer: restic restore writes nothing — it does not even take a lock. The product's restore path writes anyway, because it runs restic unlock before every restore. And restic check — the thing whose comment says it never writes — DOES take a lock.

Method. Two observers inside the container, both proven before being believed:

  1. a lock sampler running restic list locks --no-lock every ~4 s (the observer itself never writes);
  2. an argv sampler reading /proc/*/cmdline every 0.2 s, with the repo URL and sftp.command redacted at the source.

Positive control for the lock sampler — it works. The product's own integrity check ran 13:41:28 → 13:42:11:

13:41:30 locks=0
13:41:34 locks=1 ids=81fd4d4200d848466e18cb7a9d8e0d4300d95707c68417449178c3023531eecf
   ... nine consecutive samples, one lock ...
13:42:10 locks=1 ids=81fd…
13:42:14 locks=0

So restic check writes a lock file to the repository. offbox_integrity.go:255 says "It NEVER writes to the repository: check is a read verb, and nothing here prunes, forgets, unlocks or backs up." The three named verbs are correct; the sentence's headline is not. Filed as R-407.

The restore, with the same proven instrument: locks=0 across every sample inside both restore windows (13:40:11/:15/:19 for docmost, 13:42:57/:43:01/:05 for kimai). restic 0.14.0's restore does not lock the repository.

What the product runs, observed argv (redacted), one restore of opengist:

13:44:41  restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> snapshots latest --tag opengist --json
13:44:44  restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> unlock
13:44:46  restic -r sftp:<REPO-REDACTED> -o sftp.command=<REDACTED> restore b5aa8f9b --target … --include …

Three invocations. The middle one is unlockStale (offbox_restore.go:289 → offbox.go:743), which runs unconditionally before every restore and is a delete verb against locks/. With no stale lock present it deletes nothing — but it is a write-capable command on the exact path R-87 wants to run unattended. resticStep's escalation to unlock --remove-all (offbox.go:760-775) fires only on repository is already locked, which a restore cannot provoke by itself now that we know restore takes no lock.

Against §5's constraint — "R-95 still applies: that credential can delete, so a restic restore-test must never be able to write to the repo":

  • The lead in the task was right in direction and wrong in mechanism. The write is not resticStep's escalation; it is unlockStale, one line earlier and unconditional.
  • The constraint IS satisfiable, cheaply, and both mechanisms already exist in 0.14.0 and are unused: a restore-test that passes --no-lock and skips unlockStale writes nothing to the repository at all. That is a design note for whoever builds it — this spike did not fix it, per §7.

Q7 — what would an unattended restore-test catch that the weekly 100 % check does not?

Answer: of the five defects human drills found in the last six days, one. But that is the wrong scoreboard, and the right one has a much better answer.

defect would an unattended scratch-restore have caught it? how / why not
R-353 — a local unit restore reported a bare completion whether it returned a dataset or nothing NO different code path entirely (restore_unit.go). An off-site restore-test never enters it
R-354 — the off-site full restore has no named-volume replay leg: the tar reaches the scratch and is never replayed NO the defect is after the scratch. A test that stops at the scratch sees a correct scratch. Going further means replaying into a live app, which an unattended test must not do
R-356 — the off-site restore refused every app with no data drive (40 of 53 catalogue templates) YES the test calls RestoreOffboxScratch, gets a refusal, and rerr != nil. On this box five of eight apps live on /mnt/sys_drive with no drive — it would have fired on the first night
R-358 — OffboxFullScratchReady asked "non-empty directory", which is what a failed restic run leaves NO needs a restore that fails part-way. On a healthy run the broken and the fixed predicate agree
R-403 — a hollow unit mirrored over a complete one with --delete; 120 082 104 B → 7 036 B, recorded as success NO as filed — but YES for the shape R-403 destroyed a local second-drive copy. What a restore-test sees is the consequence: once a hollow unit is captured off-site, the snapshot is perfectly intact and contains nothing recoverable

One of five. If the question is "does our restore code work", the honest answer is: the drills already answer it, they answer it better, and automating a worse version of it is not worth an evening.

The scoreboard that matters is different, and the weekly check structurally cannot play on it.

restic check --read-data-subset=100% proves that the bytes we stored are the bytes we stored. It cannot tell us we stored the wrong thing.

A hollow recovery unit — no database dump, no volume tar — backs up cleanly, checks cleanly at 100 % depth, restores cleanly, and recovers nothing. R-403 proved that shape is real, on this fleet, nine days ago, measured in bytes. Nothing in the product asks the question today, on any tier, at any cadence. A restore-test is simply the cheapest place to ask it, because the manifest that answers it travels inside the snapshot.


Recommendation to Viktor — three options, with costs

Option A — do not build it. Cost: nothing. What you get: the weekly 100 % check keeps proving the stored bytes; the restore code keeps being proven by your drills. What you lose: nothing that has bitten yet — and the R-403 shape stays invisible until a customer needs the data. This is a defensible answer and it is the one R-87's original framing deserves.

Option B — a scheduled attended drill instead. Cost: one of your evenings, monthly. What you get: everything a person can see, including the R-354 class that stops at the scratch. Honest objection: this is what already happens, and it is what found all five defects. Scheduling it changes nothing except that it now has a date. Low value for the price.

Option C — build the NARROW unattended test: one app per night, rotating, restored to scratch, and checked against its own manifest. RECOMMENDED.

What it does, in one sentence: restore the newest off-site snapshot of one app into the throwaway scratch, assert that every file the unit's manifest.json declares is present, record which snapshot was proved, delete the scratch.

Measured cost, not estimated: ~4 s and ≤ 213 MB of scratch per night (one app), or 25 s for all eight. Off-site traffic: one restore's worth, ≈ the app's size. Less than the weekly integrity check already running beside it. No new metadata, no new reference, no content hashes — it reuses unitCarriesData's existing manifest read.

What it catches: R-356 outright, and the R-403 class — a snapshot that is intact and empty — which nothing else in the product asks about. What it does not catch, stated so nobody expects it to: R-353, R-354, R-358. Those stay drill work.

Three things it must be built with, all established by this spike:

  1. --no-lock, and skip unlockStale — then it writes nothing to the repository and R-95's constraint is honoured for real (Q6).
  2. Take acquireRunning — RestoreOffboxScratch does not, and the whole off-site single-writer story assumes every operation does (Q5, R-408).
  3. Record the SNAPSHOT it proved, not a timestamp — R-87's own row already says this, and R-86 built the per-archive due-ness model to copy.

If you do nothing: the weekly check keeps running and keeps being right about the bytes. The first time a hollow unit reaches the off-site store, nothing will notice, and the discovery will be a customer's restore. That is not a hypothetical shape — it is R-403, measured on 2026-08-31.

I would pick C, scoped exactly as above. It is cheaper than the check beside it, it asks a question nothing else asks, and it needs no invention.

R-87 itself should be RE-SCOPED, not built as written — from "restore-test the restic tier" to "prove the off-site snapshot still contains a recoverable unit". That is your call, so R-87 stays open carrying this verdict.


Register rows opened by this spike

id what
R-405 R-87 was mis-filed by ef6ac6f; corrected + gated (CLOSED same session)
R-406 two unrelated findings share the id R-133 in OPEN-ITEMS.md
R-407 restic check DOES take a repository lock; offbox_integrity.go:255 says it never writes
R-408 RestoreOffboxScratch takes no acquireRunning; offbox_integrity.go:28 asserts every off-site operation does
R-409 the recovery-unit manifest hashes 4 918 B of a 213 231 242 B unit — the data files have no recorded hash

Probes, teardown and side-effects

All three layers, and two of them really are "nothing was left".

layer created removed
PVE host demo-hp /root 6 helper scripts + one 0600 password file all removed; ls | grep returns nothing
guest 9201 /root, /tmp 9 helper/session files + spike-env.sh all removed; grep returns nothing
container /tmp spike-env.sh, two samplers, two logs, two run-flags all removed; /tmp lists empty

Scratch directories: four (docmost, kimai, privatebin, opengist) were created by the restores this spike drove and were removed. Three (bookstack, calibre-web, paperless-ngx) pre-date this session and were left alone — this spike never restored them.

Two deliberate state changes on the box, recorded rather than hidden:

  1. The integrity check I ran to control the lock observer recorded its verdict — the box's last_integrity_check moved to 2026-08-31T13:41:28Z, ok=true, and last_integrity_depth moved from structure to 100%. Due-ness advanced by seven days. That is the product behaving correctly; it is not a repair and it is not damage.
  2. Two logins and four restores appear in the controller log and in the customer-visible restore-op history.

Nothing was written to the off-site repository by hand. No prune, no forget, no unlock issued by me. Every write observed in Q6 was the product's own.

No controller code changed. No golden is owed. The fleet stays on v0.230.0. The one delivery debt that exists — golden_currency_gate.py red because v0.230.0 has no golden — was already red at dddcc80 before this session began, and belongs to the v0.230.0 release, not to this task.


Observations noticed and not acted on

  • paperless-ngx and filebrowser have no off-site snapshot under the tags the inventory uses — paperless returns nothing and paperless-ngx returns one; filebrowser returns none at all. filebrowser is infrastructure, so that may be correct. Not chased.
  • CLOSED-ITEMS.md has two malformed rows — R-399 and R-400 supply two columns where the table declares four, so they render with no Shipped and no Evidence. The new gate prints them as a warning rather than convicting, because an empty state cell is not an open state word.
  • Two rows carry a | inside their body (R-309, R-351), which shifts their own cells. Named as the new gate's first residual hole.
  • The controller image has no ps and no python3 — worth knowing before writing any probe that runs inside it. /proc/*/cmdline is the substitute that works.
  • git push --no-verify was used for this session's Part 1 commit — bypass #7, for the reason R-404 exists: a documentation-only push met golden_currency_gate.py. Recorded here because R-404 counts them.