Commit Graph

1130 Commits

Author SHA1 Message Date
admin 714d5bce09 REPORT: v0.245.0 — the recovery-code ask, red-proofs and the two-box live validation
gates / gates (push) Successful in 15s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 21:20:31 +02:00
admin ad398b60d9 v0.245.0 — R-543: the household is asked for the recovery code, the page says "szunetel" until then
gates / gates (push) Successful in 14s
Off-site backup is ON by default and does not RUN until the household creates its
recovery code. The pause is the zero-knowledge escrow design and is untouched here;
what was missing is that nothing ASKED, while the app-backup page promised the very
copy that had never run.

- a reminder bar on every authenticated page while the off-site tier is configured
  and its escrow is not complete, linking /backup/escrow. It is the R-241 bar, second
  instance: same session-cookie dismissal, back next visit, gone for good when
  escrowed. No second banner system. It hangs off executeTemplate, the single render
  choke point, so it cannot reach only the pages someone remembered.
- the tier-1 file sentence renders by tier3State's own vocabulary instead of the
  app's shape: active -> "vedi", escrow_pending -> "vedene ... szunetel" + the route,
  no copy at all -> says so and names both ways out.
- both fixes red-proofed: the bar test fails on BOTH pages with the hook removed; the
  sentence test quotes the exact v0.244.0 promise when the state is ignored.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 21:02:51 +02:00
admin 2f8ff2414c v0.244.0: the backup page stops promising what it does not hold (R-537/R-538/R-536)
gates / gates (push) Successful in 17s
R-537 — the contents label is now PER TIER. One string computed from the app's
shape was rendered on all three tier rows; a Tier-1 unit has no file-copy step, so
for the four class-A apps it was claiming „Adatok" for files it does not hold.

R-538 — a unit restore REFUSES before anything is touched when the unit cannot
return the app's drive-side files, and names the route that can. It runs before the
stack is stopped because the measured harm included the app's own wastebasket going
unreachable, which still held every byte.

R-536 — „Alkalmazás telepítve" moved from the deploy's acceptance to its completion,
with app_deploy_started and app_deploy_failed as the honest pair.

Each fix red-proofed: seen failing with its own sentence, passing when restored.
Requires hub v0.116.0 for the two new event types. MinAgent unchanged (0.131.0).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-16 16:55:55 +02:00
admin 383a30b3c0 REPORT: v0.243.0 FileBrowser password, per-tier backup page, tier skip, OOM visibility
gates / gates (push) Successful in 14s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 11:37:08 +02:00
admin d3eacbb7cc backup tile: unknown size shows a dash, not 0 B (R-517 follow-up, measured on 9201)
gates / gates (push) Successful in 14s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 10:53:20 +02:00
admin 843b319f35 v0.243.0: FileBrowser generated admin password (R-513); per-tier whole-guest backup truth (R-517); skip absent-storage tiers (R-518); OOM-killed worker visible (R-514)
gates / gates (push) Successful in 14s
MinAgent: 0.131.0

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-15 10:12:08 +02:00
admin 406755fa8f docs: v0.242.0 report, context; R-489 measured limit recorded (row kept open)
gates / gates (push) Successful in 14s
A volume recreated by a unit restore carries no compose label, so the
before/after difference misses it; proven on the scratch guest. Docs only,
no release.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 23:04:48 +02:00
admin d698ce343b controller v0.242.0: a removed app is listed with its kept backup; five small ones (R-487 R-491 R-490 R-489 R-476 R-456)
gates / gates (push) Successful in 14s
R-487: the local backup lists are keyed on the drives, not on what is
deployed — a removed app whose unit was kept is listed with the restore
that reinstalls it, the picker answers for it, and the restore opens the
unit where it sits. R-491: a removal clears the app's update hold.
R-490: /api/system/info reaches the API router and reads the default
storage path. R-489: volumes_removed is the real before/after difference,
[] when none. R-476: a Tier-2 copy is dated by its data, not its manifest.
R-456: the boot-orphan rule is pinned. Every fix red-proofed.
2026-09-13 22:50:18 +02:00
admin c1f62ddae8 REPORT: v0.241.0 delivered by the floor and proven live; R-491 filed
gates / gates (push) Successful in 13s
2026-09-13 21:58:02 +02:00
admin 3e813307cc controller v0.241.0: a bind-data app leans on off-site before its own unit; the hold names what the copy holds (R-479)
gates / gates (push) Successful in 13s
Operator ruling 2026-09-13. An app with classified binds walks second
drive -> off-site -> own unit (its unit holds no files); volume apps keep
2 -> 1 -> 3. RestoreHold.CopyHolds records what the chosen copy holds and
the sentence ends with it; older holds keep their tier-only sentence.
Tests on both halves; red-proof: a layout-blind order fails the bind case.
2026-09-13 21:47:33 +02:00
admin 3013a1cc93 rules: unprompted-work.md declares unconditional: true (the instructions gate requires a scope or that declaration)
gates / gates (push) Successful in 15s
2026-09-13 21:31:31 +02:00
admin 3b15ce101f rules: unprompted-work.md — the rules for goal and nightly sessions, byte-identical in all three repos 2026-09-13 21:29:43 +02:00
admin 21a9d35f70 REPORT: R-490 filed — /api/system/info shadowed, monitoring memory card never renders (from the R-465 audit)
gates / gates (push) Successful in 14s
2026-09-13 19:47:00 +02:00
admin 24d7c54c79 REPORT: v0.240.0 delivered by the floor and proven live; R-488/R-489 filed
gates / gates (push) Successful in 14s
2026-09-13 19:38:03 +02:00
admin bdcbd50b42 controller v0.240.0: seven defects from the any-tier proof and the first nightly rotation
gates / gates (push) Successful in 13s
R-486 (P1): removing an app with its backups KEPT keeps its Tier-2 record,
so the second-drive restore is no longer refused over an intact mirror.
R-484: postgis/pgvector/timescaledb images are Postgres (logical dumps).
R-485: the backup card sizes the recovery unit and the mirror(s).
R-480: a held update's sentence leaves the card once the hold is lifted.
R-477: the update's off-site lookup is one snapshots call, no stats.
R-478: a copy older than this install's deploy does not count.
R-474: "delete backups" deletes the unit, the mirror(s) and the prefs.

Tests and red-proofs per row; evidence in felhom.eu
documentation/audits/v0240-2026-09-13/ and nightly-2026-09-13-adventurelog/.
2026-09-13 19:26:50 +02:00
admin 0e3d831030 REPORT: v0.239.0 any-tier update and MinAgent header gate, proven live; R-477..R-480 filed
gates / gates (push) Successful in 13s
2026-09-13 17:54:27 +02:00
admin b93c1543da controller v0.239.0: any backup tier lets an app update (R-475)
gates / gates (push) Successful in 14s
Operator ruling 2026-09-13. The update precondition walks Tier 2, Tier 1
(own recovery unit, "helyi") and Tier 3 (off-site, 15 s bound; unreachable
counts as absent with a WARN) and leans on the first FRESH copy; the
backup_max_age rule applies to whichever tier is chosen. No copy anywhere:
back up first. Refused only when nothing exists and no backup can be taken.
RunAppBackupNow tolerates a Tier-2 failure (WARN) and marks the captured
unit proven current. The hold names the tier (második meghajtó / saját
meghajtó / távoli mentés) and the date; pre-v0.239.0 holds keep their text.
A successful off-site restore now lifts an update hold. The backups page
still uses Tier2UnitRestorePoint unchanged.

Scenarios G-M tested; red-proofs M, L, the tail and the off-site clear in
felhom.eu documentation/audits/rulings-r472-r475-2026-09-13/.
2026-09-13 17:16:24 +02:00
admin f946b0d0ca gates: the newest release header must state its MinAgent; backfill v0.233.0-v0.236.0 (R-470, R-472)
gates / gates (push) Successful in 12s
From hub v0.112.0 a floor above the golden is served only with a declared
MinAgent, read from this CHANGELOG's newest `## vX.Y.Z` header. New fast,
blocking gate minagent_header_gate.py: that block must contain a line
starting `**MinAgent: X.Y.Z**`; prose, a code span or a non-bold mention
does not count (decoy test; red-proof F: relaxing the match to a body
mention fails test_decoy_body_mention_is_not_the_line).

Backfill: v0.233.0, v0.234.0, v0.235.0 and v0.236.0 carried no MinAgent
line. Each now reads `**MinAgent: 0.129.0** (unchanged)`, the value of the
nearest earlier header (v0.232.0). Proof it did not move: zero commits
under controller/internal/agentapi since 2026-09-01 (last: bb50e12,
2026-08-14), and the highest featureMinAgent entry is 0.129.0.
2026-09-13 17:09:33 +02:00
admin cf8a371c39 REPORT: slice 4 (v0.237.0–v0.238.1) proven live — A, B, E, F, H and the restore walk
gates / gates (push) Successful in 13s
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 12:30:47 +02:00
admin cbcca03061 v0.238.1: the nightly backup leaves an app alone WHILE it is being updated, not only once it is held (slice 4 follow-up)
gates / gates (push) Successful in 13s
Found live in v0.238.0 Scenario F on demo-hp: during an update's 5-minute health wait the app is not
yet held, and the periodic recovery-unit capture at 10:17:09 wrote the never-started definition
(alpine:3.20) into its PRIMARY unit, 53 s before the hold landed. The Tier-2 mirror the hold names
survived only because Tier 2 runs daily; a nightly Tier 2 inside a verify window would have mirrored
the broken definition over the copy the customer is told to restore from.

backup.Manager.isHeld — consulted by the capture sweep, the Tier-2 run and the volume dump — is now
also true while a guarded update is moving the app, via SetUpdatingCheck wired in main.go to
stacks.Manager.IsUpdating. Test with positive control + red-proof; wiring pinned.

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 12:25:09 +02:00
admin 129201abab v0.238.0: the page follows the update, and a held app offers no way to start it (update arc slice 4 Part 4)
gates / gates (push) Successful in 13s
No behaviour change on the box — the surface only.

- Frissítés follows the job: the button shows the phase label (polling GET /api/stacks/{name}
  every 3 s) and the page reloads when updating goes false.
- An updating card offers no lifecycle button; a held card (failed update OR failed restore) shows
  the hold sentence with a Mentések link and nothing that would start it; a failed update that held
  nothing shows its sentence above the buttons. app_info shows the same three notices.
- The updating/held checks run BEFORE isOperational, which counts `restarting` as operational — how
  the 2026-09-01 spike saw a green Frissítés beside a crash loop. Pinned with StateRestarting
  fixtures; red-proofed by moving the checks after it (both tests fail).
- No new CSS, no version number.

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 12:07:35 +02:00
admin 0d402f711d v0.237.0: the Update button takes a backup first, and tells the truth (update arc slice 4 — R-448, R-443, R-439)
gates / gates (push) Successful in 13s
POST /api/stacks/{name}/update is now a guarded job answering 202:
cheap refusals (hold — R-439, busy, migration, deploying, memory via the
deploy's own memoryVerdict, a fixed 2 GB disk floor, and no restorable
Tier-2 copy) → backup-first when the proven copy is older than
update.backup_max_age (24h) → safety dump BEFORE the pin moves → pin →
pull (failure puts the pin back) → up → health (.felhom.yml check or 60 s
settle, update.health_timeout 5m). Not healthy → the app is stopped and
HELD (RestoreHold reason update_failed, same store and gate as R-379) and
the page names the backup to restore from; the pin stays. Success is only
ever update_phase=done after health (R-443). UpdateStack is deleted.

The restorable-unit predicate is EXTRACTED to backup.Tier2UnitRestorePoint
and shared with the backups page (row pinned unchanged). The copy is aged
by the last successful Tier-2 copy, not the manifest created_at — measured
on demo-hp that created_at moves only on definition changes.

Crash safety: update-journal.json before each phase; RecoverUpdates before
the boot sweep, ResumeInterruptedUpdates after the guards are wired.

Three unattended start paths ignored a hold and now honour it: the
drive-return gate (restart + boot recreate) and the nightly volume dump.
The nightly capture and Tier-2 run skip held apps so the restore point
survives. No automatic rollback — measured per-app; route back = restore.

Tests A–H across stacks/backup/api/web/cmd; six red-proofs seen to fail.

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 11:41:31 +02:00
admin 1552716722 REPORT: name the felhom.eu --no-verify push (golden-currency only) and the closed-register gate fix
gates / gates (push) Successful in 13s
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 09:07:26 +02:00
admin 7410f1c841 REPORT: v0.236.0 (R-442) proven live on demo-hp — data gone, refusal shown, SSD app not refused
gates / gates (push) Failing after 14s
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 09:05:34 +02:00
admin 42a73e667a v0.236.0: "delete my data too" deletes the data, or says that it could not (R-442)
gates / gates (push) Successful in 13s
Removal resolves the drive from the app's own app.yaml HDD_PATH (the 07 ~L437
rule), never the global cfg.Paths.HDDPath which no box sets. A data removal
that cannot be resolved, or whose drive is absent, is refused with a typed
RemoveRefusedError -> 409 + exact Hungarian sentence, before compose down, and
the app is kept. SSD app -> hdd_paths_removed: [] never null; missing folders
stated; backup-path refusals reach the response.

15 tests, two red-proofs run (pre-fix fallback -> C fails with err=nil and the
handler 200s; "no drive refuses" -> D fails).

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 08:52:25 +02:00
admin bab82c471e REPORT: v0.235.0 proven live, and the one thing the task did not anticipate
gates / gates (push) Successful in 13s
Two real catalog pushes travelling the real 15-minute cycle. The non-image change
reached the pinned app; the image change did not; and the restart that used to
take 18.3 seconds and pull a new image took 0.1 seconds, did not recreate the
container, and pulled nothing. The Update button still moves the version, with the
pin advancing 17 seconds before the pull. The frozen app read 'Frissites elerheto
- 56 napja' while the other eight read Naprakesz. Teardown returned the container
to the baseline digest byte for byte.

The correction at the top: the task's Scenario B says to freeze to the stored
definition and says nothing about keeping that store current. The store is written
when the PIN is written, so a fix delivered afterwards - Scenario A's own case -
lands in the live file and not in the store, and the first freeze reverts it. Seen
live at 08:20:29Z. Fixed in the same run; the observation is kept as its evidence.

Also named: the syncer now imports the stacks PACKAGE for two pure symbols rather
than duplicating a compose parser, which honours the task's intent and not its
letter; and syncer.Start() had to move after adoption, which the task did not say.

Three red-proofs run and reverted. A fourth defect was caught by a test: the
syncer would have written an empty compose file over a live app.
2026-09-06 10:39:13 +02:00
admin 2a56f557d0 v0.235.0: a delivered fix must also refresh the stored definition
gates / gates (push) Successful in 14s
Found by the LIVE validation on demo-hp, not by review. Scenario A passed - a
non-image catalog change reached the pinned app on the real 15-minute cycle - and
that is exactly what exposed the gap: the stored applied-compose.yml is written
when the PIN is written, so the fix landed in the live compose file and not in the
store. The first time the catalog then moved a version, the freeze would have
rendered the pre-fix definition and reverted every fix delivered since - silently
undoing the half of the operator's ruling that says fixes keep flowing.

The equal-images branch now refreshes the store as it delivers. The images cannot
move in that branch by construction, so no version moves and no intent is
rewritten. RenderPlan gains StackDir so the syncer can write it.

TestFixRefreshesTheStoredDefinition asserts both halves: the fix reaches the
store, and it survives the freeze that follows.
2026-09-06 10:04:21 +02:00
admin 8a0e0a59ad v0.235.0: freeze the version, keep the fixes flowing (operator ruling 2026-09-06)
gates / gates (push) Successful in 12s
Slice 3. R-447 was BLOCKED because R-438 established that RestartStack's use of
up -d to pick up template changes was CHOSEN and written down in its own comment.
The operator ruled Option 1, and this implements it.

The rule: while the catalog offers the same version you run, its fixes flow to
you; the moment it moves to a newer version you are frozen until you update.

NOTHING was added to any of the thirteen compose up -d call sites. Most of them
are repairs - the boot reconciler, the drive-return gate, the app-stop guard -
and a repair path that refuses to repair leaves a customer's app down, which is
worse than the problem. They are made safe by removing the reason.

app.yaml gains pinned_images: what the app is SUPPOSED to run. It is NOT
installed_images, which is an observation; letting a reading become a deployment
is the R-166 category error one field over. Four writers, each also storing the
exact definition as applied-compose.yml. UpdateStack advances the pin and
re-renders BEFORE the pull, because pull and up -d act on the file on disk, and a
pin set afterwards would pull the frozen version and report success.

The syncer renders instead of copying, through one nil-safe seam. Catalog images
equal the pin -> verbatim, so fixes and self-healing both survive; they differ ->
the WHOLE stored definition, never a substitution of refs into a newer template
(wger 2.6 needs a DB config the older template cannot supply). This is
deliberately not 'skip deployed apps', which was option B and was rejected.

AdoptPins runs once at boot after the backfill, files only, and skips loudly
rather than inventing a pin. syncer.Start() moved to after it: the initial sync
would otherwise run while every app was unpinned and overwrite a deployed app's
version once per boot.

THE BADGE HAD TO CHANGE OR SLICE 2 WOULD HAVE INVERTED SILENTLY. TemplateImages
reads the LIVE compose file, which is now the frozen one, so the comparison would
have answered Naprakesz on exactly the apps that are behind - with every test
green, because the new field has the same type. It now reads CatalogImages.

+16 tests (1729 -> 1745), 28 packages green. Three red-proofs run and reverted.
A test also caught the syncer writing an empty compose file over a live app.
2026-09-06 09:45:34 +02:00
admin 998aa31958 REPORT: v0.234.0 — the backfill proven live, and a test of mine that went red overnight
gates / gates (push) Successful in 12s
Adds section 6b (the startup backfill on both demo boxes: all nine apps already
had records by the time it was ready, so the pre-0.233.0 shape had to be
recreated on demo-hp - said plainly rather than papered over; two seeded with
digests matching independently-read ground truth, seven untouched, nine badged)
and 6c (the render test hardcoded a date and an age, was green the day it was
written and red the next morning, now derived; R-457 names six candidate files).

Also records what was deliberately NOT staged live: the backfill's refusal of a
partial observation needs a degraded app, and manufacturing one risks the false
customer email class that already cost 61 mails.
2026-09-03 12:02:30 +02:00
admin 38d28b5b62 v0.234.0: seed installed_images at startup, so the label appears on an app nobody touched
gates / gates (push) Successful in 13s
The operator looked at demo-felhom the morning after v0.233.0 and found OpenGist
- up 15 hours, running exactly the catalog pin - showing no badge at all.
v0.233.0 wrote the record only from the four bring-up paths, so an app nobody
restarts carried no record indefinitely. On a quiet box that is every app, which
is the box we most want to see. The known limitation WAS the feature not working.

BackfillInstalledImages runs once at startup, beside BackfillDesiredState and
before the boot reconciler. It READS containers: starts nothing, restarts
nothing, writes no compose file. It never overwrites an existing record.

And it REFUSES to seed a partial observation, which is why this is not a
three-line loop: the badge reads a service-count mismatch as BEHIND, so seeding a
degraded app from what is visible would render 'Frissites elerheto' over an app
that is perfectly current. The bring-up paths may write a partial because they
follow a successful up -d where a gap is real news; a backfill meets any state.
Same data, two writers, two admission rules - deliberately.

Also fixes a calendar bomb of mine: the render test hardcoded catalog_since and
the string '46 napja', but the render path reads time.Now(), so it was green on
the day it was written and red the next morning. Now derived. Filed as R-457
with six other candidate files named as unchecked, not accused.

+5 tests (1724 -> 1729), 28 packages green. Red-proof of the partial guard run
and reverted; the wiring and its ORDER pinned by an AST walk.
2026-09-03 11:56:43 +02:00
admin 32da46cd64 REPORT: the mirror base images ARE byte-identical to Docker Hub - measured, not assumed
gates / gates (push) Successful in 13s
The throttle cleared 40 minutes later, so the check was run instead of left as a
one-command IOU. 'docker pull docker.io/library/<img>' answered 'Image is up to
date' for both bases - Docker Hub's own manifest resolved to the images already
local, the ones the mirror supplied and 0.233.0 was built from. The two manifest
indexes are also identical between registries.

Also corrected in passing: the two sha256 values recorded earlier are local IMAGE
IDs, not manifest-list digests. I conflated them once during this very check, so
the report now says which is which.
2026-09-02 21:03:24 +02:00
admin 50312b63bc REPORT: the badge is proven live too, and the stale-password claim was mine to retract
gates / gates (push) Successful in 13s
Naprakesz on /stacks (x2) and /apps/bookstack; NO badge at all on /apps/docmost,
a deployed app with no record - absent is UNKNOWN, not current; and 'Frissites
elerheto - 52 napja' on both surfaces, the age real arithmetic on bentopdf's
catalog_since. Staged by editing one compose tag with no restart and no up -d,
reverted byte-identically. ASCII fragments with positive and negative controls.

Section 7 now opens with the correction rather than burying it: I reported the
vaulted password as stale on both boxes; it was fine, and I had stripped only
double quotes from a single-quoted value. The keeper is that I read the
controller's 'Failed login' past what it discriminates.
2026-09-02 20:49:01 +02:00
admin 2fb558552a REPORT: v0.233.0 live — the record is proven on metal, the badge render is not, and why
gates / gates (push) Successful in 13s
The record: proven on demo-hp through the boot reconciler (a real production
caller, no hand-set state) on a single-service AND a multi-service app, with all
three digests matching ground truth read independently beforehand.

The badge render: NOT validated. The vaulted dashboard password is stale on BOTH
demo controllers; the five attempts are listed rather than summarised, and the
controller's own log is the discriminator that says wrong password, not wrong
host header. Filed as R-453 and raised in STATUS.md item 9.

Also recorded: the build was blocked by a Docker Hub 429 and the base images came
from Google's Hub mirror, with both digests written down so the identity check is
one command when the throttle clears - KNOWN, not measured here.
2026-09-02 20:37:01 +02:00
admin 8025304acc v0.233.0: record what each compose service actually installed, and badge whether it is current
gates / gates (push) Successful in 12s
Update arc slices 1 and 2. NEITHER CHANGES ANY BEHAVIOUR — no new endpoint, no
auto-update, the three lifecycle buttons byte-identical.

Slice 1 — app.yaml gains installed_images, keyed by compose SERVICE name, each
entry carrying ref + repo digest + first-seen timestamp. Written by
Manager.recordInstalledImages after a successful compose up from StartStack,
RestartStack, UpdateStack and runComposeDeploy. Read from the CONTAINER, never
from docker-compose.yml: the syncer overwrites a deployed app's compose on a
15-minute cycle and the two disagreed for 25 minutes in the spike's own
measurement. A failed write NEVER refuses the action - the deliberate opposite
of SetDesiredState, because this is an observation and that is an intent. Not
called from StartStackServices (the R-47 DB-only window). Its own docker seam
with a context and a 30s timeout, which neither existing exec helper has.

Slice 2 — .felhom.yml gains optional catalog_since; web.updateBadge compares the
recorded ref per service against what the current template pins and returns a
*MetaBadge through the EXISTING meta_badge partial. No new markup, no new CSS.
NO RECORD RENDERS NOTHING: absent means unknown and never means current. No
version number reaches the customer and no registry is queried.

Known limitation, filed not hidden: 23 catalog pins float, so those apps can read
Naprakesz when the image behind the tag has moved.

+17 tests (1707 -> 1724), 28 packages green. Wiring proven through a real
RestartStack plus an AST walk of the four call sites. Three companion red-proofs
run and reverted.
2026-09-02 20:18:01 +02:00
admin 960d29b061 REPORT: the CI outcome - four red runs, two real causes, all now green
gates / gates (push) Successful in 12s
Adds section 11. The sharpest result of the whole sweep is in it: my own decoy-coverage gate
identified a repository by its DIRECTORY NAME and went blind the first time CI ran it, because the
act-runner checks out into a folder called hostexecutor. The gate written that morning to catch
name-for-fact was matching a name, in the first ten lines of its own main loop (R-428).

The other cause was my push ordering - two repos citing R-421 pushed before felhom.eu carried the
row - which instructions_gate convicted exactly as designed.
2026-09-01 12:48:50 +02:00
admin 22983885f1 re-run CI against a register that now carries R-421
gates / gates (push) Successful in 13s
The earlier run convicted correctly: instructions_gate found this repo citing R-421 while
felhom.eu's OPEN-ITEMS.md did not yet have the row. My ordering, not the gate's fault - the register
lives in felhom.eu, so a repo citing a new row must be pushed after it.
2026-09-01 12:45:54 +02:00
admin 670def61b4 REPORT: the decoy sweep - 29 gates, 16 fooled, 10 fixed, 6 honestly untested
gates / gates (push) Failing after 14s
Opens with the survey table. Records the three numbers, every live hole with its decoy and row, the
six gates no plausible decoy could be built for, the meta-gate's 20-name exemption list, and which of
the five defining rows actually closed (R-419; R-378 explicitly did NOT).

Includes my own mistakes by name - five decoys withdrawn as illegitimate, a 36-vs-44 arithmetic
artefact I announced before checking, an rc==0 read as a hole for a gate where rc is not the
question, a bash heredoc that ate my backticks, and an R-419 fix that was too strict and rejected
genuine markers until its own gate convicted this very report.
2026-09-01 12:42:30 +02:00
admin 681cc663ef decoy sweep: eight holes in this repo's gates, all measured, all fixed (R-421)
gates / gates (push) Failing after 13s
Every gate was DECOYED - the label constructed without the fact, the gate run, the verdict recorded.
No verdict here was reached by reading, because reading is exactly how the five prior instances hid.

SCOPE IS A FACT TOO, and it was the big one. Six gates decided what to look at with os.listdir - one
directory level. Every one was green AND CORRECT, because no template subdirectory exists today; every
one would have gone blind the moment anyone added templates/partials/, which is an ordinary act. A
single planted file carrying an emoji, a native confirm(), hand-rolled row markup, a dangling JS id
reference, a templated secret and an unregistered retrieval promise passed all six.

THE CONTROL IS WHAT MAKES THAT A MEASUREMENT: mojibake and docker-v already used os.walk, saw the
identical planted file, and convicted. So the cause was the listing, not the decoy.

COMMENTS ARE NOT CODE, AND COMMENTS ARE NOT CONTROLS. debug-routes matched `case subpath == "x"` in
raw text, so a case left in a commented-out block counted as a live handler - which is R-400's
original defect (seven dead controls on the page an operator opens when something is already wrong)
reached through the one door its own gate could not see. app-row-dedup's MUST_USE check had the same
shape: a commented-out {{template "app_list_row"}} satisfied it.

Stripping is deliberately crude in debug_route_gate, and that is correct there: its own docstring
insists on ten lines that cannot rot. A // inside a string literal truncates that line, which can
only ever HIDE a reference, never invent one - it fails in the safe direction.

NOT FIXED, and left open with its decoy rather than quietly patched: R-425, offbox-rename scans a
fixed three-entry FILES list, so banned NAS branding in a NEW offbox template passes. The scope was
correct when written and silently narrows every time the feature grows a file.

test_gate_decoys.py holds 10 decoys and declares COVERS, which felhom.eu's new decoy-coverage gate
AST-parses - a substring search for coverage would be the very shape this sweep exists to find.

No Go code. No version bump. No image. No golden owed.
Survey: felhom.eu/documentation/audits/AUDIT-gate-decoys-2026-09-01.md
2026-09-01 12:39:17 +02:00
admin 3db62fc6b2 REPORT: the CI step is MEASURED, not assumed - job 481 shows the payload works
gates / gates (push) Successful in 12s
Gitea's act-runner does populate GITHUB_EVENT_PATH with a commits array carrying per-file lists.
Job 481 read 3 commits / 12 distinct paths, classified CODE (5 document, 7 code), and ran the gates
with --scope=code. Green.

Still not observed: the docs branch in CI, and an advisory in a CI log - the second additionally
needs a golden debt to exist at that moment. Neither is being arranged artificially.
2026-09-01 12:03:28 +02:00
admin 13dd00bcff gitignore Python bytecode - the new gate test writes controller/scripts/__pycache__
gates / gates (push) Successful in 13s
test_golden_notice.py imports controller_gates.py to assert the golden-notice is registered
non-blocking, which makes CPython write bytecode beside the scripts. felhom.eu already ignored it;
this repo did not, so it would have reappeared on every test run and been one careless 'git add'
away from being committed.
2026-09-01 12:01:54 +02:00
admin a4444088ad R-404: the golden NOTICE, in the repo where the debt is created - NOT A RELEASE
gates / gates (push) Successful in 12s
No version heading on purpose. No Go code, no image, no version bump; giving this one would create
the exact golden debt the change is about.

Until today this repo - where a release actually happens - had NO golden-currency check at all,
while felhom.eu ran one on every push including documents-only ones that can neither create the
debt nor clear it. The person who could act heard nothing; the person who could not act was
blocked, thirteen --no-verify uses' worth.

golden_notice.py is ADVISORY IN EVERY CASE, and that is the only correct behaviour rather than
timidity: at the moment a release is committed the golden legitimately does not exist yet, so
blocking there would refuse the commit that STARTS the process - and blocking later is the mistake
being undone.

NO SECOND IMPLEMENTATION: it IMPORTS felhom.eu/scripts/golden_currency_gate.py and calls that
gate's own released_versions()/newest_baked(), so it is the same comparison read in the other
direction. Cross-repo shape copied from instructions_gate.py; never a copy of the script, because a
copy recreates the drift these gates exist to detect. An absent sibling clone is INCONCLUSIVE and
silent about currency - it never guesses.

controller_gates.py GAINED A FIFTH `blocking` FIELD. It could not express a reporting-only gate at
all before: every registered gate's non-zero exit failed the run, so the only way to add a notice
was to give it the power to refuse a push. The capability was added rather than the notice
compromised (R-420). False for exactly one gate, and test_golden_notice.py asserts it stays one.

Tests N1-N4 with a positive control that every other gate is still blocking. RED-PROOF RUN: making
the debt branch return 1 fails N1 - in production that would refuse the commit that starts a
release.
2026-09-01 12:01:26 +02:00
admin a017367f9f REPORT: disclose the five red CI runs, the pre-push bypass, and two instruction defects
gates / gates (push) Successful in 12s
2026-09-01 10:58:06 +02:00
admin 4e13fdaaa0 REPORT: v0.232.0 - the determination, the walk's three extra findings, the golden
gates / gates (push) Successful in 13s
2026-09-01 10:50:52 +02:00
admin 62c6a8a98a docs(v0.232.0): CHANGELOG, three CONTEXT rulings, README (R-411/408/407, R-414, R-412a)
gates / gates (push) Successful in 14s
2026-09-01 10:36:52 +02:00
admin 8b55de734c R-414: the fallback scratch must also be DELETABLE - caught by live validation
gates / gates (push) Successful in 12s
The system-data fallback resolved a scratch fine and removeProofScratch then refused to
delete it: its accepted-roots list is built from REGISTERED drives, and a driveless box has
none. Observed on demo-felhom: 'refusing to remove ... it is not inside a proof root', with
the copy still on disk. Every nightly proof would have left one behind, growing forever, on
exactly the boxes the fallback exists for.

My defect, introduced with the fallback in the same session. The unit tests missed it
because every one of them registers a drive; the new pair deliberately does not, and the
second asserts the guard still REFUSES a path outside every proof root, so the fix is not a
widening into uselessness.
2026-09-01 10:29:08 +02:00
admin fcef8e069c one writer at a time, and a check that can run (R-411, R-408, R-407, R-414, R-412a)
gates / gates (push) Successful in 12s
THE WALK FOUND THREE MORE ENTRY POINTS THAN THE REPORT DID. R-411 named one missing
acquireRunning. Fixing it and then pinning the invariant with an AST walk surfaced FOUR in
total, all of which issued restic commands with no flag:

  RestoreOffboxScratch      - the reported one
  OffboxRestorePrepareFull  - the SECOND request in the customer's own two-step full-restore
                              flow, and the one that actually shells `restic stats`. The UI
                              reaches it FIRST, so flagging only the restore would have left
                              the collision reachable by the ordinary path.
  RestoreSharesScratch      - R-411's exact shape on the shares tier: unlockStale + resticStep,
                              a live web caller, and its sibling PlaceSharesRestore has always
                              taken the flag.
  RestoreOffbox             - no production caller today, but the same dangerous pattern.
                              Flagged rather than left for a future caller to inherit.

OffsiteInventoryList is REGISTERED EXEMPT with its reason: it issues only `restic snapshots
--json`, measured on demo-hp 2026-08-31 not to take a lock, and flagging it would make
browsing a page refuse during a backup for no safety gain.

THE REAL DELIVERABLE IS THE WALK, not the acquire. offbox_integrity.go:28 asserted "Every
off-site operation takes acquireRunning" since v0.227.0, nothing checked it, and it was false
for months - the ninth instance of this project's most-repeated class. The walk is an AST
pass, not strings.Contains, because a commented-out call still contains the string.
Red-proofed twice: removing the acquire fails it naming RestoreOffboxScratch; an
unregistered fake entry point fails it naming the fake.

R-407: "It NEVER writes to the repository" corrected in place, not deleted (R-360's rule).
`check` takes a lock - and so does `restic stats`, which is the fact nobody had and the one
that made R-411 possible. Both recorded where the next reader will meet them.

R-414: the proof could not run at all on a box with no registered drive. Part 2.1's
determination came out as neither "missed" nor "deliberate": R-356's own test comments say
the scratch resolver "still resolves ... only the DESTINATION moves", so it was OUT OF SCOPE,
and it was never ruled out on state-only grounds - the one comment about a systemDataPath
fallback belonged to PlaceOffsiteRestore, concerned bulk USERDATA, and R-356 overruled even
that. So 07 section 6.3's rule applies and now has a fourth consumer.

The fallback is SCOPED, because the two callers ask different questions and one predicate
answering both is the R-356 defect itself: a UNIT-ONLY restore may fall back to the system
data path (07 section 7 records as FACT that a driveless app's unit already lives there
indefinitely, and that the same-device placement is intended); a FULL restore keeps today's
refusal, because it pulls bulk userdata onto a state-only tier.

And the silence ends either way: a proof that cannot start now records ProofResultCannotRun
rather than an Err, so last_proof_result is never ABSENT - absent already means "controller
too old", and a second meaning on the same field is the StatsKnown trap one level up. It is
recorded WITHOUT advancing per-snapshot due-ness, so the app stays retryable once a drive is
registered.

R-412 leg 1: a per-app push whose unit carried no dump and no tar now says so, at WARN.
Wording only - no guard, and the capture is untouched (08 section 8.2). Leg 2 stays OPEN.

16 new tests, 1689 -> 1705. Full suite 28 packages rc=0, all 13 controller gates OK.
Red-proofs run and reverted byte-identical for A3/B1 (twice), C1 and D1.
2026-09-01 10:19:36 +02:00
admin 9aea86cd48 REPORT: CI verdicts by run id - three controller runs green, felhom.eu 288 red on the declared golden debt only (12 of 13 gates OK)
gates / gates (push) Successful in 11s
2026-08-31 21:33:03 +02:00
admin 79f853fbbd docs: v0.231.0 REPORT + README + CONTEXT rulings + REUSE entries (R-87)
gates / gates (push) Successful in 12s
2026-08-31 21:31:04 +02:00
admin 303129e3af v0.231.0: the off-site proof gets a by-hand trigger, like its integrity sibling (R-87)
gates / gates (push) Successful in 12s
Without it the only way to see the job work is to wait for 05:30, which makes live
validation and any future diagnosis a next-day exercise. Same function as the scheduled
job - no second code path.

ONE deliberate difference from the integrity button: due-ness is NOT bypassed. There,
forcing means "check the store again", which is always answerable. Here due-ness IS the
target selection - an app is due when its newest snapshot has not been proved - so
ignoring it would mean inventing a second way to choose an app, exactly what having one
function prevents. When nothing is due the button says so, honestly.

Every other guard intact, including the single-writer flag: a hand-run during a backup
SKIPS exactly as the scheduled one would.

POST /api/debug/backup/offsite-proof, button beside "Restic integritas" on the debug page.
debug_route_gate pairs the two, so a button with no dispatch (R-400's shape) cannot ship.
2026-08-31 21:09:09 +02:00
admin e43b5ec07d v0.231.0 - the box proves its own off-site copy still holds something (R-87)
gates / gates (push) Successful in 11s
R-87 re-scoped by its own spike and built as Option C. MinAgent 0.129.0 unchanged.

THE QUESTION NOTHING ASKED. The weekly check proves the stored bytes are the bytes we
stored; it cannot tell us we stored the WRONG thing. A hollow recovery unit backs up
cleanly, checks cleanly at 100 percent depth, restores cleanly and gives the customer
nothing back - measured on demo-hp 2026-08-31, 120082104 B to 7036 B in one nightly run
recorded as a success (R-403). No tier and no cadence asked it. Now offsite-proof does,
nightly, on one app.

IT DOES NOT prove a restore puts data back into a running app. That stays drill work and
07 section 8 matrix row 4 is NOT moved.

THE ACCEPTANCE RULE HAS TWO PARTS AND THE OBVIOUS ONE IS A TRAP. "Check the unit against
its own packing list" PASSES a hollow unit, because a hollow unit declares nothing. So:
(1) everything declared is present, AND (2) the manifest declares what the app is supposed
to have. Part 2 is the whole value. RED-PROOFED: the naive rule makes the hollow-unit test
read verdict "pass".

THE EXPECTATION COMES FROM INSIDE THE UNIT, never the live box - the snapshot may predate
the app's shape, and GetDockerVolumes describes the running app. Database half is
DBServiceNames, the same discriminator RestoreFromRecoveryUnit uses. Volume half is
ParseComposeNamedVolumes as an EXISTENCE check, not a name match: tars are
<project>_<volume>.tar and ResolveDockerVolumeNames derives the project from the compose
file's parent dir, which inside a unit is the literal string "compose". Measured on all
eight real units on demo-hp the counts match exactly and the naming held every time - but
"held on eight" is not "derivable" (R-355). Half a rule that is true beats a whole rule
that is invented.

THREE OUTCOMES: pass, fail (readable and empty), cannot judge. An app that legitimately
has neither a database nor volumes PASSES. RED-PROOFED: alarming on any empty unit makes
that test read verdict "fail".

IT NEVER WRITES TO THE REPOSITORY and that is asserted on the ARGV as a non-effect:
--no-lock, no unlockStale, and m.runner() rather than resticStep so the unlock --remove-all
escalation is unreachable. RED-PROOFED: routing it the customer path's way makes the test
fail on "unlock" appearing in the argv.

IT TAKES acquireRunning ITSELF and skips rather than waits, because RestoreOffboxScratch
does not take it (R-408) while offbox_integrity.go states that invariant as universal.

DUE-NESS IS PER SNAPSHOT (R-86's model), never per clock. RED-PROOFED: recording a
timestamp fails the stored-value test AND breaks the rotation - night 2 re-picks night 1's
app.

ITS SCRATCH IS A SEPARATE ROOT (backups/offsite-proof) and that is a safety decision, not
tidiness: the job deletes its copy on every path, and sharing backups/offsite-restore/<app>
would mean a nightly background job deleting the verification copy a CUSTOMER is looking
at. It is also invisible to placement, so a proof copy can never be pushed into a live app.

SHARED RATHER THAN FORKED: offboxScratchDirIn parameterises the scratch resolver on its
ROOT builder, and unitOnlyHeadroom extracts the free-space gate, so the customer path and
the proof refuse at the same floor with the same Hungarian sentence. RestoreOffboxScratch's
behaviour is unchanged.

NEW EVENT offsite_proof_empty, severity error, operator-only - deliberately NOT
backup_integrity_failed, whose hub template says the store is DAMAGED. Here the store is
sound and the content is absent: different cause, different action. The hub half shipped
FIRST, in felhom.eu 1aeaa30 (hub v0.110.0, live and verified), because an unallowlisted
type is 400'd and vanishes.

33 new tests, all groups green; full suite 1689 tests, 28 packages, rc=0. All 13 controller
gates OK. Five red-proofs run and recorded in REPORT.md.

A golden carrying 0.231.0 is OWED - the fleet is on 0.230.0. Viktor's call (R-242).
2026-08-31 20:55:34 +02:00