Commit Graph

249 Commits

Author SHA1 Message Date
admin 300d7e87d7 REPORT: R-359/R-397 shipped and validated; the measurement changed what it is worth
gates / gates (push) Failing after 12s
The headline is not the feature, it is what measuring it revealed: THE STRUCTURE
CHECK THAT SHIPS ON DOES NOT CATCH SILENT CORRUPTION. A pack corrupted without a
size change returned `no errors were found`, exit 0. Only --read-data caught it.
So R-399 is not merely a bandwidth question -- at the shipped default a class of
damage is not checked at all.

The three numbers R-399 needed are MEASURED, not estimated: store 134.3 MB / 67
snapshots; structure check 35.0 s; and the full curve 10% 35.9 s, 50% 37.3 s,
100% 39.2 s. At this size re-reading everything costs four seconds more than
reading none. Stated limit: they do not extrapolate.

Records four things that went wrong and were caught rather than shipped:
  - R-398 was MY OWN mistaken row. The seam already existed, Part 0 was not
    built, and building it would have HIDDEN the unlock --remove-all escalation
    from the assertions that must see it. Corrected, not closed.
  - the damage classifier matched restic's ORDINARY progress output; the
    negative control caught it (v0.227.1).
  - an exit code I misread through a pipe, corrected by re-measuring.
  - wire-contract convicted a field the hub cannot decode; allowlisted WITH A
    REASON rather than skipped, because building the hub display is a decision
    R-331 already ruled belongs to the operator.

And the sweep the task asked for: EIGHT debug buttons post to endpoints that do
not exist, not one. 24 referenced, 17 dispatched. Filed as R-400.

Not live-validated and each says why: the weekly firing (a week away), read-data
on a large store, and the join between "restic catches it" and "my code
classifies it" -- both proven, the join is not, and the seam is named.
2026-08-30 21:29:15 +02:00
admin e64c84aef8 REPORT: the golden debt named in section 12 was paid the same day
gates / gates (push) Successful in 12s
Golden 0.226.1 baked, published, round-trip verified, vouched, floor raised.
golden_currency_gate.py red -> green on the same command, so the three declared
--no-verify bypasses are historical rather than standing.

The fleet state changed with it: both demo machines run 0.226.1, and
demo-felhom got there by SELF-UPDATE rather than by hand -- which is the
positive observable that the floor is acting and not merely set.
2026-08-30 20:24:14 +02:00
admin c476ae51a5 REPORT: record the 0.226.1 patch and which version the live evidence describes
gates / gates (push) Successful in 12s
The section 6 live validation ran against 0.226.0, which was already on the box
when the fallback defect was found. Saying so rather than re-attributing the
evidence to 0.226.1: they differ only by CountsUnknown and its two tests, and
neither touches any path that validation exercised.
2026-08-30 19:53:21 +02:00
admin c0c8fe67bf An unknown drawn as a zero: the defect v0.226.0's own fix introduced
gates / gates (push) Failing after 13s
Writing the REPORT's observation "the no-unit fallback already reports a zero
result, which is honest" exposed that the sentence was FALSE.

A zero UnitRestoreResult is Scenario B's shape. So RestoreFromRecoveryUnit's
fallback to RestoreApp -- which returns only an error, and whose signature is
deliberately out of scope -- would have printed "ez a mentes csak a
beallitasokat tartalmazta, adatot nem" over a restore that may have replayed the
app's entire dataset. That is an unknown drawn as a zero: the exact R-88 failure
direction this whole change exists to remove, re-introduced by the change.

UnitRestoreResult now carries CountsUnknown, the fallback sets it, and there is a
fourth sentence claiming only what is known -- the restore ran, the app is back,
and we cannot say what came back. RestoreApp's signature is untouched.

Pinned by TestUnitRestoreOutcome_NoUnitFallbackSaysUnknownNotEmpty. The A5 seam
test was corrected too: its fixture has no recovery unit, so it exercises exactly
this path and had been asserting the wrong sentence -- it now asserts the
unknown, which is what pins the fallback to it.

IT WAS THE observations GATE REFUSING THE PUSH THAT FORCED THE RE-READ. A gate
written to stop findings dying in an overwritten REPORT.md caught a live defect
instead. Also files R-397 (NotifyIntegrityOK/Failed are dead code AND the
monitoring page advertises a weekly integrity check that does not exist) and
R-398 (resticStep is not a seam, which is why R-358's ordering needed an AST
test) rather than leaving them in a file that is overwritten every session.

REPORT.md is the full run record: baselines re-confirmed, per-test results, the
five red-proofs with their observed output, the live validation with verbatim
Hungarian messages, what was NOT validated and why, teardown across three
layers, and the register 165 -> 167 -> 161.

Green gate clean: 28 packages, rc 0. All 12 controller gates OK.
2026-08-30 19:51:16 +02:00
admin e4e0aa8f46 REPORT: v0.226.0 shipped, three of four fixes proven live on demo-hp
Records what was validated and, in equal detail, what was not.

PROVEN LIVE (demo-hp, endpoints the UI invokes, evidence copied off the box):
  R-353  "A(z) opengist: 1 adatkotet visszaallitva -- az alkalmazas ujraindult."
         read off the customer's own wizard page, with real counts 1/1 volumes
         and 0/0 databases and correctly no database clause.
  R-360  refused in the exact flag state that produced the bug, and the planted
         canary file survived -- the consequence, not the branch.
  R-358  a mode=unit restore wrote {"schema":1,...,"full":false} at mode 0600
         with no .tmp left, and the gate logged place-to-live closed.

NOT live-validated, and each says why rather than being omitted:
  R-357  filling a real filesystem is a drill step, not a build step.
  R-353 Scenario B  NO app on demo-hp still has a data-less unit -- the spec
         named opengist from 21 August and it has since been recaptured (now
         1 volume dump). Manufacturing one means falsifying a manifest, which is
         the hand-set-state shortcut this project forbids.
  R-353 Scenario C and R-358's failed-download branch: unit-tested only.

Also recorded, because a near-miss that is quietly fixed teaches nobody: the
first B1 red-proof exposed a HOLLOW TEST OF MY OWN. With the gate removed the
run refused earlier, at the placement stat pre-pass, so `stops == 0` passed
against the pre-fix code. Fixture corrected and assertions reordered; only then
does the red-proof print THE APP WAS STOPPED (1 call(s)).

Register 165 -> 167 -> 161. Six rows compressed into CLOSED-ITEMS keeping title,
version, evidence and every sentence stating a rule; full original at
`git show e027b5d9`. No open row touched. ROADMAP not edited -- none of these
four ever had a row there, stated rather than silently skipped.

Teardown: this run provisioned nothing, across all three layers. Two throwaway
scripts and one canary directory were planted in guest 9201 and both removed.
2026-08-30 19:41:48 +02:00
admin e5eee501b5 R-331 (controller half): forward stats_known so the hub can tell empty from unmeasured (v0.225.0)
gates / gates (push) Successful in 12s
The hub's operator Backup card read `Snapshots 0 / Repo Size 0 MB / Integrity
Unknown` for EVERY customer, because it rendered the report's `backup` object --
whose snapshot/size/integrity fields have had NO producer since disk-tier restic
moved to the host agent (slice 8C). buildBackupReport leaves them zero
deliberately and says so. Measured on demo-hp 2026-08-30 while that night's log
said `[offbox] backup OK: 8 app(s) backed up, 67 snapshot(s), 2m14s`.

The live numbers were always in the report's `offsite` object, which the hub
already reads for its Offsite page and its fill/staleness alarms. The hub fix is
to render that -- and that made exactly ONE field mandatory that was not being
forwarded.

snapshot_count:0 means two opposite things: "holds nothing" and "never
measured". R-225 measured that confusion inside this repo (a rebuilt box
rendered 0 pillanatkep over a store really holding snapshot f3d9cd67), and
settings.OffboxTarget.StatsKnown fixed it for the controller's own UI. It was
never put on the wire, so the hub was free to make the identical mistake one
layer up -- and did. OffboxReportStatus.StatsKnown now carries it, omitempty, so
an older controller sends no key and a reader degrades to UNKNOWN, never to
EMPTY. Absence is ignorance, not emptiness.

The four dead BackupReport fields stay on the wire (historical reports in the
hub store must keep parsing) but now carry a warning naming R-331 and pointing
at Offsite. TestBackupReport_DeadFieldsStayZero fails the moment a producer
appears for one -- the prompt to update the hub card in the SAME change rather
than ship a field nothing renders.

RED-PROOF: drop `StatsKnown: t.StatsKnown` -> "a MEASURED empty repository
reported stats_known=<nil>". Tests assert the JSON the hub sees, not the Go
struct: measured-empty and never-measured must differ ON THE WIRE, which is the
entire point of the field.

Green gate clean: 28 packages, rc 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-08-30 18:38:10 +02:00
admin 45b52b6ed5 R-330: live validation on demo-hp — 3 scans inside the window, 0 events
gates / gates (push) Successful in 12s
v0.224.0 deployed to both demo boxes (both `0.224.0 … (healthy)`). Proven
through POST /api/backup/run, the endpoint the UI button invokes: 8 stacks
stopped and restarted over 87s, three dead-app scans ran INSIDE that window
(16:09:26 docmost, 16:09:56 paperless-ngx, 16:10:26 romm -- the same three apps
that alarmed the night before on 0.223.0), zero app_start_failed pushed.

The scan count is the positive control, not decoration: an absent alarm is
equally consistent with "suppressed correctly" and "the scanner stopped".

A first run is discarded IN THE REPORT rather than quietly dropped -- it fired
52s after a controller restart, inside deadAppBootGrace (90s), where the scan
returns early and could not have alarmed whatever the code did. demo-felhom is
deployed but NOT independently proven and says so: its single app cycles in ~1s,
too fast for any 30s scan to land inside.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-08-30 18:12:26 +02:00
admin 92cebb8c95 R-330: stop the backup alarming about the apps it is holding down (v0.224.0)
gates / gates (push) Successful in 11s
Measured live on demo-hp 2026-08-30 (controller 0.223.0): the nightly db-dump
and offbox-backup legs stop each stack ~13s to tar its volumes while the
deadapp-check job scans every 30s, so the scan caught whichever stack was
mid-cycle and pushed app_start_failed to the customer. 61 e-mails about apps
that were never broken.

The defect is not a missing mechanism. quiesce/suppress.go solved exactly this
in v0.179.0 and works -- but classifyRunStates read only the quiesce loop's set,
and that loop covers the WHOLE-GUEST backup. The per-app legs stop stacks
through Manager.DumpAppVolumesSafe, which registered with nothing. Two
mechanisms stop apps on purpose; only one told the alarm. Fifth instance of the
"seam built but never wired" class, and the first where the unwired half was a
consumer.

The suppression now rides AppStopGuard, which already brackets every deliberate
stop in the product (Begin before the stop, End after a successful restart) at
all three call sites, and which main.go hands as ONE object to the backup
manager and the exporter. scanDeployedAppRunStates takes the union of both sets.
All three per-app stop paths are covered, not only the reported nightly one.

It cannot latch -- End() runs only on a restart that SUCCEEDED, so unlike the
quiesce loop an open-ended hold is a real hazard here:
  1. ReleaseFailed drops the entry IMMEDIATELY on a restart that broke, wired at
     every failure path, so the app alarms on the next scan;
  2. Begin REPLACES the set (one marker file = one operation);
  3. appStopMaxHold (6h) caps a hold nothing released, logged at WARN.
Grace is 180s, deliberately quiesce's own constant and derivation. Suppression
is NOT persisted: after a crash the guard holds nothing and a down app must
alarm. ReleaseFailed keeps the durable crash marker; a test pins that.

Three companion red-proofs, each printing the pre-fix value (REPORT.md section 5):
  - drop markStopped from Begin      -> "suppressed at stop = map[]"
  - drop ReleaseFailed from the dump -> "map[bookstack:true] after a restart that FAILED"
  - pass nil instead of appStopGuard -> the AST wiring test fails
The third is load-bearing: the component was never the broken part, so a suite
that only injected it would have been green against the shipped defect.

Green gate clean: go build + go vet + go test ./... -- 28 packages, rc 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-08-30 17:58:15 +02:00
admin f8c9390946 gate 11: register the observations gate, and mark up this repo's observations
gates / gates (push) Successful in 13s
The shared gate lives in felhom.eu/scripts/observations_gate.py and is invoked
across the workspace, exactly as reuse_refs_check.py and instructions_gate.py
already are. It is never copied.

REPORT.md's observations now carry their markers. Item 1 was the finding that
had no register row - only the first broken app per hour reaches the operator -
and it is now R-389. Item 4, the golden-bake runbook's missing `pveam update`,
is R-390. Items 3 and 5 are declared NOT-A-FINDING with their reasons. The
observations' text itself is unchanged; only the markers were added.
2026-08-23 13:53:18 +02:00
admin 1da2c9c6c6 docs(v0.223.0): REPORT, CONTEXT rulings, README severity contract
gates / gates (push) Successful in 11s
REPORT overwritten: the 1.1 sweep in full (one bad severity, nine legitimate
"warn" strings that are healthcheck statuses), the hub manifest's real location
since the task's premise was wrong, all five red-proofs with the layer each
guard sits at, the live walk in six steps with the hub's own records quoted, and
the absent-intent count (0 of 8).

Three things are reported that a tidier account would omit: red-proof 5 passed
first time because the mutation was INERT; Scenario G was silently refused twice
behind an HTTP 200; and the live Scenario A does NOT prove the customer gate,
because demo-hp has no prefs row at all.

CONTEXT records the severity vocabulary as a ruling with its mechanism, the
intent ruling with its three-way handling of unknown, both fences, and two traps
worth more than the fixes: a 200 can be a refusal, and a passing red-proof can
mean an inert mutation.

README: the event table said `app_start_failed | warn` - the defect, written
down as if correct. Now `warning`, with the vocabulary contract and who receives
what. `disk_critical` also corrected from `error` to `critical`, which is what
fillwatch has always sent.
2026-08-23 12:06:45 +02:00
admin 14137efac5 docs(v0.222.0): REPORT, CONTEXT decisions, README state table + the ordering
gates / gates (push) Successful in 11s
REPORT.md overwritten with the full run: baselines and the hub's four numbers,
the four red-proofs with the mutation and observed text for each, the five
IsDownState consumers walked and named, the live walk in full with the old and
new heartbeat lines quoted side by side, and the halt.

CONTEXT records the decision - a dead supervised member is asked about before a
failing healthcheck, because they are different questions and the second was
answering the first - plus the fence that IsDownState did not move, the trap
that three existing subtests pinned the defect, and R-386.

README gains the `degraded` row, which the state table never had, and a note
that the ORDER is load-bearing. Points at the new alarm-ladder architecture doc.
2026-08-23 07:58:09 +02:00
admin f7881787f4 R-361 docs: CONTEXT decision, README, REPORT
gates / gates (push) Successful in 11s
Records the db_dumps decision with every consumer named, the trap that a stable
db_dumps lets CaptureRecoveryUnit's already-current early return fire (so
per-capture housekeeping must sit above it), and the NEGATIVE that a held app
does not raise the dead-app alarm - measured, not reasoned, so nobody re-derives
it.
2026-08-23 00:26:32 +02:00
admin cbc3fa589c docs: R-356 CONTEXT + REPORT (v0.219.0, proven live on demo-hp)
gates / gates (push) Successful in 11s
2026-08-22 13:25:36 +02:00
admin f94543ee5c v0.217.0: prefill from the app's own backup, where-the-data-goes on deploy, bounded inventory fan-out
gates / gates (push) Successful in 10s
Completes R-351 and ships R-352's visibility half. Gates 11/11 OK, suite 28 packages ok,
go vet clean, -race clean on the changed package - all run and read BEFORE this commit.

PART 2 SCENARIO A - the deploy page prefills the address and data folder from the app's OWN
backup. backup.RecordedUnitForStack scans every readable namespace root (the app is NOT
installed in this case, so there is no own drive to ask) and reads manifest.json plus the
captured compose/app.yaml. Local file reads only: no network, no restic, no restore.
RecordedAddress.Known() requires BOTH halves on purpose - an absent SUBDOMAIN makes the live
deploy path substitute the CATALOG default (stacks/deploy.go:88-90), and offering that back as
"what your backup says" would be a fabricated fact. The prefill is labelled as coming from the
backup and stays editable: a memory, not a lock.

PART 1 VISIBILITY (R-352) - the deploy page now states where the app's data will live before
the button is pressed. Measured 2026-08-21: 13 of 53 catalogue templates declare a storage
field; the other 40 have none and their data goes to the system drive, which no screen said.
Metadata.HasDeployField answers "does this app have somewhere to PUT a recorded value?" - for
the 40-class a recorded placement is a fact to state, never a value written into a field that
does not exist. NO PLACEMENT CHANGED. NOTHING MIGRATED. The rest is a filed specification.

PART 4 - measured before theorising, on the live off-site target:
  snapshots --json 2605 ms once; stats 2697 ms PER APP, sequential, 5 app tags
  => 2605 + 5*2697 = ~16.1 s, matching the reported ten-to-fifteen seconds.
The cause is the shape already on file, so the per-app size calls now run concurrently,
BOUNDED TO 4. The bound is the safety property, not the speed one: the repository is a Hetzner
Storage Box with a session cap, and a refused size call returns SizeBytes 0 - a silent
UNDER-REPORT of the customer's data rather than a visible failure. Peak-in-flight is asserted.
OffsiteInventoryList had no test at all before this.

TEMPLATE SAFETY - every Restore* key is set UNCONDITIONALLY in the deploy handler, because a
template doing index/eq against an undefined key errors at RENDER time: green build, green vet,
green suite, 500 on the page. Four render tests, one per branch, because the existing deploy
render test only renders AutoFields and never reaches these blocks.

RED-PROOFS, mutation asserted applied then reverted to 0:
  A   three template guards dropped (count asserted 3) -> the blank form returned
  P4  inventorySizeConcurrency = 1 -> "peak in flight was 1", elapsed 282ms = sequential

DOCS: CHANGELOG v0.217.0 (MinAgent 0.129.0 unchanged), CONTEXT (the restore's own memory +
what is next), controller/README.md (Backup System), REUSE.md (4 new rows), REPORT.md
overwritten - the previous REPORT preserved to audits/REPORT-v0.216.0-2026-08-14.md first.

NOT fixed here, filed as R-353 and named the next session's first item: a restore whose unit
carries no db_dumps and no volume_dumps still reports a bare completion.
2026-08-21 21:29:01 +02:00
admin 2fa1efc5e5 docs(REPORT): confirming cycle on v0.216.0, and persistence proven live
gates / gates (push) Successful in 9s
- 09:31:35Z on 0.216.0: '2 disk(s) evaluated, 0 alert(s)' — the count now
  matches the 2 persisted records, closing the disagreement that exposed R-335.
- The 0.215.0 -> 0.216.0 redeploy replaced the container and the state file
  came back with a changed_at written by the PREVIOUS version, so the new
  container loaded the pre-restart record instead of re-baselining. Scenario L
  observed on real hardware, not just through the production-path unit test.
- R-332 narrowed accordingly: what remains unproven is an already-ALERTED disk
  not re-alerting after a restart.
2026-08-14 11:33:11 +02:00
admin 330e4a051e docs(REPORT): v0.215.0 -> v0.216.0 run report
gates / gates (push) Successful in 9s
Includes the two clean live cycles, the warning-vs-warn notification_log proof,
the 13 red-proof outcomes (A reported as a finding — the spec's mutation for it
is not a valid red-proof), and section 14 on R-335, the aliasing defect found
live in v0.215.0 and fixed in v0.216.0.
2026-08-14 10:33:22 +02:00
admin 3e3ee94b7b REPORT: live 422 proven on hardware; R-308 withdrawn (my quoting bug, not a stale credential)
gates / gates (push) Successful in 12s
2026-08-12 19:05:23 +02:00
admin ae10f64806 REPORT: controller v0.214.0 — the screen stops hedging, and the claim guard grew a surface
gates / gates (push) Successful in 18s
2026-08-12 18:49:56 +02:00
admin 3168a78935 REPORT: v0.213.0 pinned-fingerprint condition, red-proofs, claim guard
gates / gates (push) Successful in 12s
2026-08-12 15:39:08 +02:00
admin 1b66010298 REPORT: v0.212.0 orphan card second promise
gates / gates (push) Successful in 16s
2026-08-12 14:05:45 +02:00
admin f87be3575f REPORT: correct installer publication status
gates / gates (push) Successful in 12s
2026-08-10 14:20:22 +02:00
admin 38f4535bfa REPORT: golden 0.211.0 baked and published; only the Day-0 vouch remains
gates / gates (push) Successful in 14s
2026-08-10 14:20:08 +02:00
admin 397d62136f REPORT: v0.211.0 written, not delivered - bake and Day-0 approval outstanding
gates / gates (push) Successful in 13s
2026-08-10 13:58:45 +02:00
admin 37b5ba08a7 REPORT: v0.208.0 — both R-254 sites, the guard's measured holes, and what the live read could not prove
gates / gates (push) Successful in 13s
Records the deliverables, and is explicit about the limit on the live half: the
curl of an app info page could not be done, and names exactly what was tried —
crafty-controller is the only app declaring initial_credentials and is deployed
nowhere, and demo-hp's dashboard password in ~/.config/credentials no longer
authenticates (200 with no session cookie). A probe of the new routes was
discarded because its control killed it: real and bogus paths both 302 behind the
auth middleware.

§7.2's answer including the part that contradicts the task's premise: no line in
the repo says 'no silent auto-fill'; the rule is CONTEXT.md:2070 about accidental
EMPTY-password deployments. The hidden input is deliberate and untouched.

§7.4's measurement: the gate covers all 36 templates and catches a launder through
a local variable, but is blind to a secret under a neutral page-data key — the
exact shape of site two. Runtime coverage is 4 of 27 pages. Filed as R-255 rather
than described as complete.

§7.3: no evidence of actual exposure on the fleet, with the limit stated — it is a
current-state measurement and nothing recorded reads, which was part of the fault.

Also corrects v0.207.0's report: html/template STRIPS HTML comments; they do not
ship in the response body. Measured.
2026-08-07 21:29:13 +02:00
admin 62998aab4f REPORT: v0.207.0 — R-249 before/after on a live box, the census, and what was NOT proven live
gates / gates (push) Successful in 18s
Records the deliverables: the raw response body before (1 occurrence, v0.206.0)
and after (0, v0.207.0) with a positive control in both directions; the §7.1
census finding two more instances of the render-then-hide pattern (R-254, one a
real per-install secret); §7.2's decision and why the promise was the wrong half;
every changed Hungarian string; all eight tests with their red-proof outcomes.

States plainly what was NOT proven live: R-252/R-253's notices could not be
rendered on VM 325 because both states are rebuild-only and the box re-registers
a drive on restart — the live run therefore exercised Scenario E instead, and the
notices are pinned at the template + predicate level with red-proofs.

Also records that red-proof D caught a fault in my own work: the explanatory HTML
comment quoted the old sentence, and HTML comments ship in the response body, so
the contradiction was still on the page and the assertion forbidding it could
never fail.
2026-08-07 18:33:37 +02:00
admin 3d3b4496f3 REPORT.md — R-241 fixed, deployed, live-validated on both demo boxes
gates / gates (push) Successful in 22s
Scenario A's live result first: on demo-hp in the rebuilt shape, no key was
minted on the real start-up offsite-apply path, and the hub received the state
it reports instead - offsite.state=awaiting_recovery_key with enabled:false.
Key restored byte-identical afterwards.

Includes Q4's seven rows mapped to the three states, the SEC 7.2 choice and
why, SEC 7.3's answer on the new-code button, every changed Hungarian string
quoted, all nine red-proofs with what was mutated, the R-245 reasoning, and
three observations noticed but not acted on.
2026-08-07 12:23:55 +02:00
admin c6b69d888e v0.205.0 — a run that skipped an app the customer selected is not successful (R-234)
gates / gates (push) Successful in 21s
THE VERDICT. The R-203 block already said "a warning beside a success is read as a
success" and applied it to ONE of the two shapes it describes: an app missing a
declared mandatory FOLDER made the run incomplete, while an app skipped ENTIRELY
still reported ok. Both do now. Which skips count, decided by measurement:
selected+deployed with no recovery unit YES; selected but NOT deployed no (named,
with what to do — a box left amber by an app somebody removed is a status nobody
reads); disconnected/decommissioned drive no (own signal); nothing selected no.
LastSuccess and SnapshotCount still record what WAS captured.

THE FILED MECHANISM WAS NOT THE MEASURED CAUSE, and saying so is the point. §3
stated that toggling an app on leaves it without a bundle so the first run skips
it. Measured on demo-hp: the run's own pre-dump phase calls captureAllRecoveryUnits
for every DEPLOYED stack, through admitApp, before the push — a unit moved aside
was RECREATED and the run reported ok. That state does not survive a run.

What actually produced the 2026-08-06 sequence: the manual run was dropped by the
single-flight while an earlier run was still going. runOffboxBackup returned nil,
the handler had already answered "A tavoli mentes elindult", and the card then
showed the PREVIOUS run's green verdict — read as covering the app just selected.
The decision is now taken synchronously in the handler and a dropped request says
so. The nightly path still returns nil on purpose: nobody asked, and it retries.

§7.3 measured before deciding: CaptureRecoveryUnit writes a few KB of compose +
manifest, only ENUMERATES dumps rather than creating them, is idempotent and does
NOT stop the app — and already runs inside the off-site run. So there is no wait to
remove for a deployed app and NOTHING was built.

28 packages ok, 9/9 gates. Four red-proofs, each asserted to have applied. Fixture
note: the shared provider's ListDeployedStacks returned nil, so Scenario A first
passed for the wrong reason; fixed with an opt-in deployed set that defaults to nil.
2026-08-06 21:58:21 +02:00
admin 53e9bf0224 v0.204.0 — the restore list is keyed on the store (R-237); the size gate stops refusing in silence (R-238)
gates / gates (push) Successful in 26s
R-237: /backups/restore listed apps that are CURRENTLY DEPLOYED and CURRENTLY
TOGGLED ON for future off-site backups. A rebuilt box has neither, so a household
that had just lost everything was shown nothing to restore while the repository
held their snapshots — measured live on the R-201 re-walk. To restore an app you
had to select it, to select it you had to have installed it, and to know what to
install you had to see the backup you could not see.

The store is now the source of the list (offsite_restore_list.go), built on the
existing R-193 OffsiteInventoryList. Installed-ness became a property OF a row,
never a filter on it. Every case is answered rather than hidden: a snapshot for an
app that is not installed is offered and says it will reinstall first; an installed
app with no snapshot is shown as having nothing; an unreadable store renders as
UNKNOWN (R-225's rule, one screen over) AND keeps the action, because "we could
not look" is not "there is nothing"; no-target is its own state. The felhom-offbox
and _shares marker tags are excluded from the app list.

R-238 classified as a HARNESS ARTIFACT: mode=full without confirm=1 is step 1 of a
deliberate two-step — it starts no job by design and redirects carrying
&full_prep=<app>, which deriveWizardStep requires to reveal the commit. A driver
that did not carry it forward landed back on the intent step. The operator's
browser run completed the same restore. The wizard's precedence rules were NOT
re-keyed: a stale ?full_prep= must never resurrect a commit button mid-restore.

The residue WAS real and is fixed: neither branch of that step wrote anything to
the log, so a refusal — including by the headroom gate — left no trace on the box.
Both branches now log, and so does the concurrent-op refusal.

resolveWizardApp is removed: it was dead once the gate moved, and its test pinned
the defect's behaviour (an untoggled app refused), which would have read as policy.

28 packages ok, 9/9 gates OK. Three red-proofs, each asserted to have applied.
2026-08-06 16:44:05 +02:00
admin 4d349d1106 REPORT + CONTEXT for v0.203.0: the retry shape, the marker answer, R-220's shape
gates / gates (push) Successful in 10s
Records the decisions rather than only the code:
- POLL not ACK, decided on Scenario B against the ACTUAL promises — the
  no-target message gives no deadline and the card says 'within a day', so a
  5-minute tick is inside both and no text needed changing. If either promise
  tightens to minutes, go ACK-driven.
- The marker question: applied_marker lives in the guest's DataDir, which a
  rebuild destroys, so it cannot suppress a legitimate re-run. Left alone.
- R-220 candidate (b), corroborated rather than a wider prefix, reading
  /proc/mounts because the lsblk args are pinned in sudoers.

Live: Scenario C proven on demo-hp WITH a positive control — the job ran once
and logged nothing. A first reading counted 2 lines that turned out to be the
start-up reconcile, not the retry; the instrument was corrected before the
conclusion. Scenarios A and E are deliberately NOT live-proven here: both need a
rebuilt box, and that state arises naturally in Part 4.
2026-08-06 13:05:30 +02:00
admin a62bb3874b REPORT + CONTEXT: v0.202.0's rule, its live proof, and what was NOT verified live
gates / gates (push) Successful in 9s
CONTEXT gains the rule so it outlives the bug: on the unlock path the customer
is blamed only after a real attempt REFUSED their code; every other outcome,
including an unclassifiable one, says something else. Plus the two things that
must not be 'fixed' into it — elapsed time is never a classifier, and the error
TEXT is never read (when the distinction was not a value, the agent was changed
to provide one).

REPORT states the split honestly: the AGENT half is proven live on the venue
(400 -> 502 -> 400, same wrong code, only the hub's reachability changed), while
the controller's message selection rests on handler tests and red-proofs,
because /recovery correctly redirects since F7 set the old data aside and
restoring that state is the reconfiguration §11 forbids. Also records that the
correct codes were shredded by the previous session, so the live re-run used a
WRONG code — which makes the test harder, not weaker.

Two venue changes stated because they were not asked for, both restorations: a
fresh dashboard password (the previous session shredded it, leaving the box
impossible to log into) set through the supported --print-reset-code escape
hatch, and one normal off-site run to populate stats_known.
2026-08-06 08:34:19 +02:00
admin 05cf352a2f docs: CAMPAIGN-11 fix pass — REPORT + CONTEXT (controller v0.201.0)
gates / gates (push) Successful in 8s
2026-08-05 18:03:50 +02:00
admin a315d623b8 docs: R-193 CLOSED — CONTEXT + REPORT (controller v0.200.0)
gates / gates (push) Successful in 10s
2026-08-05 12:56:43 +02:00
admin be3c5fa7f6 docs: R-204 item 4 (box half) — CONTEXT + REPORT (controller v0.199.0)
gates / gates (push) Successful in 10s
2026-08-05 11:06:19 +02:00
admin 68f195676b docs: R-204 items 1 & 3 — CONTEXT, REPORT, README (controller v0.198.0)
gates / gates (push) Successful in 9s
2026-08-05 07:37:25 +02:00
admin f4796e0d00 docs: R-203 contract + report (controller v0.197.0, proven live)
gates / gates (push) Successful in 10s
2026-08-04 18:53:47 +02:00
admin 532f5712a8 docs: R-200 Part 0 shipped; R-203 recorded (mandatory userdata dir missing from the offsite snapshot while the run says ok)
gates / gates (push) Successful in 9s
2026-08-04 15:00:40 +02:00
admin bdab80c933 docs: R-200 diagnostic — CONTEXT + REPORT (proven live on demo-felhom)
gates / gates (push) Successful in 10s
2026-08-04 13:56:01 +02:00
admin 0887fd676d REPORT: R-182 — the run digest, the live proof, and the red-proof that did not fail first time
gates / gates (push) Successful in 9s
2026-08-03 13:59:42 +02:00
admin db0d4b129d REPORT: R-181 — the reserve, the live proof, the du measurement and the teardown
gates / gates (push) Successful in 9s
2026-08-03 11:37:33 +02:00
admin d5be67b913 REPORT: CI run ids and conclusions (all three commits green)
gates / gates (push) Successful in 9s
2026-08-02 23:57:07 +02:00
admin 95eb5c2c1a REPORT: record every CI run id, run number and conclusion
gates / gates (push) Successful in 9s
2026-08-02 20:38:52 +02:00
admin e6311f9fbc docs: CONTEXT + REPORT for v0.190.0 (R-157 A / R-170 / R-171)
gates / gates (push) Successful in 8s
2026-08-02 20:34:48 +02:00
admin 3446609420 REPORT: record all three CI run IDs and their conclusions
gates / gates (push) Successful in 9s
2026-08-02 18:58:56 +02:00
admin a8f7c61d41 docs: CONTEXT + REPORT for v0.189.0 (R-166)
gates / gates (push) Successful in 9s
2026-08-02 18:58:11 +02:00
admin e7c44c0e0f docs: CHANGELOG + REPORT for the CI workflow (no version bump)
gates / gates (push) Successful in 9s
2026-08-02 16:35:42 +02:00
admin eaded79b18 REPORT: gate enforcement session (no version bump) 2026-08-02 15:37:02 +02:00
admin 4115e88f68 REPORT: D5 — restore from the drive alone (v0.188.0), proven live 2026-07-30 17:00:32 +02:00
admin 2f27a363d5 R-108: network storage may not host an app's data namespace (v0.187.0)
This is D5's precondition and it is now met.

An app's namespace root IS its backup root: namespaceRoot returns a non-system
drive path as-is, so the recovery unit lands at <HDD_PATH>/backups/primary/<stack>/.
On a NAS that sits inside the share, which FileBrowser binds WHOLE — share root,
:rslave, download:true.

The bind was NOT narrowed, and establishing why inverted the fix. The share-root
:rslave bind is load-bearing (a 2026-07-22 probe proved an in-container access
through it wakes the idle automount trigger), and scoping is undefinable anyway:
apps on a share store at <share>/<app>, there is no userdata/ layer, and creating
one would write Felhom convention onto a customer's own NAS, which R-67 forbids.
So the browsing surface cannot be narrowed and the backup tree must never be
placed under it. Operator ruling: refuse the placement, keep the browse bind.
Tier 2 already refuses network targets for this reason (F-6C-1).

Nothing stranded: zero apps on network storage across all six hub customers
including Peti. R-67's browse capability is byte-identical.

FIVE surfaces, not the four the register named — settings.RefuseAsAppNamespace is
the single predicate. The deploy POST is the real boundary (it accepts any
caller-supplied HDD_PATH; DeployStack validates only os.Stat). Surface 4,
handleStorageDecommission mode=migrate, guarded only its SOURCE, so a whole
namespace could be decommissioned ONTO a NAS — that one is not in the register.

Fails closed: /mnt/felhom-drives holds both kinds, Kind exists only on a
registered path, so an unregistered path under that root refuses.

Supersedes README's "NAS backup locality — decision A" (v0.118.0).

9 tests, all non-effect (nil stackMgr, so a guard that misses panics rather than
passing). 4 red-proofs, each mutation asserted to have landed.
Suite rc=0, 27 packages, 0 FAIL. vet rc=0. Template + emoji gates OK.
2026-07-30 14:10:20 +02:00
admin b331f18424 v0.186.0 — R-114 + R-112: tell the truth about the backup target, then show it
Two defects E-2d found on a real box, fixed in this order deliberately: the
message is corrected BEFORE it is put on screen, because switching on a banner
that lies is worse than a silent one.

R-114 — the third state. resolveBackupTargetState had two outcomes: a disk
claims the target (healthy), or nothing does (degraded, "the backup is on the
system disk"). The state "configured, and its drive is gone" had no branch, so
it fell into the second and inherited its message AND its offer. Observed live
with the target detached: degraded:true, target:"felhom-backup" plus the
system-disk copy (false -- the backup was on a drive that had vanished) plus
offer_path naming that same vanished drive as the remedy.

New BackupTargetState.TargetAbsent discriminates. Degraded keeps its meaning
("is there a problem") so the wire contract is unchanged for every consumer;
TargetAbsent answers "which problem", because the two have opposite remedies --
attach any second drive, versus reconnect THAT one. Copy routed through
degradedMessageFor so one place still decides what a customer reads. The offer
is suppressed on the branch itself, NOT left to firstOfferableDrive's
Disconnected skip: that flag is set by the agent-side gate in another repo
(R-113), and this state must be correct independently of it.

R-112 — the state finally has a consumer. The endpoint was byte-correct and
nothing in the product ever asked for it: templates fetch 18 distinct
/api/storage/* endpoints and backup-target[/assign] were the only two with zero
references. Server-rendered on /backups now, following the existing
SingleCopyWarning banner pattern -- not a 19th JS fetch, because a banner that
needs JavaScript to appear is one more thing that can silently not happen.
backupTargetView returns nil for healthy and unknown so those render nothing at
all. The offer control POSTs to the existing assign endpoint behind the standard
inline confirm, never auto-submits, and surfaces restart_required honestly
instead of adding a self-restart.

Scenario E (the seam test) drives backupsHandler over httptest and asserts the
RENDERED HTML -- handler -> view -> resolver -> template. It deliberately does
not call the resolver and assert a string, which would prove the resolver that
was never broken. Deleting the one line that sets data["BackupTarget"]
reproduces the R-112 state and fails every render assertion.

Tests 326 -> 338 (+12) in internal/web; suite green (27 packages); both template
gates pass. Three red-proofs run and reverted, files byte-identical after.

MinAgent unchanged at 0.113.0: R-114 reads BackupTarget/MountPath/GuestPath/Role,
none of which R-113 altered (it changed BoundUnderParent, which this code does
not read). demo-hp on agent 0.113.0 is not held.

The absent copy is verbatim the hub's customerMessages["backup_target_absent"]
so the banner and the email tell one story -- filed as a two-repo drift risk,
not solved.

NOT LIVE-VALIDATED. Scenario C cannot occur on a healthy box; Session C proves it.
2026-07-29 19:21:32 +02:00
admin d8b3279731 REPORT + CONTEXT: R-101 + F-DIAG (v0.182.0), rendered dialog proven live 2026-07-28 16:46:16 +02:00