Commit Graph

1256 Commits

Author SHA1 Message Date
admin 6aaa3a4a34 BIGNIGHT phase 4: R-517 (P1, backup page claims a failed PBS tier is current and present), R-518 (apps down ~8 min on 'a few seconds'); backup evidence
gates / gates (push) Successful in 20s
2026-09-14 21:16:16 +02:00
admin e0366f1f05 BIGNIGHT phase 3 done: 12 apps seeded and used, 0 interventions; R-515, R-516 filed; box logs
gates / gates (push) Successful in 20s
2026-09-14 20:58:22 +02:00
admin af256795b5 BIGNIGHT phase 3: R-514 filed (paperless OOM on a 20-document upload, silent); immich, jellyfin, mealie evidence
gates / gates (push) Successful in 19s
2026-09-14 20:47:01 +02:00
admin 5e8bff6808 BIGNIGHT: R-513 filed (P1 security: FileBrowser admin/admin on every box; demo-hp login page public)
gates / gates (push) Successful in 21s
2026-09-14 20:32:03 +02:00
admin 8a12c9a1bc BIGNIGHT phase 3: immich + vaultwarden seeded; R-512 filed (vaultwarden open signup, read-only control)
gates / gates (push) Successful in 18s
2026-09-14 20:27:52 +02:00
admin 3a7bbd2f6b BIGNIGHT phase 3: bookstack, docmost, privatebin, gokapi, nextcloud seeded and used; evidence
gates / gates (push) Successful in 20s
2026-09-14 20:25:27 +02:00
admin 4d92127f1a BIGNIGHT phase 2: claim with the mailed code, gate FAIL (R-510), R-511 filed (DR tier stuck), drive enrolled, box logs
gates / gates (push) Successful in 19s
2026-09-14 20:13:51 +02:00
admin 9993f7813e BIGNIGHT: R-510 row itself (previous commit's insert failed on a quote; journal was already pushed)
gates / gates (push) Successful in 20s
2026-09-14 20:09:49 +02:00
admin 39a2627f40 BIGNIGHT: R-510 filed (tester-1 route lacks noTLSVerify -> 502) before intervention I2; bind + reenroll-mail evidence
gates / gates (push) Successful in 19s
2026-09-14 20:09:29 +02:00
admin d9522c38fb BIGNIGHT: R-509 filed (no self-bind mail for an existing customer) before intervention I1; journal + screens so far
gates / gates (push) Successful in 19s
2026-09-14 19:55:33 +02:00
admin a4d684412b R-505: tunnel had no published route; operator added *.enkicsifelhom.hu -> https://traefik, verified with a throwaway connector; day-0 A.1 names the exact route
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-14 19:18:31 +02:00
admin 6e4d372720 doorstep teardown complete: host tester-1-8603a2 deleted, ep0 peer gone, customer kept; ep0 DR data retained for ruling
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-14 18:52:09 +02:00
admin 65790672d5 doorstep walk on ISO 1.27.x: 1 intervention (R-505), STOP before publish
gates / gates (push) Successful in 17s
ISO 1.27.1 gated PASS and proven live: first-boot console Felhom-only,
pvebanner masked across a proven reboot. Hub v0.113.0 hand-over copy live
(R-497 closed). Full first hour walked again on customer tester-1 (three
disks + one disk): deploy, use, backup, removal, byte-identical restore,
power cut, typo all PASS. The tunnel gives a fresh box no routes: 12/12
503 from DooPlex (R-505); the record has no e-mail (R-508). Rows R-507,
R-508 filed; R-496/R-495 fixed/answered awaiting publish; day-0 A.1 no
longer claims the controller creates hostnames (R-506). NOT PUBLISHED.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-14 18:25:01 +02:00
admin 27e8ec860c iso 1.27.1: the FIRST boot is Felhom's — postinst masks pvebanner and writes /etc/issue
gates / gates (push) Successful in 18s
The 1.27.0 proof install (VM 331, screen s20) still showed Proxmox's
":8006" block on the first boot: pvebanner.service ran before
felhom-bootstrap could mask it, and the Felhom text lost its o/u double
acutes (painted before the Latin-2 font loads).
- postinst: mask pvebanner by symlink (a file act, valid in the chroot) and
  write /etc/issue; both guarded, still exit 0 (G8 self-checks unchanged).
- bootstrap: same text, byte-identical, no o/u double acutes (second line).
- harness: PI scenario (postinst in a container) + no-accent checks, red
  first against the 1.27.0 code; 55/55 green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-14 17:55:08 +02:00
admin 63f29c6ad8 hub v0.113.0: deploy (R-497 passphrase hand-over copy)
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-14 17:28:53 +02:00
admin 6fd8c87516 doorstep: console is Felhom's (ISO 1.27.0 source), passphrase hand-over copy (hub 0.113.0 source), rulings
gates / gates (push) Successful in 19s
Phase 0: the public ISO never auto-installs by construction (no answer.toml,
G1); the operator re-affirmed the interactive installer 2026-09-14.
- felhom-bootstrap.sh: mask pvebanner.service, write a Hungarian /etc/issue
  (no :8006 admin URL); pairing banner names the Tulajdonosi jelmondat and
  paints through the CONSOLE_DEV seam (R-496). Harness: 8 checks, red first;
  fake hub now sends a pairing code (the banner was never tested, R-502).
- hub: created flash + Credentials block tell the operator to hand the phrase
  over; the self-bind mail names the operator (R-497). Tests red first.
- iso-release-gate G14-G16; domain ruling in 01-topology + CONTEXT; R-494
  narrowed to P3; R-502..R-504 filed; volunteer guide and day-0 A.2 aligned.
ISO_VERSION 1.27.0 (not built, not published).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-14 17:26:53 +02:00
admin 8c7f882d1c drill 0242: teardown complete in three layers; R-501 filed
gates / gates (push) Successful in 19s
Machine: VM 330 destroyed with its disks. Host: ISO and scratch removed,
~8.5 GiB returned on nvme-scratch. Hub: customer drill0242 DELETED via the
cascade (journal #17), pages 404; its ep0 WireGuard peer dropped at the next
full-list push. Evidence pulled before the destroy. R-501: the documented
CI-check recipe reads only the last jobs page, which is not in id order.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-14 16:40:59 +02:00
admin 38848ffbeb drill: a stranger's first hour on 0.242.0 — 1 intervention, not ready for a volunteer
gates / gates (push) Successful in 21s
Golden 0.242.0 baked, round-trip verified and vouched (cadence rule, R-468).
Fresh box from the public ISO on demo-hp: landed on the vouched set, two apps
deployed and used, backup, remove, byte-identical restore, power cut and code
typo all PASS. Stopped for a volunteer by R-493 (no instructions) and R-494
(the setup mail's dashboard link has no DNS; intervention I1). R-493..R-500
filed. Capability map: first-hour row added (PARTIAL), journey row scoped.
Stopgap Hungarian volunteer guide written. Hub teardown layer pending.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-14 16:17:47 +02:00
admin 41590f8ee6 second night: scratch guest 9202 built (R-481 CLOSED, persists); controller v0.242.0 delivered (R-487 R-491 R-490 R-476 R-456 CLOSED, R-489 re-scoped); R-492 filed; rotation restarted from bentopdf; morning note
gates / gates (push) Successful in 19s
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 23:06:21 +02:00
admin 72ee053a9e evening: rules file in place; R-483 CLOSED (catalog, operator-confirmed); R-479 CLOSED (controller v0.241.0, proven live); R-481 blocked on the one-host-one-customer model; R-491 filed
gates / gates (push) Successful in 18s
2026-09-13 21:57:47 +02:00
admin 550fd84754 rules: unprompted-work.md declares unconditional: true (the instructions gate requires a scope or that declaration)
gates / gates (push) Successful in 19s
2026-09-13 21:31:33 +02:00
admin 59bc1636a4 rules: unprompted-work.md — the rules for goal and nightly sessions, byte-identical in all three repos 2026-09-13 21:29:45 +02:00
admin 8914ab089e R-456 (doc half): a partly-dead stack is not a boot orphan — the rule written in 02-controller-module-map.md; the test pin is owed to the next release
gates / gates (push) Successful in 19s
2026-09-13 19:48:11 +02:00
admin bcb65984cd R-465 audited and closed (one inert, unreachable reader); R-490 opened: the monitoring memory card never renders
gates / gates (push) Successful in 18s
2026-09-13 19:46:37 +02:00
admin 321770d9d6 R-452 CLOSED: the catalog-since gate (hook-enforced); 09 §8.2 limitation lifted; STATUS note updated
gates / gates (push) Successful in 18s
2026-09-13 19:42:42 +02:00
admin 5e8a82c3c4 night 2026-09-13/14: first "be a customer" rotation (adventurelog) — 7 defects found, 13 rows closed
gates / gates (push) Successful in 18s
New runbooks/nightly-rotation.md; observations_gate.py reads every section
(R-471); target-selection.md names real paths (R-461); R-93 carries the
fact that drill-r50 is gone. Register: R-473/R-474/R-466/R-471/R-453/R-461
and v0.240.0's R-477/R-478/R-480/R-482/R-484/R-485/R-486 closed; R-481,
R-483, R-487, R-488, R-489 opened. 09 §6.1, 07 §6, CONTEXT, STATUS note.
Evidence: audits/nightly-2026-09-13-adventurelog/, audits/v0240-2026-09-13/.
2026-09-13 19:38:00 +02:00
admin 681c3d6a6d docs: rulings 7 and 8 shipped and proven live (R-470/R-472/R-475 CLOSED); R-477..R-480 opened
gates / gates (push) Successful in 21s
Hub v0.112.0 serves a floor above the golden with a declared MinAgent;
controller v0.239.0 reached both demo boxes by that floor in 14 s and 15 s
and updates on any backup tier. 09 §3 decisions 7 and 8, §6/§6.1; 07 §6
line; capability map row; STATUS items 15/16 done and the cadence line
corrected; CONTEXT; register: R-470/R-472/R-475 compressed to CLOSED-ITEMS
(full text at 2f5d3af), R-477..R-480 opened, R-474 reproduced a third
time. OPEN-ITEMS 431689 -> 432156 bytes, CLOSED-ITEMS 118051 -> 120598.
Evidence: documentation/audits/rulings-r472-r475-2026-09-13/.
2026-09-13 17:54:23 +02:00
admin 2f5d3af6f9 deploy: hub 0.112.0 (R-472 declared MinAgent floor)
gates / gates (push) Successful in 20s
2026-09-13 16:49:15 +02:00
admin f181efd6a7 hub v0.112.0: a floor carries a declared MinAgent past the golden (R-472)
gates / gates (push) Successful in 18s
Operator ruling 2026-09-13. Above the vouched golden, a floor saved with a
declared MinAgent is served under the same agent comparison; an undeclared
one is still held beyond the golden. The declaration is stored beside each
floor as FLOOR=MINAGENT so it never carries to a later floor. Both floor
forms require min_agent above the golden (flash floor_needs_min_agent,
nothing stored). The Hosts page and the API log name the source.

Vouch path and R-120 gate untouched. Scenarios A-E tested; red-proofs A
and C in documentation/audits/rulings-r472-r475-2026-09-13/.
2026-09-13 16:47:11 +02:00
admin 5ef0f52bcd Slice 4 shipped (R-448/R-443/R-439 CLOSED, proven live); R-472..R-476; the floor-between-bakes claim corrected
gates / gates (push) Successful in 19s
Controller v0.237.0-v0.238.1: the Update button is a guarded job — refusals, backup-first when the
proven Tier-2 copy is stale, safety dump, pin, pull (pin back on failure), health, HOLD on failure.
Proven live on demo-hp: A, B, E, F, H and the restore walk (audits/slice4-2026-09-13/).

Correction to this morning's pages: between golden bakes the hub HOLDS a floor above the vouched
golden, so a release does not reach the fleet by floor (R-472, operator decision). Corrected in the
runbook, STATUS, CONTEXT, R-468 and the gate docstring.

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 12:30:17 +02:00
admin abe567e14d REPORT: felhom.eu CI run 537 green for 4727aaa
gates / gates (push) Successful in 18s
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 10:20:48 +02:00
admin 4727aaa5ad register: compress R-459 and R-467 to CLOSED-ITEMS (full text at ae59c31); REPORT sizes + CI
gates / gates (push) Successful in 20s
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 10:15:46 +02:00
admin ae59c31a84 R-459 CLOSED (MariaDB converts itself, proven by harness + live), golden 0.236.0 (R-467), the golden waiver (R-468)
Operator rulings 2026-09-13, both shipped the same day:
- MariaDB finishes its own conversion (catalog eec1228/bd32830/3525e35). Harness E3/E3b `proven`
  with engine_state_after "already upgraded to 12.3.3-MariaDB [exit=1]", the skip line gone, C3
  still `failed`; landed on demo-hp through the real 15-min cycle, nothing recreated, one deliberate
  restart logged "MariaDB upgrade not required" with the app serving. Evidence:
  documentation/audits/r459-close-2026-09-13/. The engine-major rule + gate keep every engine
  inside its major until Slice 4 (R-448) — removal tracked as R-469.
- Goldens on a cadence, not per release. golden_currency_gate.py reads a dated waiver
  (documentation/tests/golden-waiver.yml, <= 14 days, row-bound): valid + BEHIND -> loud advisory,
  exit 0; expired -> red again naming the date; UNRECORDED (R-385) never covered; malformed -> 2,
  never 0. Tests cases 5-15 incl. the R-421 decoy; red-proof old-vs-new on the real behind tree.
  R-242's vouch half stays open. Cadence in RUNBOOK-manual-build.md §4.2 + the checklist.
- Golden 0.236.0 baked, round-tripped, vouched, floor raised 0.232.0 -> 0.236.0
  (documentation/tests/golden-0.236.0-2026-09-13/) — the last per-release bake; the waiver was
  issued AFTER it landed. No --no-verify anywhere in this session.

Rows: R-459 CLOSED, R-467 CLOSED, R-242 narrowed; R-468/R-469/R-470/R-471 opened. 09 §3 gains
decisions 5 and 6; STATUS items 11 and 12 closed; CONTEXT records the cadence ruling.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 10:14:37 +02:00
admin 4b2e5608c2 R-442 CLOSED (controller v0.236.0): "delete my data too" deletes it or refuses; R-465..467 opened
gates / gates (push) Failing after 18s
- OPEN-ITEMS: R-442 row removed; R-465 (six remaining Paths.HDDPath readers — audit),
  R-466 (recovery-unit residue after "delete backups"), R-467 (v0.236.0 owes a golden).
- CLOSED-ITEMS: R-442 compressed, reasoning kept, fleet shape now established.
- 00-capability-map: lifecycle row narrowed (Campaign 3 proved remove removes the APP,
  not the data) and re-proven from audits/R442-2026-09-13/.
- STATUS: item 13 in plain language; item 7 closed (ruled 2026-09-02, 09 §3).
- audits/R442-2026-09-13/: live evidence (A, C, D bodies, controls, log window, teardown).

Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 09:07:08 +02:00
admin d6837d98ee SPIKE R-459: the skipped MariaDB conversion is stable, and the trade it implied does not exist
gates / gates (push) Successful in 20s
Outcome A, qualified. Not B and not C.

It does not degrade: 5 of 5 restarts of 12.3 on an 11.6 datadir, readback passed
every time, mariadb_upgrade_info unchanged, the entrypoint line never escalated
past [Note]. It also never heals - the engine answers 'Major version upgrade
detected from 11.6.2-MariaDB to 12.3.3-MariaDB. Check required!' on every start
and will forever.

The trade R-459 was expected to produce is not real. Converting properly SUCCEEDS
across the multi-major jump, takes 7 seconds, backs up the system database
unasked - and putting 11.6 back afterwards STILL starts and serves the data. So
the operator is being handed a cheap correction, not a choice between a correct
engine and a reversible one.

The exit-code polarity was measured rather than read: 0 means the upgrade IS
needed, 1 means it is not. Assuming either the flag name or the polarity would
have inverted the headline. And run without credentials the same command returns
a confident-looking FATAL ERROR that is an auth failure.

R-464: after converting and going back, the entrypoint prints 'MariaDB upgrade
not required' on a state the same engine calls an unsupported downgrade. The
obvious cheap instrument for R-459 would have been to grep for that line, and it
would have reported fine for the broken case.

R-463: the PostgreSQL analogue, deliberately NOT measured here. 11 templates, 8
on postgres:16-alpine, register grep for pg_upgrade returns zero. The two engines
fail in OPPOSITE directions - MariaDB skips quietly, Postgres refuses to start -
so that one cannot hide; it presents as eight apps down at once.

No template changed. Teardown all three layers, hub checked rather than asserted,
local-lvm 30.53 percent before and after.
2026-09-06 17:42:19 +02:00
admin a1a6c73fe1 SPIKE: an upgrade test that runs again — and a real defect in our own bookstack template
gates / gates (push) Successful in 19s
R-449. Until today one app upgrade out of 53 had ever been measured, by hand, and
the whole update arc was designed against that single data point.

C3 first: the negative control, whose TO image exits immediately, came back
failed. That is what makes the greens mean anything, and it cost 556s because a
negative is only honest if it waits out the full settle window.

Seven edges, three apps. All five real catalog upgrades kept the customer's data.

The finding that changes an assumption the arc was carrying: whether an upgrade
can be UNDONE is a property of the individual APP, not of upgrades. Docmost
refuses - 'corrupted migrations: previously executed migration
20260213T085259-notifications is missing' - and privatebin does not. That
reproduces the Nextcloud result on a second app by a DIFFERENT mechanism, so the
struck word 'rollback' now rests on two measurements instead of one.

The finding nobody was looking for, R-459: our own bookstack template moves
MariaDB across a major and sets no MARIADB_* env at all, so the engine logs that
the datadir upgrade it requires is being skipped, and serves anyway. The cause is
assigned rather than guessed - the app half alone produces no upgrade line, both
edges that move the engine produce it - which is exactly what decomposing E3 into
E3a and E3b was for. It also explains why E3's abort looked like it worked: the
datadir was never converted. Whether that ever breaks is NOT established, and the
row says so.

Also opened: R-460 (bookstack's file half cannot be seeded headlessly), R-461
(target-selection.md names a venue that does not exist and fences a VM that is
gone), R-462 (the widening, costed with this run's real numbers - and the cost is
dominated by fixtures, which do not amortise).

Teardown all three layers, hub checked rather than asserted. local-lvm read 30.50
percent before and after. The capability map was deliberately NOT edited: this
measured apps, not the product.
2026-09-06 11:48:57 +02:00
admin 417df06f35 slice 3 docs: the ruling, the shipped mechanism, and four rows closed
gates / gates (push) Successful in 17s
09-update-architecture.md gains the fourth dated operator ruling (2026-09-06,
Option 1) and its section 5 is rewritten from a proposed shape into the shipped
one: the pin, the stored definition, the render table, the four writers, the
startup ordering, and the trap this slice set for slice 2 - the live compose file
is now the frozen one, so a badge comparing against it would answer Naprakesz on
exactly the apps that are behind.

02-controller-module-map.md said 'copy compose + .felhom.yml'. That stopped being
true today, so it is corrected, and the two sections describing the old seam now
carry a banner saying they describe v0.234.0 and below - kept because every box
under v0.235.0 still behaves that way and because they are the measured account
of why it changed.

R-447, R-441, R-438 and R-455 closed and compressed into CLOSED-ITEMS; R-458
opened for the .felhom.yml asymmetry, with what would settle it by measurement.

Live evidence: two real catalog pushes travelling the real 15-minute cycle, both
reverted, the tree byte-identical afterwards. The restart that used to take 18.3
seconds and pull a new image now takes 0.1 seconds and pulls nothing.
2026-09-06 10:37:33 +02:00
admin bc47dd4ef9 v0.234.0: a known limitation written on 2026-09-02 was a defect by the next morning
gates / gates (push) Successful in 18s
The operator looked at demo-felhom and found OpenGist - up 15 hours, running
exactly the catalog pin, showing no badge at all. 09-update-architecture.md had
recorded that as an accepted limitation the day before: 'the fleet view fills in
gradually'. On a quiet box gradually means never, and a feature that fills itself
in on an event nobody triggers is, on the quiet installations, not shipped. That
limitation row is now struck with the reason kept.

The living document gains slice 1b, the two admission rules of the backfill (it
never overwrites, and it refuses to seed a partial observation because the badge
reads a service-count mismatch as BEHIND), and the note that the same field having
two writers with two different admission rules is deliberate.

Live evidence added: all nine apps already had records by the time 0.234.0 was
ready, so the natural fleet state could no longer exercise the new code - said
plainly rather than papered over. The pre-0.233.0 shape was recreated on demo-hp
by stripping two records; the backfill re-seeded exactly those two with digests
matching independently-read ground truth and left the other seven alone.

The refusal half was deliberately NOT staged live: it needs a degraded app, and
manufacturing one risks the false-customer-email class that already cost 61 mails
(R-330). Unit-tested with a red-proof, and recorded as unproven-live.

R-457: a test that hardcodes a date and asserts an age derived from it is green
only on the day it is written. Mine was, and it went red overnight. Six other
files carry both a date literal and time.Now() - named as candidates, not accused.
2026-09-03 12:01:44 +02:00
admin 7941b0c159 R-455: the mirror base images are MEASURED byte-identical to Docker Hub
gates / gates (push) Successful in 19s
The row carried 'KNOWN, not MEASURED' because the throttle was still in force. It
cleared 40 minutes later and the check was run: docker pull from Hub answered
'Image is up to date' for both bases - Hub's own manifest resolved to the images
already local, the ones v0.233.0 was built from - and the manifest bodies are
identical between registries.

This matters for the row's own decision: the mirror being a sound source is what
makes 'sanction the mirror in build.sh' a real option next to 'get a Docker Hub
login', rather than a hope.
2026-09-02 21:03:40 +02:00
admin 0705942783 the badge IS proven live, and the 'stale password' finding was mine, not the box's
gates / gates (push) Successful in 16s
I reported that the vaulted dashboard password no longer worked on either demo
box, and quoted the controller's own 'Failed login' as the discriminator. The
password was fine. ~/.config/credentials quotes its values with SINGLE quotes and
my sed stripped only double quotes, so the quote characters went out as part of
the password. The operator corrected it in one line; one retry returned 302.

The instrumentation lesson is the finding and R-453 now carries it: 'Failed login'
separates wrong-password from wrong-Host-header, and that is ALL it separates. It
cannot tell a wrong password from wrong password HANDLING, and I read it as if it
could. This is the second time this file's quoting has produced a confident wrong
verdict, so the fix is one shared extraction helper, not a resolution to be careful.

With the session recovered, the badge is validated on live pages: Naprakesz twice
on /stacks and on /apps/bookstack; NO badge at all on /apps/docmost (a deployed app
with no record - absent is UNKNOWN, not current); and 'Frissites elerheto - 52
napja' on both surfaces, the age being real arithmetic on bentopdf's catalog_since.
The behind state was staged by editing one compose tag, with no restart and no
up -d, and reverted byte-identically (sha256 equal, diff empty, container never
touched). Capability-map row upgraded to PROVEN-LIVE with the one unexercised
badge state named. STATUS item 9 now needs nothing from the operator.
2026-09-02 20:47:53 +02:00
admin e86cf42e0b three more register rows: the observations gate refused a report that filed none
gates / gates (push) Successful in 18s
R-454 five gofmt-unclean internal/web test files at the BASELINE, with no gate
that would ever notice; R-455 DooPlex has no Docker Hub login and the
unauthenticated ceiling now blocks a BUILD, not just the catalog's resolvability
gate; R-456 a partly-dead stack is not a boot orphan and that rule exists in no
document, so this session re-derived it by watching a repair not happen.

All three were observations in felhom-controller/REPORT.md, which is overwritten
every session. The controller's observations gate refused the push until each one
either named a row or declared itself not a finding - which is exactly its job.
2026-09-02 20:36:52 +02:00
admin 6035dfcc3a 09-update-architecture.md: the update path finally has a document, and it is a living one
gates / gates (push) Successful in 17s
R-438's document half. It records how an update works AS MEASURED, quotes the
RestartStack comment that proves the restart half was CHOSEN (a design decision
is not a defect), carries the three operator rulings of 2026-09-02, strikes the
word 'rollback' (once a migration has run the old image will not start), states
the target shape, and lists the seven slices with a status each.

R-438 and R-440 amended and BOTH STAY OPEN: the mechanism is documented, not
changed. Nothing closed, so CLOSED-ITEMS.md is untouched.

Eight new register rows, 194 -> 202: R-446 (Naprakesz can be false for the 23
floating pins), R-447..R-451 (one per remaining slice, with a rank and an owner),
R-452 (no gate enforces catalog_since - the runner fetches at --depth 1), and
R-453 (the vaulted dashboard password is stale on BOTH demo boxes, which is what
stopped the badge render from being validated live).

Live evidence for slices 1 and 2 in documentation/tests/. The record is PROVEN
LIVE through the boot reconciler on demo-hp - one entry per compose service,
digests matching ground truth read independently. The badge RENDER is not, and
the five attempts are listed rather than summarised.
2026-09-02 20:32:30 +02:00
admin 56c7e373a3 SPIKE: what an app update actually does, and which other paths do it too
gates / gates (push) Successful in 18s
THE GATE IS ANSWERED: YES. compose up -d upgrades an app whose compose file has
already moved, and the Restart button does it — 18.3s with a network pull when
the target image is absent, 0.5s when present, against a negative control that
did not even recreate the container. The boot reconciler does the same thing
unattended when an app fails to come back (bootrecon.go:269 -> StartStack).

AND ONE FEAR IS SMALLER THAN THE BRIEF CLAIMED: a plain power cut upgrades
nothing. Docker restores the old containers and the reconciler logs 'no
boot-orphaned apps (nothing to start)'.

AND ONE IS BIGGER: app data CANNOT be rolled back. Once a migration has run,
the old image refuses to start — Nextcloud: 'the version of the data (32.0.9.2)
is higher than the docker image version (31.0.14.1) and downgrading is not
supported'. 'Rollback' is the wrong word for this arc and is struck.

Phases 0-6 all run, on demo-hp (Tier 0, disposable). Phase 6 run on operator
confirmation. Peti's box was never contacted. No production code written in any
repo: felhom-controller is at 960d29b0612c before and after, tree clean, and
build/vet/test are green — run at the end to prove exactly that.

Register: R-438/439/440 updated with live evidence; R-441..R-445 opened
(restore-vs-sync conflict; remove_hdd_data inert with no paths.hdd_path, 128 MB
left behind; update reports success over a broken app; no fleet fstrim; hub
telemetry outlives the app). Capability map gains three measured rows. The
mechanism is written down in 02-controller-module-map.md as MEASURED BEHAVIOUR,
not [DESIGN] — the operator has not ruled. STATUS.md carries the one decision.

Five claims in the brief are named as wrong, including two of my own method.
2026-09-01 21:35:32 +02:00
admin ac079b8c43 backlog: file R-438/R-439/R-440 — the app-update spike's three register rows
gates / gates (push) Successful in 19s
R-438 (P1-HIGH, Viktor rules): the catalog syncer rewrites a DEPLOYED app's
docker-compose.yml on a 15-minute cycle with no deployed check, and no
architecture document records the consequence. Mechanism behind R-40.

R-439 (P3-LOW, CC): Router.actionStack checks the R-379/R-380 restore hold for
start and restart only; update falls through to compose up -d.

R-440 (P2-MEDIUM, CC): 23 of 79 catalog image lines carry a tag with no patch
version, so an update is not reproducible. Measured over 29edad9c5bf4.

Phase 0 of SPIKE-app-update-2026-09-01. Documents only; no code touched.
2026-09-01 19:32:09 +02:00
admin 1d59353df4 provider questions: arm them against the Storage Box / Storage Share conflation (operator-found)
gates / gates (push) Successful in 18s
The operator noticed the "can I restore specific files from within a backup?" FAQ lives under
storage-share, not storage-box, and asked which product it covers. It is Storage SHARE only —
a managed Nextcloud — and it never mentions Storage Box. Its own text gives it away: Nextcloud's
data cache, a database dump, the konsoleH web interface.

The two products document OPPOSITE answers:
  Storage BOX   (ours) "You can download individual files or entire directories as usual"
  Storage SHARE (not)  "we only support restores for the full backup ZFS snapshot"

That matters because a web search for the obvious phrasing surfaces the SHARE page and it reads
like a definitive NO — so a support agent could answer Question 1 from the wrong page and push
R-95 to the top of the register for no reason. Question 1 now names the product, quotes the
Storage Box line, and states up front that we know what the Share FAQ says. A warning block at
the head of the file tells the reader to check which product any full-snapshot-only answer is
about before acting on it.

Verified by grep: nothing in this repository ever leaned on the Share claim. The only vendor
line cited anywhere is the Storage Box one.

R-436 strengthened from the same source the operator supplied: the rclone-over-SSH restic
backend is OFFICIALLY DOCUMENTED, not merely advertised in a shell banner --
"we support the restic backend, which is provided by Rclone over SSH". And the same page settles
that the docs cannot answer the caveat: neither its Rclone nor its Restic section mentions
append-only at all, so nobody need re-read the documentation hoping for it. Unlooked-for
corroboration: that page's port-23 command table matches, item for item, the help output
measured live on our own sub-account.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-09-01 18:59:37 +02:00
admin f380c6d43c REPORT: correct the register line count (620 -> 688, one measure) and record the CI run ids
gates / gates (push) Successful in 18s
Both earlier commit messages carried a wrong line count - 621->688 and 688->700 - because I
mixed wc -l with a Python line split and then carried the error forward. Corrected in the
report rather than by rewriting history, and named there.

CI verified by run id against head_sha: 500, 501, 502 all success.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-09-01 18:43:20 +02:00
admin 1a1b32b3dd REPORT + R-437: the beta stopping line recorded, and the live trigger declined with its reason
gates / gates (push) Successful in 19s
REPORT.md carries the deployed sentence quoted from the RUNNING binary (kubectl cp + byte
grep, both controls), the stopping line as it reads in all three places, the enumerated
deferred set, the two provider questions, the register census, and the ArgoCD verification.

R-437 filed: the register compression sweep is OWED and was deliberately not run here.
Measured first — 12 of 181 rows / ~25 KB of 316 KB (about 7%) carry a closed leading
verdict — so it buys little and touches everything, and it is the exact operation that
misfiled seven rows in August (R-378; the seventh, R-87, sat wrong for nine days, R-405).
The row carries the scope so it can be picked up cold.

The live alarm trigger was NOT run and the report says so in its own section rather than
substituting quietly: this alarm only fires on a real fall in a real customer's snapshot
count, so firing it means either deleting real backups or POSTing a falsified report
claiming demo-hp lost its own. That would write a fabricated point into a customer's report
history, move its latch and baseline, and mail the operator a second alarm about a real box
hours after the first one already confused him. Covered instead by the deployed-bytes proof
plus three red-proofed tests driving saveReport -> Check -> notify. What remains unproven is
named: that the dispatcher delivers THIS wording to a mailbox.

Register 688 -> 700 lines; 182 rows; open-state 170.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-09-01 18:41:50 +02:00
admin 0f65f7a197 manifests: hub 0.111.0 -> 0.111.1 (R-434, the alarm's withdrawn promise)
gates / gates (push) Successful in 18s
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-09-01 18:35:58 +02:00
admin db38f4c800 hub v0.111.1: the alarm stops promising a rescue that does not exist, and the arc is closed for beta
gates / gates (push) Successful in 17s
R-434 CLOSED — and the row's own "blocked on R-433" verdict was wrong, which is the point.
The fix is a DELETION, not a replacement: withdraw the promise instead of swapping it for a
new one, and the sentence is true under every possible answer to the provider questions, so
it never needs a second rewrite. A replacement would have been blocked; a withdrawal is not.

  was:  "...still hold the older copy, so this is recoverable file-by-file; it is NOT
         confirmed data loss. Check whether a deletion ran on the box before restoring."
  now:  "...still hold the older copy. The route back out of them is not yet established,
         so treat this as neither confirmed data loss nor confirmed recovery. Get in touch
         before restoring anything, and check whether a deletion ran on the box."

It must not swing the other way either: "your backups are gone" is still usually false.
Clause (a) — the box cannot WRITE into the snapshot area — stands and is re-confirmed.

Tests: offsite_r434_test.go, three, all driving the production path so they assert the
sentence an operator RECEIVES. ASCII-only fragments, positive and negative controls.
RED-PROOF: restoring the v0.111.0 sentence failed all three, on every fragment, with the
offending sentence printed. TestR431_FiresOnAMassDeletion asserted "NOT confirmed data
loss" and caught this fix correctly; its wording fragment is REMOVED rather than updated,
so the wording keeps ONE home.

R-435 written into the detector's own documentation, no threshold changed: it sees a mass
deletion, not one app being wiped (69 across 9 apps -> ~35 needed, one tag is ~9, and
forget --prune groups by host,tags). Says explicitly not to lower the numbers.

THE STOPPING LINE, in all three places — register, 07 section 8 head, STATUS.md.
Deferred set ENUMERATED, not described: 07 section 8 rows 4, 8, 9, 10, 11 (+11b), 12,
each tagged [BETA-DEFERRED]. A number in the brief was wrong and is corrected in place:
six rows are DEFERRED, ELEVEN carry a blank RTO (4,5,8,9,10,11,11b,12,13,14,15); the other
five are blank for reasons that are not deferred work, and row 15 is an open DEFECT (R-104)
that the stopping line does NOT cover. NO STATUS MOVED — nothing was proven today.

Two provider questions drafted, not sent, no API called (11-D stands):
documentation/runbooks/provider-questions-2026-09-01.md, linked from R-95 and R-433, and
tracked by a dated DUE-CHECKS row (2026-09-15) — the 2026-07-27 check that sat unconfirmed
for 36 days is the scar that block exists for.

R-95, R-433 BLOCKED-ON-PROVIDER. R-95's one-day demotion on a clause that did not hold is
recorded; the proposal to rank it back near the top is stated and NOT acted on. R-430 marked
LATENT with its trigger: it becomes live the moment delete is withdrawn, so it is a
precondition on the R-95 build, not a follow-up. The stale ranking paragraph ("armed",
"zero snapshots") is corrected in place, order unchanged.

Register 621 -> 688 lines; 181 rows throughout; open-state 170 -> 169.
No controller or agent change. No golden owed, no floor change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-09-01 18:34:52 +02:00
admin 10c223bdfe DRILL R-95: the recovery route does not exist — stopped before the destructive phase
gates / gates (push) Successful in 19s
The drill was to delete demo-hp's off-site history and get it back out of a Storage
Box snapshot, filling row 10's blank RTO. Phase 1 found there is nothing to get it
back from: no snapshot is reachable from a sub-account BY ANY NAME.

Measured, read-only, no delete verb issued against any live store:
- 777,600 exact names in the vendor form YYYY-MM-DDTHH-MM-SS, nine full days at
  second granularity, plus 126 alternative shapes -> ZERO hits.
- The control is what makes that mean anything: the identical 600-name batch shape
  with one real path appended returned it, 6 of 6.
- Structural cause: /home (u629488-sub3) is st_dev 0,82; /.zfs/snapshot is st_dev
  0,276; /home/.zfs does not exist. A snapshot under /.zfs/snapshot belongs to a
  different dataset than the one holding felhom-repo.
- Three tools agree with controls in the same run: SFTP, the port-23 shell,
  rsync --list-only.

So yesterday's re-scope splits: clause (a) "the box cannot write into the snapshot
area" STANDS and is re-confirmed; clause (b) "the rest is recoverable file by file"
is NOT SUPPORTED. STOPPED before Phase 2 on the operator's ruling — with no recovery
leg the deletion would have destroyed real history to buy only an alarm test that
could not fire at the specified size. Store verified untouched at 69 snapshots.

R-432 ANSWERED (negatively; its panel-read next step withdrawn as unnecessary).
R-433 no snapshot reachable by any name — decides R-95's remedy and its rank.
R-434 the drop alarm's text promises a file-by-file recovery that cannot be performed.
R-435 the drop detector is blind to a single-app deletion (>50% of 69 needed, ~9 given).
R-436 LEAD: the provider offers `rclone serve restic --stdio` and restic 0.14.0 speaks
      `rclone:` (measured, controlled) — real prevention may need no new machine, IF
      the vendor pins --append-only. Ask before building.

07 §8 row 10: text corrected, status NOT moved, RTO still blank.
No code, no version bump, no image, no golden. REPORT.md's only copy of the R-331
report preserved as REPORT-r331-backup-card.md before overwrite.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-09-01 16:55:50 +02:00