From 51ea1dbaade1069b04cbd99be4d0fd98945d9b60 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Wed, 7 Oct 2026 07:11:39 +0200 Subject: [PATCH] R-892: SSH works (operator key); the walk needs the box's dashboard password Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- STATUS.md | 2 +- .../night-burndown-2026-10-06/MORNING-NOTE.md | 3 +-- .../night-burndown-2026-10-06/NIGHT-LOG.md | 1 + .../s3/s3-tester1-readonly.txt | 16 ++++++++++++++++ documentation/backlog/OPEN-ITEMS.md | 2 +- 5 files changed, 20 insertions(+), 4 deletions(-) create mode 100644 documentation/audits/night-burndown-2026-10-06/s3/s3-tester1-readonly.txt diff --git a/STATUS.md b/STATUS.md index 21b0cdae..3c304e6a 100644 --- a/STATUS.md +++ b/STATUS.md @@ -19,7 +19,7 @@ held catalog branch).** **Needs you (none urgent):** 1. Release order today: hub, agent, controller, then the catalog branch. My pick: yes, after Parts B and D. -2. The Tester 1 key: `ssh-copy-id -i ~/.ssh/id_ed25519.pub root@192.168.0.154` from this server. +2. The Tester 1 box: the key works now; the update test needs the box's dashboard password in `/home/kisfenyo/.felhom-tester1/.ctlpw` (0600). ## Night (2026-10-06): the night-free parts done; the rest after 08:30 diff --git a/documentation/audits/night-burndown-2026-10-06/MORNING-NOTE.md b/documentation/audits/night-burndown-2026-10-06/MORNING-NOTE.md index 3ab84feb..1238f72c 100644 --- a/documentation/audits/night-burndown-2026-10-06/MORNING-NOTE.md +++ b/documentation/audits/night-burndown-2026-10-06/MORNING-NOTE.md @@ -41,8 +41,7 @@ Rows before: **138**. Rows after: **137**. Opened: **2**. Closed: **3**. ## 5. What failed, and why -- **The Tester 1 box: still no entry.** Your key copy at 20:20 used the key of your own login, not this server's - key. Fix: run `ssh-copy-id -i ~/.ssh/id_ed25519.pub root@192.168.0.154` from this server as kisfenyo. +- **The Tester 1 box:** your key works now (07:10). The update test still cannot run: it signs in to the box's dashboard, and this server has no password for it. Put that password in a 0600 file, e.g. `/home/kisfenyo/.felhom-tester1/.ctlpw`. - **The hub release did not run.** The safety check of my session refused the image build at 00:12. I stopped there. The hub still runs 0.140.0, unchanged. - **I could not read the hub's list of apps per box.** The same safety check refused it. So every catalog change diff --git a/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md b/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md index 0fda40ad..3422aba0 100644 --- a/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md +++ b/documentation/audits/night-burndown-2026-10-06/NIGHT-LOG.md @@ -54,3 +54,4 @@ no reboot. | R-330 | **wire half built** in three repos (agent sends 187/188/199, controller decodes, hub models; carried only — no verdict reads them); 5 red-proofs | 15 | hub/agent/controller (this batch) | | (hub release) | **NOT DEPLOYED.** 00:10 the release commit (`hub/CHANGELOG.md` v0.141.0: R-366, R-31, R-330 wire) was pushed, green; at 00:12 the image build (`build/felhom-hub/build.sh 0.141.0 --push`) was **refused by the session's permission check** („Production Deploy"). Stopped there (brief Rules 3), no workaround; the manifest still names 0.140.0, nothing changed on the hub. The CHANGELOG says so. The day session builds and deploys 0.141.0 | 5 | (this batch) | | R-872 | **closed** — the 05:00 run judged Tester 2 on the longer lines (dump missed=1, backup missed=0 correctly) and the mail reached the operator inbox (second channel) | 10 | (this batch) | +| R-892 | 07:10 (operator: key added, „continue"): SSH works, app list read; **the walk is blocked again** — no dashboard password for this box on DooPlex. Nothing changed on the box | 6 | (this batch) | diff --git a/documentation/audits/night-burndown-2026-10-06/s3/s3-tester1-readonly.txt b/documentation/audits/night-burndown-2026-10-06/s3/s3-tester1-readonly.txt new file mode 100644 index 00000000..59648dd2 --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-06/s3/s3-tester1-readonly.txt @@ -0,0 +1,16 @@ +2026-10-07T07:10:37+02:00 +$ ssh root@192.168.0.154 hostname +felhom +$ pct exec 9201 -- docker ps --format "{{.Names}} {{.Image}} {{.Status}}" +bookstack lscr.io/linuxserver/bookstack:26.09.1 Up 3 hours (healthy) +bookstack-db mariadb:12.3 Up 3 hours (healthy) +cloudflared cloudflare/cloudflared:2026.9.3 Up 45 hours (healthy) +felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.301.0 Up 15 hours (healthy) +filebrowser gtstef/filebrowser:1.5.6-stable Up 45 hours (healthy) +paperless-postgres postgres:18-alpine Up 3 hours (healthy) +paperless-redis redis:7-alpine Up 3 hours (healthy) +paperless-webserver ghcr.io/paperless-ngx/paperless-ngx:2.20.15 Up 3 hours (healthy) +privatebin privatebin/pdo:2.0.6 Up 3 hours (healthy) +traefik traefik:v3.7.13 Up 45 hours +$ ls /opt/felhom/stacks + diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 516a0335..94f71a4e 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -283,7 +283,7 @@ stopping line that lies. |---|---|---|---|---|---|---|---| | **R-733** | Process & tooling | P3 | **[P3-LOW] The test bench has NO swap and the boxes have 512 MiB — so a box proof can pass on swap where the bench fails, and nobody records whether a customer guest has swap.** MEASURED 2026-09-30 (R-732): immich's first start was OOM-killed 61–104 times on the bench (swap 0) and passed on 9202 by swapping ~108 MB; the bench given 512 MiB swap passed too. demo-hp 9201, 9202 and demo-felhom 9201 all read `swap: 512`; the golden's guest config is not recorded in its bake evidence, so a customer guest's swap is NOT measured. The harness's memory watch judges `anon` against the limit and never reads `memory.swap.current`. **Needs:** the golden's `swap` read and recorded; the box walk and the harness report `memory.swap.peak` beside `anon`; a decision whether proofs run with swap off (the stricter venue, as R-732's fix was proven). | **READY — rank P3-LOW; owner: CC (harness + golden evidence)** **2026-10-06 night: NARROWED** — the harness records each container's `swap_peak` and the venue's swap (catalog branch `night-held-2026-10-06`, `3f4611c`; SwapRecorded red-proved; bench SwapTotal 0 kB, 9202 524288 kB). LEFT: the golden's swap in its bake evidence, the box walk's record, and whether proofs run with swap off. | — | — | CC | | **R-887** | Process & tooling | P3 | **Some CI jobs are never run, and Gitea fails them ~10–13 minutes later with no log.** Seen 2026-10-05: felhom.eu job 1361 (commit `1122b5c`) and felhom-controller job 1357 (`114ff27`): every step reads `failure`, including the first fetch, the log API answers `file does not exist`, and the runner pod's log has no `task` line for them (its task ids are job id + 1). A re-run through the API ran the controller job normally (success in 32 s) but the felhom.eu job was again never picked up and failed after ~12 min. The runner pod (`gitea-system/act-runner`, image `felhom-act-runner:0.1.0`) had restarted 5 times ~142 min earlier, around the Longhorn instance-manager restart (R-882). Suspected, NOT measured: a stale runner registration claims jobs it never runs — the session's Gitea token cannot list runners (`read:admin` scope). Consequence: a red CI verdict that is not about the code, and **no failure mail** (the alarm step never runs either), so only the pull check sees it. **CORRECTED 2026-10-05 18:21 (operator's screenshot of Gitea → Site Administration → Runners): ONE runner only — ID 2, `felhom-gates-runner`, v0.6.1, label `felhom-gates`, Idle, last online „now". There is no old registration; the stale-registration guess (this row's first text and the reviewer's) was WRONG.** **RE-DIAGNOSED 2026-10-05 (round 2), from the logs that survive:** (1) **„lost in a runner restart" does NOT fit** — the runner pod last restarted 13:24:42Z (`restartCount 5`, all around the 13:20Z Longhorn restart), the lost attempts started 1.5–2.5 h later. (2) **FOUR attempts were lost, not two:** controller job 1357 (start 15:05:41Z → failed 15:18:38Z), felhom.eu job 1359 (15:15:37 → 15:28:38 — the previous session blamed that one on the BusyBox fault; the runner never ran it), job 1361 (15:33:21 → 15:43:38) and its API re-run (15:46:48 → 15:58:38). None has a `task` line in the runner log; every one was failed at a :38-second mark on a 5-minute step, 10–13 min after it was handed out — **the shape of Gitea's periodic „zombie task" stop** (a task assigned to a runner that never reports is failed after ~10 min; no log exists because none was written). (3) The runner's task ids are NOT job id + 1 (the controller re-run was task 1363). (4) **Gitea's own log for the window is gone** — the pod log starts 16:16:32Z (rotated), so the assignment side cannot be read. **Likely mechanism, NOT proven:** the runner's fetch-task request timed out on its side after Gitea had already assigned the task, so the task was orphaned. In that same hour this session polled Gitea's jobs API hard (15 pages every 15 s per wait loop) and Gitea logged „slow" requests — a plausible load cause, and the session's own. Mitigation taken: the session's CI waiter now polls once a minute. **Nothing changed on DooPlex.** **MECHANISM SEEN 2026-10-05 17:15–17:28Z, with Gitea's own log (round 2):** catalog run 1368 (`4828dc7`) — 17:15:14 the job is marked started; 17:15:16 `router: slow POST /api/actions/runner.v1.RunnerService/FetchTask for 10.42.0.42 (the runner), elapsed 3192ms`, then `UpdateRepoRunsNumbers … context canceled` and `GetActionWorkflow: EOF` — **the runner abandoned its fetch after Gitea had assigned the task**; the runner log has no line for task 1371; 17:28:39 `actions/clear_tasks.go:174 stopTasks() [W] Cannot transfer logs of task 1371` — Gitea's zombie-task stop. **The load at that minute:** an outside crawler (216.73.216.78) walking commit pages and `archive/*.tar.gz`, and THIS session's CI waiter, whose 15-page job listings took 13–31 s each. An API re-run passed in 7 s. **Done in-session:** the waiter now asks `GET …/actions/runs?head_sha=` once a minute (1 s). **Not done (DooPlex, the operator's):** the runner's fetch timeout and Gitea's exposure to the crawler. | **OPEN** **DATED CHECK 2026-10-12 (DUE-CHECKS):** if no job was lost since 2026-10-05 16:00Z (no completed job whose runner log has no `task` line / whose log API answers `file does not exist`), close. **NIGHT WATCH 2026-10-05/06 (burn-down night): 2 jobs lost of ~30 runs** — felhom.eu run 1384 (`4aa4d837`, 21:23→21:33Z, no log) and felhom-controller run 1401 (`c67b26be`, 00:45→00:58Z, no log); each re-run once through the API and each passed (2 m 05 s, 57 s). The night's waiter made one filtered call a minute. So the 2026-10-12 close condition („no job lost since 2026-10-05 16:00Z") is already NOT met. `audits/night-burndown-2026-10-05/r887-lost-jobs.txt`. **NIGHT WATCH 2026-10-06/07 (second burn-down night): 0 jobs lost** — every push checked by its commit (felhom.eu, controller, agent, catalog main and the held branch), each completed `success`, one filtered call per check. | — | Operator: decide whether to raise the act-runner fetch timeout and/or rate-limit the public Gitea pages the crawler walks; meanwhile re-run a lost job via `POST /repos/admin//actions/runs//rerun`. Keep the 2026-10-12 check | operator | -| **R-892** | Process & tooling | P4 | **The update test's box walk cannot reach the Tester 1 box, so decision 149's admin seed there cannot be used.** `app-catalog-felhom.eu/scripts/box_walk.py` drives guests only on demo-hp (its `HP`, `ssh` + `pct exec`) and reaches the app by the guest's LAN address. Read 2026-10-06 (evening): the Tester 1 box (hub host `tester-1-d70be4`) has no SSH alias in DooPlex's `~/.ssh/config` and no entry in `operations/nodes.md`. **Corrected 2026-10-06 18:24 (`09` §3 decision 158):** its Proxmox host IS known — it runs as **VM 341 on the HP box** (`ssh hp`), recorded in `audits/catchup-2026-10-05/tester1/vm341-was-stopped.txt`; the session that filed this did not find that file. Not a 30-minute fix: it needs the box's location and an operator-approved route first. | **OPEN — filed 2026-10-06** **2026-10-06 18:24: operator ruling — yes, CC may reach it by SSH (decision 158).** **2026-10-06 (night): the route is BUILT in the box walk** (catalog `d63ea35`: `TARGET=tester-1`, `ssh -J demo-hp root@192.168.0.154`, guest 9201, `felhom.enkicsifelhom.hu`; `BOX_ADMIN_SEED_GUESTS` has it) and the identity is matched (the agent's report `host.node=felhom` = VM 341's certificate; the guest answers its domain 200, demo-hp's 404). **Blocked:** DooPlex's key is not authorized on VM 341 (`Permission denied (publickey,password)`); the VM has no guest agent and a disk edit needs a VM stop (a reboot, not allowed); fetching its vaulted password from the hub was refused by the session's permission check. `audits/readback-2026-10-07/` **2026-10-06 20:27 (night brief §3): still blocked** — the operator ran `ssh-copy-id` at 20:20, but it copied the key of the operator's own SSH agent; this shell has no agent, and DooPlex's two keys (`id_rsa`, `id_ed25519`) are still refused, direct and through demo-hp (`audits/night-burndown-2026-10-06/s3/`). The fix: run `ssh-copy-id -i ~/.ssh/id_ed25519.pub root@192.168.0.154` from DooPlex as kisfenyo. | DooPlex's key on VM 341 | Operator: authorize DooPlex's public key on VM 341 (one line in `/root/.ssh/authorized_keys`, from its console), or allow CC to read the vaulted password; then prove one step there | operator | +| **R-892** | Process & tooling | P4 | **The update test's box walk cannot reach the Tester 1 box, so decision 149's admin seed there cannot be used.** `app-catalog-felhom.eu/scripts/box_walk.py` drives guests only on demo-hp (its `HP`, `ssh` + `pct exec`) and reaches the app by the guest's LAN address. Read 2026-10-06 (evening): the Tester 1 box (hub host `tester-1-d70be4`) has no SSH alias in DooPlex's `~/.ssh/config` and no entry in `operations/nodes.md`. **Corrected 2026-10-06 18:24 (`09` §3 decision 158):** its Proxmox host IS known — it runs as **VM 341 on the HP box** (`ssh hp`), recorded in `audits/catchup-2026-10-05/tester1/vm341-was-stopped.txt`; the session that filed this did not find that file. Not a 30-minute fix: it needs the box's location and an operator-approved route first. | **OPEN — filed 2026-10-06** **2026-10-06 18:24: operator ruling — yes, CC may reach it by SSH (decision 158).** **2026-10-06 (night): the route is BUILT in the box walk** (catalog `d63ea35`: `TARGET=tester-1`, `ssh -J demo-hp root@192.168.0.154`, guest 9201, `felhom.enkicsifelhom.hu`; `BOX_ADMIN_SEED_GUESTS` has it) and the identity is matched (the agent's report `host.node=felhom` = VM 341's certificate; the guest answers its domain 200, demo-hp's 404). **Blocked:** DooPlex's key is not authorized on VM 341 (`Permission denied (publickey,password)`); the VM has no guest agent and a disk edit needs a VM stop (a reboot, not allowed); fetching its vaulted password from the hub was refused by the session's permission check. `audits/readback-2026-10-07/` **2026-10-06 20:27 (night brief §3): still blocked** — the operator ran `ssh-copy-id` at 20:20, but it copied the key of the operator's own SSH agent; this shell has no agent, and DooPlex's two keys (`id_rsa`, `id_ed25519`) are still refused, direct and through demo-hp (`audits/night-burndown-2026-10-06/s3/`). The fix: run `ssh-copy-id -i ~/.ssh/id_ed25519.pub root@192.168.0.154` from DooPlex as kisfenyo. **2026-10-07 07:10: SSH works now** (the operator added the key; `hostname` → `felhom`; the guest app list read, `audits/night-burndown-2026-10-06/s3/s3-tester1-readonly.txt`). **Still blocked:** the walk signs in to the box's dashboard, and DooPlex holds no dashboard password for this box (the claim of 2026-10-04 kept none; the four old `.ctlpw` files predate the install). Needs: the operator puts the box's dashboard password in a 0600 file for `SC=… TARGET=tester-1`. | DooPlex's key on VM 341 | Operator: authorize DooPlex's public key on VM 341 (one line in `/root/.ssh/authorized_keys`, from its console), or allow CC to read the vaulted password; then prove one step there | operator | | **R-206** | Process & tooling | P4 | **The build-cache cap and the weekly prune exist only as a hand-edited `/etc/docker/daemon.json` on DooPlex — not in Ansible, so a rebuild loses them.** The `node_housekeeping` role must also carry the prune, which today it is forbidden to run | **READY (M) — NEW 2026-08-05** | — | **The spike validated the recipe; this row builds it.** Three parts. **(a) Template `/etc/docker/daemon.json`** with the **`policy` array** form — **the flat form (`{"gc":{"reservedSpace":…}}`) is SILENTLY IGNORED**, measured: the daemon starts, logs nothing, and `docker buildx inspect` still reports the built-in defaults. **The oracle is `docker buildx inspect`, never `dockerd --validate`** — the validator returned `configuration OK` for a bogus key AND for the config that then **crashed the daemon** (`filter` takes one value per policy entry, not an array; `error initializing buildkit: filters expect only one value`). **(b) Narrow the role's Docker ban** (`node-housekeeping.sh.j2:10-14`) to permit exactly `docker builder prune -af` and nothing else — the ban's stated premise ("Docker here runs only unrelated jarr-* dev containers") is obsolete: the growth is Felhom Go build cache. **The measured prune is SYNCHRONOUS** (150.35 GB back at t+0, two consecutive polls <1 MB apart within 60 s) — **unlike containerd's image GC, so it needs no `settle_imagefs` equivalent**, but it MUST measure the filesystem rather than trust the command: `prune` claimed **156.9 GB** and the filesystem returned **150.35 GB**, the 6.5 GB gap being layers still shared with images. **(c) A restart-safety note in the role:** a bad `daemon.json` takes the daemon down AND leaves the `unless-stopped` dev containers stopped — they needed a manual `docker start` — so the role must restart-and-verify, not validate-and-assume. Recipe + every measurement: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` | CC | | **R-209a** | Process & tooling | P4 | **The SSD2 move has NOT survived a reboot, so by this project's own standard it is not fully validated** | **WATCHING — NEW 2026-08-05** | the next DooPlex reboot | **Operator ruled explicitly: do NOT reboot DooPlex.** Uptime verified unbroken (7 weeks 6 days, since 2026-06-10). **The distinction is stated rather than glossed: the MECHANISM is proven** — the guard is wired into both units and containerd refuses to start when a required mount's device is absent — **but the CONSEQUENCE is not**: that a real boot mounts `/mnt/ssd_2` before containerd starts, in this host's actual ordering. Mount-ordering reasoning is precisely the class this project has been burned by (`RequiresMountsFor` RE-MOUNTS rather than refusing — the ep0 lesson), and `CLAUDE.md` prefers a consequence assertion over a mechanism one. **Two deliberate consequences: (1)** the rollback copy `/var/lib/containerd.pre-move-2026-08-05` (**34.3 GB on `/`**) **STAYS** until a reboot validates — which is why `/` sits at 54% and not lower; deleting it now would trade a cheap 34 GB for the only cheap way back. **(2)** validation is **automatic and needs no one to remember it**: `felhom-store-postboot-check.service` (oneshot, enabled, dry-run PASS at install) runs at **every** boot and writes `RESULT: PASS`/`FAIL` to `/var/log/felhom-store-postboot-check.log`, asserting positively that `/mnt/ssd_2` is mounted, that containerd's root is on it, that **`/var/lib/containerd` does NOT exist** (the empty-store trap), that ≥100 images are visible and that both dev containers run. **Next action: after the next reboot — planned or not — read that file; on PASS, `rm -rf /var/lib/containerd.pre-move-2026-08-05` returns ~34 GB to `/`** **P3's prune already removed the urgency: `/` went 86% → 53% used and SSD1's Longhorn disk went `Schedulable=False (DiskPressure)` → `Schedulable=True` (18.85% → 50.32% available).** The move was ruled "cap then move"; the cap is in and the pressure is gone, so this is now a deliberate choice rather than a rescue. **The numbers, measured (`Crucial-SSD-240G`, `/mnt/ssd_2/data/longhorn`, `storageMaximum` 235,148,750,848):** available today **214,958,080,000 (91.41%)**; 25% floor **58,787,187,712**. Moving the whole containerd tree at steady state (~31.5 GB images + ≤30 GB cache ≈ 65 GB) leaves **63.77%, i.e. +38.8 pp above the floor — comfortably safe as measured.** **But `storageScheduled` on SSD2 is 139,586,437,120 while `df` says only 20,094,939,136 is actually used** — Longhorn has overcommitted 6.9× — and if those volumes ever inflate to their scheduled size, the same disk lands at **13.00%, i.e. 12 pp BELOW the floor → `Schedulable=False`**, which is exactly the failure that just took SSD1 out. **Recommendation: do the move only together with setting `storageReserved` on SSD2 to cover the containerd tree (~80 GB); SSD2 reserving zero while HDD2 and HDD4 each reserve 500 GB is an anomaly in its own right.** Mechanism, if it goes ahead: **containerd's `root` in `/etc/containerd/config.toml`** (the key is present but commented out) — **not** Docker's `data-root`, which would move only 0.62 GB. Guard: `RequiresMountsFor=/mnt/ssd_2` on `containerd.service` **and** `docker.service`, remembering that **`RequiresMountsFor` RE-MOUNTS rather than refusing** ([[ep0-datastore-volume-move-2026-07-27]]) — so it must be tested with a genuinely absent device, and the move is not validated until it has survived a **reboot**. Full pre-analysis: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` §P6 | operator + CC | | **R-230** | Process & tooling | P4 | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor | — | — | operator |