2026-09-16T22:10:43Z ================ ROUND 7 : use nextcloud, while: internet-gone-10min ================
2026-09-16T22:10:46Z --- BEFORE --- containers=26  cloud=200  status=200  paste=200
2026-09-16T22:10:46Z     (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
2026-09-16T22:10:46Z --- ACTION: use on nextcloud ---
2026-09-16T22:10:46Z     cloud read 1 -> 200
2026-09-16T22:10:47Z     cloud read 2 -> 200
2026-09-16T22:10:47Z     cloud read 3 -> 200
2026-09-16T22:10:47Z --- ACCIDENT: internet-gone-10min (injected after the action started) ---

## MID-CUT measurements (22:11:15Z) — the block does exactly what it claims
Asked of the box while its network was cut:
    from the box:  internet **blocked** (ping 1.1.1.1 fails) · LAN **reachable** (ping 192.168.0.180 ok)
    containers     **26** — every app still running; they do not care that the internet is gone
    free `/`       **7573M**, unchanged
    diskguard      active, **0 log lines** — it has not needed to fire
    from DooPlex:  **LAN 301** (traefik answers on the box) · **public 502**

**The two paths separate cleanly, and this is the first round where that matters.** The apps are all
up and serving locally; what is unreachable is the way in from outside. Note the public code differs
from round 4's: **530** when cloudflared was dead (Cloudflare had no tunnel at all), **502** now
(the tunnel process is alive but cannot reach anything). Two different failures of the same journey,
and the box reports them differently without being asked to.

**The injector behaved as designed** — it blocks off-LAN traffic on this VM's tap only and leaves the
LAN alone, which is what made these measurements possible at all: I can still reach the box to ask
it questions while it cannot reach the world.

## „vzdump procs: 2" was MY OWN COMMAND — the twelfth self-inflicted reading tonight
Asking for the actual command lines instead of a count showed exactly one match, and it was mine:

    120847  00:00  bash -c echo START echo "--- actual vzdump command lines ---" ps -eo pid,etime,args
                   | grep "[v]zdump" … echo "vzdump procs: $(…)" …

The `[v]zdump` bracket trick stops the PATTERN matching itself, but my own **label** — the literal
text `"vzdump procs: "` that I echoed — contains the word, so `ps` listed my shell and the grep
counted it. Every „vzdump procs: 2" reading tonight was **noise from my own command line**, including
the one I cited when deciding the local tier might retry mid-round.

**What it changes:** no second backup was ever running. The box was idle after the off-site leg
finished at 22:08:27Z, and it is idle now, mid-cut: no `.tar.dat` being written, free `/` unchanged at
7573M, `diskguard` active with **0 log lines**.

**What it does not change:** the disk guard stays. The local tier really did announce
„retrying with backoff (next attempt in 15m0s)", and that retry is still due; the guard is cheap
insurance either way and has cost nothing.

Same family as the `pkill -f` that killed my own watcher earlier: **a pattern that matches the hand
holding it.** The cure is the one that worked here — ask for the command lines, not the count.
START
--- controller: any event pushes attempted, and did they fail? ---
2026/09/16 22:08:42 [INFO] [scheduler] Running job: offsite-credential-retry
2026/09/16 22:08:42 [INFO] [scheduler] Job offsite-credential-retry completed (took 0s)
2026/09/16 22:08:42 [INFO] [scheduler] Running job: hub-report
2026/09/16 22:08:42 [INFO] [report] Building system report
2026/09/16 22:08:43 [INFO] [report] Hub report pushed successfully (15934 bytes)
2026/09/16 22:08:43 [INFO] [scheduler] Job hub-report completed (took 1.122s)
2026/09/16 22:13:42 [INFO] [scheduler] Running job: offsite-credential-retry
2026/09/16 22:13:42 [INFO] [scheduler] Job offsite-credential-retry completed (took 0s)
--- is the box still blocked right now? ---
  internet: still blocked
--- containers / free / guard ---
containers=26  free=7573M  guardlines=0
END

## INTERIM: the cut may be shorter than the box's own cadence — so nothing even tried to send
Read from the box at 22:13:49Z, while it was still blocked:

    22:08:42Z  [scheduler] Running job: hub-report
    22:08:43Z  [report] **Hub report pushed successfully (15934 bytes)**   <- BEFORE the block
    22:13:42Z  [scheduler] Running job: offsite-credential-retry — completed (took 0s), no push
    (no hub-report job since 22:08:42Z; no push failures; no dropped events logged)
    containers **26** · free `/` **7573M** unchanged · diskguard **0 lines**

The block began about 22:10Z. The box's last contact with the hub was a minute and a half BEFORE
that, and in the ten minutes since, its scheduler has not tried to reach the hub at all.

**So round 7 may not test what I expected it to test.** The interesting question — are alarms raised
while the hub is unreachable retried and then silently dropped? — needs the box to actually attempt a
push during the outage. If its reporting interval is longer than the outage, a ten-minute cut can
pass entirely between two reports and the hub never notices anything happened.

That would itself be a finding worth having, and a reassuring one: **a short internet outage is
invisible to the fleet view not because anything is hidden, but because nothing was due to be said.**
The interval is measured from the box's own log rather than recalled from documentation, and the
post-block check confirms whether any push was attempted, failed, or lost.

## The cadence, MEASURED — and what it means for this round and the next two
From the box's own log, not from documentation:

    21:53:41Z  [scheduler] **Registered periodic job: hub-report (every 15m0s)**
    21:53:49Z  [report] Hub report pushed successfully (30535 bytes)
    22:08:43Z  [report] Hub report pushed successfully (15934 bytes)      <- exactly 15 min later

The block runs ~22:10Z → ~22:20Z. **The next report is due ~22:23:43Z, after the cut ends.**

**Conclusion for round 7, stated plainly: this round does not test what I hoped it would.** The
question „are alarms raised while the hub is unreachable retried and then silently dropped?" needs
the box to attempt a push during the outage. Here the outage falls entirely between two reports, so
nothing was attempted, nothing failed, and nothing was lost. The correct verdict is **not exercised**
— not „passed".

**And it is a real, if quiet, finding of its own:** on this box a ten-minute internet outage is
invisible to the fleet view, because the box had nothing due to say. The hub's staleness thresholds
(30 minutes to `node_stale`) are set well beyond that, so the two mechanisms agree.

**What it implies for rounds 8 and 9, written before they run:** both are also ten-minute cuts, drawn
at ~25-minute spacing against a 15-minute report cycle, so one of them will almost certainly contain a
scheduled report and will exercise the drop behaviour properly. **The rounds are NOT re-timed to make
that happen** — re-timing a round to obtain a better result is choosing the night after the fact. If
it happens naturally, it is measured; if it does not, that is recorded too.
    2026-09-16T22:10:47Z ACCIDENT=internet-gone-10min round=7
    2026-09-16T22:10:48Z blocking the box's traffic off-LAN at the HOST, on tap336i0; the LAN stays up
    2026-09-16T22:10:48Z blocked (LAN allowed, everything else dropped) — 10 minutes
    2026-09-16T22:20:48Z unblocked; host sysctl restored to 0 and both rules removed
    -P FORWARD ACCEPT
    2026-09-16T22:20:48Z accident internet-gone-10min complete
2026-09-16T22:20:48Z --- AFTER: what the box did BY ITSELF ---
2026-09-16T22:20:50Z     t+603s containers=26 (before 26)
2026-09-16T22:20:50Z     STEADY after 603s
2026-09-16T22:20:51Z     front doors: cloud=530  status=530  paste=530  wiki=530
2026-09-16T22:20:51Z     household lines this round: 10  failures: 0
2026-09-16T22:20:51Z --- alarms ---
  | Time | Severity | Type | Message | Source
  | Sep 16 21:59 | error | whole_guest_backup_failed | Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s) | controller
  | Sep 16 21:53 | info | controller_started | Controller elindult (0.245.0) | controller
  | Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller
  | Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller
  | Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller
  | Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller
  | Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller
  | Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller
2026-09-16T22:20:51Z ================ END ROUND 7 ================

## The block, and the fence check that proves the host was left as it was found
    22:10:48Z  blocked (LAN allowed, everything else dropped) — ten minutes
    22:20:48Z  unblocked; host sysctl restored to 0 and both rules removed
    22:20:48Z  accident internet-gone-10min complete

**Verified independently on demo-hp at 22:21:08Z, against the baseline taken BEFORE the round:**
    iptables -S FORWARD   ->  `-P FORWARD ACCEPT`   (baseline: identical)
    physdev rules         ->  **0**                 (baseline: 0)
    net.bridge.bridge-nf-call-iptables = **0**      (baseline: 0)

Byte for byte the state I recorded before touching anything. That matters more than tidiness: demo-hp
is a Tier-0 host that also carries guests **9201 and 9202**, so an abandoned FORWARD rule would be a
fence breach rather than an untidy drill. The unconditional cleanup trap added before this round was
not needed in the end — the happy path ran — but it was the right insurance for a ten-minute sleep on
a shared host.

## A reading that measures nothing, and is labelled as such
The runner's own front-door sample at **22:20:51Z** shows `cloud/status/paste/wiki = 530` — taken
**three seconds** after the block lifted, before Cloudflare could re-establish the tunnel. It is not a
recovery measurement and is not treated as one; the recovery time comes from a later reading.

## ROUND 7 — the five things
**action:** `use` nextcloud (drawn as `update`; ran as `use` because the catalog bump was never
pushed — the catalog gates returned INCONCLUSIVE and undetermined is never a pass)
**accident:** the box's traffic blocked off-LAN for ten minutes (22:10:48Z → 22:20:48Z)

1. **What the customer saw.** Depends entirely on where they were standing. **At home: nothing** —
   traefik answered `301` on the box throughout and every app kept serving. **Away from home: ten
   minutes of nothing** — the public route returned **502** while the tunnel was alive but could reach
   nothing, then **200** again about a minute after the block lifted.
2. **What the box did by itself.** Kept all **26** containers running and needed no repair. When the
   block lifted it re-established the way in **unaided**: 22:20:51Z still 530 (three seconds after
   unblocking), **22:21:52Z all four apps 200** — so **~64 seconds** from network restored to front
   doors serving.
3. **Time to steady.** The apps never left steady (26 before, 26 after). The only thing that broke and
   healed was the path in: **~64 s**.
4. **Alarm fired / true?** **None fired, and none should have.** The ladder puts `node_stale` at 30
   minutes; a ten-minute outage is far inside that. Nothing false was raised either.
5. **Should have fired and did not.** **None.**

**Household loop: 10 operations, 0 failures — and again narrower than it looks.** It does not follow
redirects, so it measured the box (up throughout) and was blind to the public outage. Its zero is not
evidence the household was unaffected.

## The question this round was supposed to answer, and honestly did not
Are alarms raised while the hub is unreachable retried and then **silently dropped**? **Not
exercised.** The box reports every **15m0s** (measured: pushes at 21:53:49Z and 22:08:43Z), and the
block fell entirely between two reports — the next was due ~22:23:43Z, after it lifted. Nothing was
attempted, so nothing could be lost. Recorded as *not exercised*, never as *passed*.

**Rounds 8 and 9 are also ten-minute cuts** and, at ~25-minute spacing against a 15-minute cycle, one
of them should contain a scheduled report naturally. They are not re-timed to force it.

## The fifth mistimed reading — caught by the measurement labelling itself
I checked for the post-block hub report at **22:23:28Z**. It was due at **22:23:43Z**. Fifteen seconds
early, and the newest entry was still the pre-block push at 22:08:43Z.

What stopped it becoming a false conclusion („the box never resumed reporting") is that the command
**printed its own timestamp beside the answer**, with the due time written into the label:
    reading taken at 2026-09-16T22:23:28Z   (too early if before 22:23:43Z)
    guest clock: 2026-09-16T22:23:29Z

That is the fix that has now worked twice, after three earlier clock errors where I estimated instead
of reading: **make the measurement state its own precondition, so a premature reading announces itself
rather than being mistaken for a result.** The same principle as the marker blocks that caught the
password-file check and the command lines that caught the `vzdump` self-match — in each case the cure
was making the instrument say what it had actually done.

## ROUND 7 CLOSED — the box resumed reporting exactly on time, and nothing was lost
    22:23:42Z  [scheduler] Running job: hub-report
    22:23:43Z  [report] **Hub report pushed successfully (10538 bytes)**
    (reading taken at 22:24:06Z, guest clock 22:24:08Z — after the due time, and labelled as such)

The cadence held to the second: pushes at **21:53:49Z · 22:08:43Z · 22:23:43Z**, fifteen minutes apart
each time, straight through a ten-minute network cut that sat between the second and third.

**So the complete answer for round 7 is:** the box lost its way out for ten minutes, kept every app
serving at home, restored the public path unaided in ~64 s, raised no alarm (correctly — the staleness
threshold is 30 minutes), **attempted no hub contact during the outage because none was due**, and
then reported on schedule as if nothing had happened. Nothing was dropped, because nothing was sent.

That is a good result and a narrow one, and the narrowness is the point: this round did **not** test
what happens to an alarm raised *while* the hub is unreachable. Rounds 8 and 9 are also ten-minute
cuts, and at ~25-minute spacing against a 15-minute cadence one of them should contain a scheduled
report. If it does, the drop behaviour finally gets measured; if it does not, that is recorded too.
