Files
felhom.eu/documentation/audits/update-night-2026-09-21/phase3_b4.py
T
admin da20722e76
gates / gates (push) Successful in 27s
Update night 2026-09-21: Phase 0 and Phase 1 evidence, the drill method, and two instrument fixes
INTERIM CHECKPOINT — evidence off the machine at the end of the phase that produced it (R-320),
not at the end of the session. Phases 2-5 follow in a later commit.

Phase 0, all three mechanisms proven with their controls:
- the fleet floor to 0.261.0 with its declared MinAgent — both demo boxes in 13 s, the hub
  logging `managed floor SERVED ... from declared (golden 0.258.0)`.
- a PRIVATE DRILL CATALOG (admin/app-catalog-drill), so that broken, dummy, cross-repo and
  engine-major edges can be measured without the live catalog ever carrying one. Positive
  control quoted, and two negative controls: the live catalog's main and both real boxes'
  caches unchanged.
- a throwaway image store on the scratch guest, which is what makes an UNATTENDED HOLD
  measurable at all: an edge that PASSES the within-a-major test and still fails.
  CompareImageRefs was proven to order host:port/ references by RUNNING it (4 positive cases
  + 1 negative control), not by reading it.

Phase 1: real within-a-major upstream edges walked on guest 9202 through the product's own
guarded Update, each app seeded and read back through its OWN front door (R-156), with a
per-edge verdict record in 09's shape. `inconclusive` is never collapsed into `failed`.

TWO INSTRUMENT FIXES, both in this repo's own evidence code:
- 00-api-recipe.md said the app page is /app/<n>; it is /apps/<n>, and every call it described
  404s. Corrected, with the session-expiry note that cost the same time.
- unattended-caller.py's follow() read update_phase/updating off the API ENVELOPE, so both were
  always None and EVERY followed update ran to its 900 s timeout and was then recorded
  `timeout` and never-press-again. Fixed before B1 relied on it. R-623.

No controller, agent or hub code was written. The live catalog carries no broken reference.

Gates: repo_gates.py --fast — all 15 OK, exit 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 21:17:46 +02:00

144 lines
5.8 KiB
Python

#!/usr/bin/env python3
"""Phase 3 leg B4 — two Updates within one second, then five.
Is there a single-flight, or do they run together? What does memory do? Do they all end honest?
**Why the images come from the local store.** Each app is put on `drill/<x>:2.0.0` and the edge is
`:2.0.1` — the SAME real image under two tags. The update is real and the pull costs nothing, so
what is measured is CONCURRENCY and not download speed. An edge that could fail on content would
confuse the two.
"""
import json, os, sys, time
import concurrent.futures as cf
HERE = os.path.dirname(os.path.abspath(__file__))
sys.path.insert(0, HERE)
import walk as w # noqa: E402
from phase3_legs import out # noqa: E402
# app -> (subdomain, drill repo, the ref the template carries today)
APPS = {
"privatebin": ("paste", "paste"),
"bentopdf": ("pdf", "pdf"),
"wishlist": ("wishes", "wishes"),
"uptime-kuma": ("status", "status"),
"opengist": ("gist", "gist"),
}
A = "localhost:5000/drill/%s:2.0.0"
B = "localhost:5000/drill/%s:2.0.1"
def current_ref(app):
p = f"{w.DRILL}/templates/{app}/docker-compose.yml"
for line in open(p):
s = line.strip()
if s.startswith("image:"):
return s.split("image:", 1)[1].strip()
return None
def prep(apps):
"""Put each app on the drill A tag and deploy it."""
w.say(f"==== B4 prep: putting {apps} on the local store's A tag")
for app in apps:
repo = APPS[app][1]
cur = current_ref(app)
if cur != A % repo:
if w.drill_bump(app, cur, A % repo) is None:
w.say(f" [prep] could not point {app} at the drill A tag (current={cur})")
return False
w.sync_rescan()
for app in apps:
st = w.stack(app)
if st.get("deployed"):
inst = {k: (v.get("ref") if isinstance(v, dict) else v)
for k, v in ((st.get("app_config") or {}).get("installed_images") or {}).items()}
if list(inst.values()) != [A % APPS[app][1]]:
w.say(f" [prep] {app} is deployed on {inst} — removing so it comes up on the A tag")
w.remove(app)
if not w.deploy(app, APPS[app][0]):
w.say(f" [prep] {app} never came up on the A tag")
return False
w.backup_now(apps[0])
return True
def publish_b(apps):
w.say(f"==== B4: publishing the B tag for {apps} in ONE drill commit")
for app in apps:
repo = APPS[app][1]
w.drill_bump(app, A % repo, B % repo)
w.sync_rescan()
for app in apps:
st = w.stack(app)
w.say(f" [b] {app}: installed={ {k:(v.get('ref') if isinstance(v,dict) else v) for k,v in ((st.get('app_config') or {}).get('installed_images') or {}).items()} } "
f"catalog={st.get('catalog_images')}")
def fire(apps, label):
d = out("B4-concurrent-updates")
w.say(f"==== B4 [{label}]: pressing Update on {apps} at once")
t0 = time.time()
presses = {}
with cf.ThreadPoolExecutor(max_workers=len(apps)) as ex:
fut = {ex.submit(w.ctl, "POST", f"/api/stacks/{a}/update"): a for a in apps}
for f in cf.as_completed(fut):
a = fut[f]
c, b = f.result()
presses[a] = {"http": c, "at_s": round(time.time() - t0, 3),
"reason": (b.get("data") or {}).get("reason") if isinstance(b, dict) else None,
"error": b.get("error") if isinstance(b, dict) else None}
w.say(f" press {a:14} +{presses[a]['at_s']:.3f}s http={c} "
f"reason={presses[a]['reason']!r} :: {(presses[a]['error'] or '')[:90]}")
spread = round(max(p["at_s"] for p in presses.values()), 3)
w.say(f" all {len(apps)} presses issued within {spread}s")
# were they RUNNING together, or one at a time? sample the phases
samples, mem = [], []
t1 = time.time()
while time.time() - t1 < 900:
row = {}
for a in apps:
st = w.stack(a)
row[a] = (st.get("updating"), st.get("update_phase"))
samples.append({"t": round(time.time() - t1, 1), "state": row})
n_updating = sum(1 for v in row.values() if v[0])
if len(samples) % 4 == 1:
w.say(f" +{samples[-1]['t']:>6.1f}s updating={n_updating} " +
" ".join(f"{a}={row[a][1]}" for a in apps))
mem.append(w.guest("free -m | sed -n 2p"))
if all(not v[0] for v in row.values()) and time.time() - t1 > 5:
break
time.sleep(2)
max_concurrent = max(sum(1 for v in s["state"].values() if v[0]) for s in samples)
w.say(f" MOST UPDATES IN FLIGHT AT ONCE: {max_concurrent} of {len(apps)}")
fin = {}
for a in apps:
st = w.stack(a)
fin[a] = {"state": st.get("state"), "update_phase": st.get("update_phase"),
"update_error": st.get("update_error"), "hold_reason": st.get("hold_reason"),
"observables": w.observables(a)}
w.say(f" {a:14} ended phase={fin[a]['update_phase']} err={fin[a]['update_error']!r} "
f"hold={fin[a]['hold_reason']!r} pinned={fin[a]['observables']['pinned_images']}")
rec = {"leg": f"B4-{label}", "apps": apps, "presses": presses,
"press_spread_s": spread, "max_concurrent_updates": max_concurrent,
"samples": samples, "memory_samples": mem, "final": fin,
"elapsed_s": round(time.time() - t1, 1)}
p = f"{d}/{label}.json"
json.dump(rec, open(p, "w"), indent=2, ensure_ascii=False)
open(f"{d}/log.txt", "w").write("\n".join(w.LOG) + "\n")
w.say(f" -> {p}")
if __name__ == "__main__":
w.login()
cmd = sys.argv[1]
apps = sys.argv[2].split(",")
if cmd == "prep":
prep(apps)
elif cmd == "publish":
publish_b(apps)
elif cmd == "fire":
fire(apps, sys.argv[3])