Session handover — 2026-07-19¶
The overnight corpus run was found dead — and the failure was the box, not
the corpus. Root-caused to the minute from the journal (networkd declared
ens5 Failed under load at 16:32; DHCP lease lapsed ~17:30; box healthy but
dark 20 h until IT's console stop/start), mitigated on the box (self-heal
watchdog + fstab swap), and run 2 relaunched ~14:45 with the five
degradation-window crawls forced. By close, 12 scorecards were banked and they
already carry the corpus's headline findings: the first-ever clean blind
SHIP, inline-left confirmed as the portfolio norm (×5), and a second
mesh-generation site (blind mesh count now 2/9).
Where to look¶
- scope-2026-07-18-cms-batch-runner
— new §2026-07-19 incident: the measured failure chain, the two box
changes now live (
/etc/cron.d/networkd-selfheal→ logs to/var/log/networkd-selfheal.log;/swapfile2now in fstab), the force-recrawl recovery discipline (braylaunderette 31.8s clean vs 858s degraded — the poisoning was real), and the open chore: the instance's ID/type/account are recorded nowhere (172.26/16 hostname + IP surviving stop/start suggest Lightsail, unconfirmed — get it from IT). - known-issues — corpus status block: run-1 lost
(its 34 failure rows are incident artifacts, not corpus data), run-2
partial scores + veto decomposition. Mesh entry: bluestarsearlyyears is
mesh site #2, plus the null-axis composite-100 instrument bug (S/G
null → composite 100 from T alone; floors blocked the ship; verify
batch-scorecardexcludes the row before the fit). Header fast-follow #3: inline-left ×5 — now the highest-leverage veto-clearing fix. - known-patterns — "A batch whose failure MODE
migrates upstream over time is the ENVIRONMENT dying" (the diagnostic
chain:
--list-bootsfirst, sshd bot-noise as a free cut timestamp, overnight cron durations vs starvation; degradation-window "ok" artifacts are poisoned). - Commits:
26910f2(incident + mitigations + pattern), plus this session-end. Run 2 log:~/corpus-run-0719.logon EC2 (EXIT:marker).
What shipped¶
- Incident diagnosis end-to-end from inside the box (no console access):
external check-nodes proved the darkness global; the journal's boot list
proved the kernel never crashed; sshd bot-noise silence timestamped the
cut; overnight cron at normal speed eliminated starvation; the networkd
ens5: Failedline closed it. IT's only needed action was the stop/start. - Box hardening: networkd self-heal cron (this failure class is now a
≤5-min blip) and
/swapfile2in fstab (a reboot had silently dropped 6 G swap → 2 G — an unasserted box fact, the baked-seed disease). - Run 2 relaunched with resume semantics: 8 banked sites skip, bmac
resumes at cms-build, 5 suspects forced
crawl,theme, ~28 fresh. Zero failed stages at close. - Early score read (12 banked): brianlarkinsolicitor 87.9 SHIP, no vetoes (first blind ship — the calibration fit now has a ship-positive class). S 72–94 / T ~100 on non-mesh sites; bane 91.4 and braylaunderette 93.8 beat both references. The 10 holds decompose into few producers: inline-left ×5, header-bg-vs-light ×7 (measurement only, remedy unscoped), criticalSectionDropped ×3, contentDestroyed ×1, known asset debts, mesh ×2.
The honest parts¶
- The run is in flight at close (~5 of the relaunch's headers done, ~28 to go, ETA early hours). Everything below assumes it survives the night — the watchdog covers the known killer, nothing covers an unknown one.
- The bluestars mis-read was caught one command before the wiki. The
fake composite-100 was first attributed to the recorded home/index alias
gap — measurement killed it (both sides key
index; the real cause is 0 sections). The alias gap remains open but is NOT what bit here. - chrome bg ×7 is a measurement, not a scoped fix — live headers paint colour/media/dark bands, ours renders light. Deliberately left as evidence
- reopen; the remedy (new bounded dim vs existing tone threading) is for a session that measures first.
- Wiki scale surfaced as a real constraint (known-patterns ~96k tokens now exceeds a single read; growth 788 lines in May vs 3,299 in July). Discussed, not executed: prune known-issues' 11 closed-but-retained entries per its own header rule, plus a generated one-line index (the MEMORY.md pattern) — an embedding index is NOT the next move (chunk retrieval is hostile to this wiki's correction-layered structure); its trigger: a session measurably failing to find an existing entry by grep+index, or the index outgrowing a comfortable load (~500 entries).
Next¶
- Morning:
EXIT:in~/corpus-run-0719.log→ if died, relaunch (the resume absorbs it) →node scripts/batch-scorecard.mjs calibration/corpus-hosts.txt— and verify the bluestars partial is loud-skipped, not silently folded into the matrix. - Count the 0-section sites (mesh prevalence — already 2/9 blind; the number decides mesh-segmenter vs manual-tier routing).
- Cathal labels ship / needs-polish / hold per site (desktop-only instrument; mobile concerns noted separately).
fit-calibration→ review → promote only iffitted:true+ precision.- Ranked fixes from the corpus evidence: inline-left render branch (×5, shared-canonical → zero-regression A/B on garvanbay + WCP), the header-bg dimension (×7, measure first), criticalSectionDropped triage (×3), null-axis composite bug.
- Record the box's instance ID/type/account in the EC2 runbook (from IT).
- Optional slice while the corpus is consumed: the wiki prune + one-line index (session-scoped offer, not committed work).