Skip to content

Session handover — 2026-07-19

The overnight corpus run was found dead — and the failure was the box, not the corpus. Root-caused to the minute from the journal (networkd declared ens5 Failed under load at 16:32; DHCP lease lapsed ~17:30; box healthy but dark 20 h until IT's console stop/start), mitigated on the box (self-heal watchdog + fstab swap), and run 2 relaunched ~14:45 with the five degradation-window crawls forced. By close, 12 scorecards were banked and they already carry the corpus's headline findings: the first-ever clean blind SHIP, inline-left confirmed as the portfolio norm (×5), and a second mesh-generation site (blind mesh count now 2/9).

Where to look

  • scope-2026-07-18-cms-batch-runner — new §2026-07-19 incident: the measured failure chain, the two box changes now live (/etc/cron.d/networkd-selfheal → logs to /var/log/networkd-selfheal.log; /swapfile2 now in fstab), the force-recrawl recovery discipline (braylaunderette 31.8s clean vs 858s degraded — the poisoning was real), and the open chore: the instance's ID/type/account are recorded nowhere (172.26/16 hostname + IP surviving stop/start suggest Lightsail, unconfirmed — get it from IT).
  • known-issues — corpus status block: run-1 lost (its 34 failure rows are incident artifacts, not corpus data), run-2 partial scores + veto decomposition. Mesh entry: bluestarsearlyyears is mesh site #2, plus the null-axis composite-100 instrument bug (S/G null → composite 100 from T alone; floors blocked the ship; verify batch-scorecard excludes the row before the fit). Header fast-follow #3: inline-left ×5 — now the highest-leverage veto-clearing fix.
  • known-patterns — "A batch whose failure MODE migrates upstream over time is the ENVIRONMENT dying" (the diagnostic chain: --list-boots first, sshd bot-noise as a free cut timestamp, overnight cron durations vs starvation; degradation-window "ok" artifacts are poisoned).
  • Commits: 26910f2 (incident + mitigations + pattern), plus this session-end. Run 2 log: ~/corpus-run-0719.log on EC2 (EXIT: marker).

What shipped

  • Incident diagnosis end-to-end from inside the box (no console access): external check-nodes proved the darkness global; the journal's boot list proved the kernel never crashed; sshd bot-noise silence timestamped the cut; overnight cron at normal speed eliminated starvation; the networkd ens5: Failed line closed it. IT's only needed action was the stop/start.
  • Box hardening: networkd self-heal cron (this failure class is now a ≤5-min blip) and /swapfile2 in fstab (a reboot had silently dropped 6 G swap → 2 G — an unasserted box fact, the baked-seed disease).
  • Run 2 relaunched with resume semantics: 8 banked sites skip, bmac resumes at cms-build, 5 suspects forced crawl,theme, ~28 fresh. Zero failed stages at close.
  • Early score read (12 banked): brianlarkinsolicitor 87.9 SHIP, no vetoes (first blind ship — the calibration fit now has a ship-positive class). S 72–94 / T ~100 on non-mesh sites; bane 91.4 and braylaunderette 93.8 beat both references. The 10 holds decompose into few producers: inline-left ×5, header-bg-vs-light ×7 (measurement only, remedy unscoped), criticalSectionDropped ×3, contentDestroyed ×1, known asset debts, mesh ×2.

The honest parts

  • The run is in flight at close (~5 of the relaunch's headers done, ~28 to go, ETA early hours). Everything below assumes it survives the night — the watchdog covers the known killer, nothing covers an unknown one.
  • The bluestars mis-read was caught one command before the wiki. The fake composite-100 was first attributed to the recorded home/index alias gap — measurement killed it (both sides key index; the real cause is 0 sections). The alias gap remains open but is NOT what bit here.
  • chrome bg ×7 is a measurement, not a scoped fix — live headers paint colour/media/dark bands, ours renders light. Deliberately left as evidence
  • reopen; the remedy (new bounded dim vs existing tone threading) is for a session that measures first.
  • Wiki scale surfaced as a real constraint (known-patterns ~96k tokens now exceeds a single read; growth 788 lines in May vs 3,299 in July). Discussed, not executed: prune known-issues' 11 closed-but-retained entries per its own header rule, plus a generated one-line index (the MEMORY.md pattern) — an embedding index is NOT the next move (chunk retrieval is hostile to this wiki's correction-layered structure); its trigger: a session measurably failing to find an existing entry by grep+index, or the index outgrowing a comfortable load (~500 entries).

Next

  1. Morning: EXIT: in ~/corpus-run-0719.log → if died, relaunch (the resume absorbs it) → node scripts/batch-scorecard.mjs calibration/corpus-hosts.txt — and verify the bluestars partial is loud-skipped, not silently folded into the matrix.
  2. Count the 0-section sites (mesh prevalence — already 2/9 blind; the number decides mesh-segmenter vs manual-tier routing).
  3. Cathal labels ship / needs-polish / hold per site (desktop-only instrument; mobile concerns noted separately).
  4. fit-calibration → review → promote only if fitted:true + precision.
  5. Ranked fixes from the corpus evidence: inline-left render branch (×5, shared-canonical → zero-regression A/B on garvanbay + WCP), the header-bg dimension (×7, measure first), criticalSectionDropped triage (×3), null-axis composite bug.
  6. Record the box's instance ID/type/account in the EC2 runbook (from IT).
  7. Optional slice while the corpus is consumed: the wiki prune + one-line index (session-scoped offer, not committed work).