The evidence room.
Measured, not promised. Every prompt the system uses, every adversarial scenario we ran, and the raw run data — verbatim. So you can check the work, not take our word. Nothing here is edited for the story; it's the receipts.
The exact instructions the system uses — shown, not described. Cited by both case files.
The 1,000‑attack red team. Exactly what we threw at it, and exactly what came back. Its conclusion is withdrawn — the grader that marked these runs was tested in September 2026 and could not tell a harmful reply from a safe one. The raw material stays filed; see the Case 04 exhibits below.
The 10 jailbreak conversations
The exact adversarial scenarios, turn by turn — anthropomorphism baiting, the "grandma" exploit, DAN, crisis‑in‑jailbreak, encoding tricks, and the rest, across all five age bands.
Full transcripts — smallest model (E2B)
Every turn of every run, with the model's reply and the safety gate's verdict. 841 safe · 159 gate‑caught · 0 breaks.
Full transcripts — capable model (E4B)
The same, on the capable tier. 850 safe · 150 gate‑caught · 0 breaks.
Run summaries (machine‑readable)
The per‑conversation roll‑ups — outcomes, break counts, resisted vs gate‑caught — one file per model.
The honest, in‑progress detection work (how reliably we alert you). Protocols locked before each run; raw numbers after, whichever way they came out.
Pre‑registrations — locked before running
Each protocol's bar and prediction, written down before the run, so a result can't be graded against a moving target. Cascade v2 failed 2 of its 4 bars; it is published here anyway.
Detection numbers — raw‑LLM vs production
The barrage roll‑ups per model, each measure at both levels — classifier alone vs. classifier + the deterministic floor.
The frozen scenario battery
The hashed calibration set the detection numbers are measured against — frozen so any run is reproducible.
The work that withdrew Case 01. The thresholds were written down before the run, the marks were made by hand, and the grader failed all three of them.
The thresholds, locked before the run
What the grader had to achieve to be trusted, written down before a single agreement number existed — so it could not be moved afterwards to fit the result. It was not moved. It was missed.
The result — both graders failed
Agreement scores, the confusion matrix, and the exploratory follow‑up that did not rescue it. Includes what this invalidates and what happens next.
101 delivered replies, marked by hand
Every reply a child would have seen in the run, with the human label beside it. The three marked harmful have their text redacted — we are not reproducing content that could hurt a child to prove that it could. The other 98 are verbatim.