Helix AI Labs · Research 4 case files · 15 exhibits · 1 withdrawn

The case files.

We test two different things, and we don't mix them up. First, what the child is actually shown. Second, how reliably we alert you, the parent. Both are open work, and one of the files below is a claim we withdrew ourselves after finding that the grader behind it could not tell a harmful reply from a safe one. Measured, not promised, highs and lows alike — and this is what the lows look like.

Case 01 · withdrawn
Case 01 Kids Mode / adversarial Withdrawnthe grader failed

Withdrawn: “nothing harmful reached the child.”

This case concluded that 1,000 hostile turns per model produced nothing harmful. In September 2026 we tested the grader that produced that verdict and found it was marking every reply safe without reading it. The finding stands withdrawn until it is re‑run with a grader that passes validation. The page is kept up, uncorrected below its notice, because deleting it would be the dishonest option.

0.00 grader agreement 3 replies later found harmful re‑run pending
Models
gemma‑4 E2B & E4B · on‑device
Method
10 techniques × 5 age bands × 20 runs
Withdrawn
2026‑09‑06 · see Case 04

Read the withdrawal notice →

Open investigations

The living logs of how we work, mostly the alerting side we're still improving. Published while in flight, wrong turns and all.

Case 02 Kids Mode / alerting Openprotocol locked first

Improving the alerting, how well we tell a parent

The honest, in‑progress work on the other layer: how reliably our on‑device system flags a concern to the parent. A weak model went 20% → 95.7% on the detection task by changing the prompt, wrong turns, a nearly‑published false conclusion, and the honest ceilings included.

Protocol locked
2026‑07‑12 · re‑ask v1
2026‑07‑12 · cascade v2
2026‑07‑14 · cascade v3
Last entry
2026‑07‑14

Read the log →

Case 04 Kids Mode / measurement Openthresholds locked first

We tested our own safety test. It was answering “safe” to everything.

Every safety number we had published rested on an AI grader nobody had ever checked. We marked 101 real delivered replies by hand and asked the graders to mark the same ones. The grader agreed with a human no more than chance would, and missed all three replies that should never have been shown. This is the correction, the three failures in plain language, and what changed because of them.

Thresholds locked
2026‑09‑06 · before the run
Marked by hand
101 delivered replies
Outcome
both graders failed · Case 01 withdrawn

Read the correction →

Case 03 Report card Pending · not yet filed

Report Card 01, Kids‑Safety, small model vs capable model

The parent‑facing verdict: which on‑device model is fit for Kids Mode, in plain language. The cascade experiment has now resolved (see Case 02); the plain‑language conclusion is being written.

Status
In preparation
Why a case gets withdrawn. Case 01 led this file as a settled result. It is now withdrawn, because in September 2026 we tested the grader underneath it and the grader could not tell a harmful reply from a safe one (Case 04). We have left the page up with its notice rather than quietly deleting it. Every protocol above was locked before its run and the raw numbers published after, whichever way they came out — and this is what that commitment costs when it comes out badly. See the evidence room.