hanah ← Trust Center
Australia · What changed

VALIDATION TRANSPARENCY

What changed.

Our previously published benchmark against the one we publish today, on the same corpus, in this region. Both reports remain available in full — the numbers below are computed from them, not typed in.

The grading panel changed too. Our earlier run was graded by a smaller, cheaper pair of jury models. We found they made avoidable mistakes — miscounting sentences, misreading document structure — so the current run is graded by stronger models. We have deliberately not re-graded the older run to match. Its numbers stand as they were measured on the day. That means part of the coverage movement below reflects better grading rather than a better product, and we would rather say so than quietly restate history.

What we changed

Why we publish this

Model providers change what sits behind an API without announcing it, and a vendor claiming an upgrade "will be better" is not evidence. We run our own benchmark, per region, through the same browser path a clinician uses, and publish what we saw on that date. If a future run regresses, that will be published too — either we can explain it, or we have found something to fix.