AI health tech should be transparent. Most of it isn't.
AI health tech should be transparent. Most of it isn't.
So here is something most companies in our position would never put in writing: Hanah does not pass our own internal benchmarks with perfect accuracy. In our latest evaluation run, across 106 tests, it made 3 mistakes. It pains me to publish that — and that is exactly why I'm publishing it.
Every instinct the traditional software playbook drills into a founder says to bury that number. And if burying it feels too dishonest, there is always the softer option: massage it. You loosen the way you score. You pick a friendlier set of test cases. You choose the accuracy metric that flatters you, normalise away the errors that count against you, and quietly tidy the reference transcripts until the gap disappears. There are a dozen respectable-looking ways to turn a real 97% into a marketed 100%. We know all of them — which is precisely why we refuse to use any of them.
A number you can't reproduce isn't a benchmark
It's marketing wearing a lab coat. The moment we let ourselves quietly adjust the scoring until the result looked good, the benchmark would stop measuring the product and start measuring our willingness to flatter ourselves. That is worthless to a clinician deciding whether to trust the tool in front of a patient.
So we hold the line the other way. Our word-error rate is scored strictly and transparently, on a versioned corpus with the whole method laid out on our Australian benchmark page — read it and check our working. We don't hand-edit transcripts to lower it, and we don't reach for a lenient normalisation to make a bad week look like a good one. When the number moves, it moves because the product actually changed — for better or for worse — and we can point to why.
Health tech is not productivity software
This distinction matters more here than almost anywhere else, because the price of being wrong is not symmetric.
Hanah is not a project-management app, where a bad autocomplete costs someone a few seconds. It is not even a governance or compliance platform — the Delve-style GRC category whose entire job is to be trusted — where the worst case is a failed audit and an awkward quarter. This is a clinical record. It sits between a clinician and a patient, at the exact point where a quiet error can propagate into a decision about someone's care.
"Fake it till you make it" is a perfectly good motto for a photo-sharing app. In health tech it is negligence dressed up as ambition. The stakes don't forgive the bravado.
What accountability actually looks like
It is not a slogan on a pricing page. For us it means a few concrete, uncomfortable commitments:
- We measure ourselves against a fixed, strict benchmark and publish the full methodology, so the number can be checked rather than taken on faith.
- Every output is reviewed and signed off by a clinician before it lands in a record — the human is the backstop, by design, not an afterthought.
- When we get something wrong, we'd rather tell you where than let you find out yourself. Those 3 mistakes are not a footnote we're hoping you skip. They are the whole point.
We would rather show you an honest 97% than sell you a fabricated 100%.
The uncomfortable truth is that radical transparency is a competitive disadvantage right up until the moment it becomes the only thing anyone trusts. In a field where everyone claims to be accurate and no one shows their working, being the company that shows its working is not a marketing angle. It's the product.
We think the rest of the industry should be measured the same way — openly, reproducibly, and including the times it falls short. We're happy to go first.