Updating to new model and transcription algorithm.
We've updated the model that drafts your notes, along with the algorithm that records and transcribes a consult. This post explains what changed, what to expect, and what stays the same.
What our benchmarking shows
Measured against our Australian corpus, comparing this run with the one we published on 29 June:
Word error rate
1.6% −4.3pt vs 29 JuneInvented details, confirmed
0 −3 vs 29 JuneClinical facts captured
94.7% +3.2pt vs 29 JuneView data and results in more depth.
From exensive testing and Googles guidance, the new Gemini Flash 3.6 model may capture slightly fewer details but hallucinations such as fabricated details are far less likely. We feel in a clinical context completeness should not come at the cost of accuracy.
Templates Remain the Same
Your templates are unaffected. The same templates should produce the same structure, the same sections, and the same fields. Nothing you have set up needs revisiting, and no action is required from you.
What changed
We updated both the note generation model from Gemini Pro 3.1 to Gemini Flash 3.6 and updated our transcription algorithm.
The transcription algorithm. Reworked how speech is detected and how the decoder keeps pace with a long consult, so more of what is said reaches the transcript โ particularly at the start of a sentence and through the end of a longer recording. We also improved name detection, and udpated our benchmarking corpus to include lesser known names.
The note-generation model. Updated to Gemini Flash 3.6. Combined with the improved transcription accuracy this led to fewer hallucianations and tightened language around note generation.
The numbers, and how to check them
Both runs are published in full: the current report, the side-by-side, and past results. You can read every transcript, every generated note, and every grader's reasoning.
Important note: We upgraded the LLM Jury to use larger models for increased confidence in our benchmarking. We found running smaller models multiple times often produced worse benchmarking reports than large models once.
Why we publish this at all
We believe software for healthcare providers should not be treated like SaaS 'move fast and break things'. We are brutally transparent and share any changes we make or plan to make to ensure you are adequately informed about how we process data.
The results from these updates are overall quite positive, however as we wrote at the start of this month, the number only means something if we also publish when they are bad and that transparency is our commitment to you.
If you feel we have regressed, or believe our benchmarking corpus is not reflective of your practice please don't hesitate to reach out at hello@hanah.health.