How we measure

Every number we publish carries its qualifier. This page is the qualifier's long form: what our benchmarks are made of, how they're scored, and where their limits sit. We'd rather show you a flag than a wrong note — and we'd rather show you this page than an unqualified percentage.

What we test on

Never real patients. No clinical audio from real consultations is ever used for benchmarking or training — not once, not de-identified, not with consent. Our accuracy work runs on synthetic Australian-accent clinical audio (scripted consultations and medication lists rendered with Australian voices), on acoustically degraded copies of the same material (room reverberation, noise, distance — because clinics are not studios), and on public speech corpora for general-speech robustness.

Medication coverage is measured against a gazetteer built from the medicines Australian clinicians actually prescribe, spanning common day-to-day scripts through to specialty drugs — spoken in sentence context, not read as a list.

How we score

Strictly. A medication counts as recognised only on an exact whole-word match, with Australian spellings and their accepted variants normalised first. We learned early that lenient scoring flatters everything and hides real failures, so partial credit doesn't exist here. Word-error rates are computed conventionally; drug recognition, insertion behaviour, and speaker attribution are scored separately, because a system can be good at one and poor at another.

Through the real pipeline. We measure engines as deployed — the same audio path, the same serving stack, the same correction layers a clinician's consult passes through — because an engine that shines in isolation can behave differently in production, and the production behaviour is the only one that matters. Where a change is evaluated, held-out material the system has never seen is kept apart from anything used in tuning, permanently.

Insertions are counted as failures in their own right. A transcription system that invents a medication that was never spoken is worse than one that misses a word, so we test on medication-free conversational audio specifically to catch invented drugs — and treat any invention as disqualifying for the pathway that produced it.

The limits, plainly

Synthetic audio is not a clinic. Our degraded material approximates real rooms rather than reproducing them; a small number of scripted consultations cannot cover every accent, register, or interruption pattern; and internal benchmarks are exactly that — internal. They are how we compare candidate engines honestly against each other, not a promise of a specific error rate in your room. That's why the product is built to flag uncertainty in the moment rather than lean on an average measured elsewhere, and why every draft requires your review and confirmation.

Questions about any published number — or a case where the product fell short of one — are welcome at contact. Related: privacy by design · responsible use.