← All news
product

Who said it matters as much as what was said

Dr Tyson K··4 min read
Illustration of an older woman and a younger woman sitting side by side in a consulting room, both looking towards a clinician whose back is to the viewer.

Take a single line out of a consult: denies locking.

If the patient said it, it’s history — a symptom asked about and ruled out, and it belongs in the record as such. If you said it, it’s you summarising aloud, or asking, or thinking through the differential in the room. Same four syllables. Two entirely different facts.

A transcript that gets the words right and the speaker wrong is not a small error. It is a note that reads perfectly and says something nobody said.

The failure mode nobody looks for

Clinicians reviewing an AI note tend to read for content. Is the drug right, is the dose right, is the plan right. That is the correct instinct, and it catches a lot. What it doesn’t catch is attribution, because a mislabelled line has no tell. The grammar is fine. The clinical content is plausible — it was, after all, said by someone in the room. It sits in the right section. There is nothing on the page asking you to slow down.

This is the same problem we wrote about with uncertainty, pointed at a different target. A confident error is the one you sign without noticing, and a confident speaker error is the most invisible kind of all.

Binary, on purpose

Most transcription systems try to answer “how many people are speaking, and which is which?” — and hand you Speaker 1, Speaker 2, Speaker 3, leaving you to work out who’s who. In a consult room that’s a guess dressed as data. Rooms have interpreters, parents, students, a spouse who answers on the patient’s behalf, someone in the corridor.

Akoua asks a narrower question, because it’s the one that actually determines how a line should be recorded: is this the clinician, or is it not?

That’s binary attribution. Your voice on one side, everyone else’s on the other. It’s a smaller claim, and it’s a claim the system can stand behind. The note doesn’t need to know whether the second voice was the patient or their daughter to know that it wasn’t you — and that distinction is what decides whether a line is history or your own reasoning.

When it can’t tell, it says so

Sometimes attribution genuinely isn’t clear. Two people talk over each other. Someone is far from the microphone. A voice is unfamiliar and close in register to yours.

Akoua’s answer to all of these is the same as everywhere else in the product: say so rather than guess. The line stays in the note as a highlighted block carrying attribution unclear · review, and it waits for you. You read it, you know instantly who said it, you settle it in a second or two.

That is a deliberate trade. A system that always assigns a speaker looks more finished and is more dangerous. We would rather hand you a short list of lines that need a human than a clean note with a silent mistake in it. In every benchmark condition, Akoua produced zero confident speaker mislabels — when it can’t tell who spoke, it flags instead of committing.

Internal benchmarks on synthetic Australian-accent clinical audio and public speech corpora. We publish our methodology — and we’d rather show you a flag than a wrong note.

The voiceprint stays put

Telling your voice from everyone else’s means Akoua holds a voiceprint for you. That is biometric data about a person, so it’s treated like it.

Your voiceprint is encrypted at rest with AES-256, alongside transcripts, notes and other personal information. No stored voiceprint ever leaves Akoua. It isn’t shared, it isn’t sold, and it is never used to train models — the same rule that covers audio, transcripts and notes.

Only yours is enrolled. Akoua does not build voiceprints for patients; it doesn’t need to. Recognising one voice is enough to answer the only attribution question the note depends on, and it means the people who never consented to anything beyond the consultation itself have nothing enrolled at all.

Why this is a documentation problem, not a clever one

There is a version of speaker attribution that chases accuracy for its own sake — more speakers, finer labels, better numbers. That isn’t the goal here.

The goal is a note whose every line can be traced to a person who actually said it. Attribution is upstream of everything else Akoua does: the structure of the note, what gets recorded as history, what’s marked [not stated]. Get it wrong and the honest parts downstream are honest about the wrong thing.

So Akoua keeps the claim small, marks the doubt, protects the biometric, and leaves the judgement with you. It’s a documentation aid, not a diagnostic device — it makes no clinical inferences and writes down only what was said. It just insists on being right about who said it, or admitting that it isn’t sure.

productsafetyai medical scribe