Skip to content
CogniYukti

Getting a call in

Nobody has ever read a transcript of speaker one and speaker two.

They open it, scroll, realise they would have to reconstruct who was talking from context, and close it. Which means the transcript your recording tool produces is technically present and practically unread.

The fix is not better audio analysis. It is looking at the meeting.

The difference

Two columns, and only one of them gets opened by a manager.

Separating voices from audio is a solved problem and it produces labels, not names. Turning those labels into people is usually left to the reader, or guessed from a voice profile that has to be trained.

The bot is already in the meeting, so it can simply watch: which participant was speaking, sampled several times a second, kept alongside the recording.

From the audio aloneWith who was on screen
  • Speaker 1Rin Takahashi
    So the migration window is the part I keep coming back to.
  • Speaker 0You
    Let me show you what a staged cutover looks like.
  • Speaker 2Dr Amara Okafor
    We have been burned on that before.
Nobody reads a transcript of speaker one and speaker two. Put the names on it and a manager will actually open it — which is the difference between a feature and a habit.
Illustrative
  • It votes twice, and the second vote has a reason

    Once within each stretch of speech, and again across everything attributed to the same voice. The second pass exists to survive a momentary flicker, or a participant list that lagged and briefly named the wrong person.

  • Every failure falls back silently

    No record of who was speaking, a malformed one, an empty one, or no voice separation at all — each returns to anonymous labels rather than blocking the transcript. A name is an improvement, never a dependency.

  • A call that did not come from the bot keeps its labels

    An uploaded file has no participant list to consult, and the product does not invent one.

Why it matters more than it sounds

Every downstream fact inherits the name.

An objection is only useful if you know who raised it. A commitment matters entirely because of who made it — a promise from the person who signs is a different object from the same words out of somebody who was invited late.

So this is not a readability improvement. It is what makes the evidence attributable, and the difference between "somebody said they were worried about the migration" and knowing it was the person whose team has to do it.

Does it work on a call we uploaded?

No — there is no participant list to read. The names come from the bot having been in the meeting.

What about a dialled call?

Two parties, both known: you, and the person you rang. The problem this solves is a meeting with several people in it.

Can we correct a name it got wrong?

The attribution is conservative by design — where it is unsure it falls back rather than guessing. A wrong name is the failure mode worth avoiding hardest, which is why there are two votes rather than one.