Home › Guides › Speaker labels in transcription
A transcript without speaker labels is a wall of text where you cannot tell who said what. Speaker labels fix that, turning the raw words into a readable conversation. Here is what they are, how the software figures out who is talking, and the easiest way to get a speaker-labeled transcript.
A speaker label is the tag that marks who said each line, like Speaker 1, Speaker 2, or a person's name. Compare these two transcripts of the same exchange:
so are we agreed on Friday yes Friday works can you send the deck I will send it tonight
...versus the same audio with speaker labels:
Maria: So are we agreed on Friday?
James: Yes, Friday works. Can you send the deck?
Maria: I will send it tonight.
Same words, completely different usefulness. That is what speaker labels buy you.
The technical name is speaker diarization. The software analyzes the audio, detects how many distinct voices are present, and groups each segment by who spoke, producing the generic Speaker 1 / Speaker 2 tags. Those tags stay as Speaker 1, Speaker 2, and so on in the transcript itself — diarization identifies voices, not names. It is a separate step from transcription: transcription turns sound into words, diarization decides who owns each stretch of words.
Most enterprise speech-to-text services support diarization, but they are built for developers. For a normal conversation, the simplest path is an app that does it automatically. With Attesta on iPhone you just record:
Not in the transcript itself. The labels come straight from diarization — a voice, not a name — so the transcript always reads Speaker 1, Speaker 2, and so on. What you can do is use Refine to tell Attesta who's who in plain language. Say "Speaker 1 is Yazan" and that name carries through automatically to the summary, the action items, and the people list, so the parts you actually read afterward are attributed to real people even though the transcript's own labels stay generic.
Diarization is not limited to a two-person back-and-forth. It separates however many distinct voices are actually in the recording, so a three-person interview, a small panel, or a study group comes back labeled the same way — Speaker 1, Speaker 2, Speaker 3, and so on, in the order each voice first speaks.
Speaker 1: So the review turned up three open risks in the draft.
Speaker 2: Can you walk through the second one?
Speaker 3: Sure — it's the vendor dependency I flagged last week.
Speaker 1: Right, let's put that on Friday's agenda.
Accuracy depends more on the audio than the headcount. Two people a few feet apart, taking clear turns, is the easy case. Five people around one phone, with two of them talking over each other, is the hard case — the same things that help with two speakers (clean audio, distinct voices, no talking over each other) matter even more as the group grows.
Get a speaker-labeled transcript automatically, plus a summary and action items, from a single tap.
Download on theApp StoreThe tag that marks who said each line in a transcript, like Speaker 1, Speaker 2, or a name. It turns a wall of text into a readable back-and-forth.
Software uses speaker diarization to tell voices apart and group each segment by who spoke. The labels stay generic — Speaker 1, Speaker 2, and so on — since diarization identifies voices, not names. Clear audio and distinct voices make it far more accurate.
Name or tag, then a colon, then the words, with a new line each time the speaker changes (for example, "Maria: Let's ship on Friday."). Timestamps at each turn are common for interviews and meetings.
Use a tool that supports diarization. Attesta does it automatically on iPhone: record the conversation and you get a speaker-labeled transcript plus a summary, with no setup. Tell it who's who afterward with Refine, and the real name shows up in the summary and action items.
Not in the transcript itself — the labels stay as Speaker 1, Speaker 2, and so on. Use Refine to tell Attesta who's who (for example, "Speaker 1 is Yazan") and the real name carries through automatically to the summary, action items, and people list.
Yes — it separates however many distinct voices are actually in the recording, not just two. A panel, a three-person interview, or a study group all come back labeled Speaker 1, Speaker 2, Speaker 3, and so on. Audio quality matters more as the group grows: distance, background noise, and people talking over each other get harder to untangle the bigger the group gets.