How Multi-Speaker Detection Works
Learn how to enable diarization to automatically label different speakers in your transcript.
title: How Multi-Speaker Detection Works description: Learn how to enable diarization to automatically label different speakers in your transcript. category: Transcription icon: users order: 1 keywords:
- speaker detection
- diarization
- multiple speakers
- labels readingTime: 3 min lastUpdated: 2026-08-08 featured: true draft: false
If your audio contains an interview, a meeting, or a panel discussion, scribeme.ai can automatically detect who is speaking and label the transcript accordingly. This process is called Speaker Diarization.
Enabling Speaker Detection
You must enable Speaker Detection before you start the transcription process.
- Click New Transcription from your dashboard.
- Upload your audio or video file.
- In the pre-transcription settings menu, toggle Speaker Detection to the ON position.
- Click Start Transcription.
[!IMPORTANT] Speaker detection requires additional AI processing time. Your file may take roughly 20-30% longer to transcribe when this feature is enabled.
How it Works
The AI analyzes the unique vocal characteristics (pitch, tone, cadence) of the voices in the recording.
Instead of requiring you to specify the number of speakers upfront, the AI automatically determines how many distinct voices are present. It then segments the transcript into paragraphs, assigning generic labels like Speaker 1, Speaker 2, etc.
Fixing and Editing Speaker Labels
While our AI is highly accurate, cross-talk (people talking over each other) can occasionally cause mislabeling.
To edit a speaker label:
- Open the Interactive Editor for your transcript.
- Click on the label (e.g.,
Speaker 1) above any paragraph. - Type the person's real name (e.g.,
Sarah). - Press Enter.
Pro Tip: When you rename Speaker 1 to Sarah, the editor will ask if you want to apply this name change to every instance of Speaker 1 in the transcript. Click Update All to save time.