Know who said what
The recording is split by voice, so every line of the transcript carries the person who said it — and the subtitles can too.
- Fourteen languages
- You pay for speech, not silence
- No monthly fee
Who this is for
A council session where it matters who moved what. A panel discussion that reads as one long paragraph without it. Interviews where the question and the answer have to be told apart. Any recording with more than one person in it that somebody has to quote from.
How it works
- 1
You upload the recording and tick that you want speakers separated.
- 2
Before transcribing, the whole file is listened to for voices — the whole file, because that is what makes the voice at minute 3 and the voice at minute 40 the same person.
- 3
Each stretch of speech is matched to a voice, and the voices are grouped.
- 4
The transcript comes back with a speaker on every turn, and the subtitles carry it too.
What you get
- Every turn labelled with who spoke it.
- The same labels in the text, the subtitles and the JSON.
- Labels once per turn in subtitles, not on every line.
- How many distinct voices were found.
- It works across the whole recording, not chunk by chunk.
- Priced once, however many languages you asked for.
Being straight about this: It is right about 85% of the time on a decent recording, measured against an annotated corpus. Two people with similar voices get confused, and people talking over each other get confused more. Treat it as a very good first pass that saves the typing, not as a court record.
What it costs
+2 € per hour of speech
On top of the transcription or translation. Charged once and not per language: separating voices happens before any translating, so it costs the same whether the transcript goes out in one language or fourteen.
See pricingQuestions
- Does it know people's names?
- No. It tells voices apart and calls them Speaker 1, Speaker 2 and so on, in the order they first talk. Putting names to them is a find-and-replace you do once.
- How many voices can it handle?
- It is built around small groups — a panel, a meeting, an interview. A hall with thirty people taking turns is not what it is for.
- What if two people talk at once?
- That is where it struggles, and where everybody struggles. Overlapping speech is the hardest part of this problem and we do not pretend otherwise.
- Does it slow the job down?
- It adds a pass over the whole recording before transcription starts. On an eighteen-minute meeting that pass is under a minute.
- Do the subtitles carry it?
- Yes, when there is more than one voice. The label goes on the first subtitle of each turn and not on every line, which would fill the screen with what you already know.
Try it on a meeting
Something with three or four people in it. That is where the difference between a wall of text and a readable transcript actually shows.
Get started