.webm

How to transcribe a WEBM

Upload the WEBM, get the text back in the language that was spoken — and in as many more as you need, with the timings.

WebM comes out of browsers and web recorders — a call recorded in a tab, a video downloaded from the web.

From a WEBM to text, step by step

  1. 1

    Upload the file

    Drag the WEBM in. Nothing to convert first and nothing to install — the format is taken as it is.

  2. 2

    We measure the speech

    Before anything runs we measure how much of it is actually speech, and that is the price you see. A two-hour recording with forty minutes of talking costs forty minutes.

  3. 3

    It is transcribed

    Word by word, with the timings, in the language it was spoken in. Turn on speaker separation and each line says who said it.

  4. 4

    You get it back

    Text, subtitles, or both. If you asked for other languages, they come out of that same listen and not from listening again.

What you get

  • The transcript in the language spoken, with timings.
  • SRT and VTT, ready to drop on a video.
  • Any of the fifteen languages, out of that same audio.
  • Who said what, if you asked for it.

What it costs

€2 per hour of audio

Transcription only. With translation it is €4 per hour of audio, and €2 more for each extra language. Speech, not file length: silence is not charged, so an hour of recording with twenty minutes of voice pays for twenty minutes.

Up to four hours of audio, one hour of video, and 800 MB per recording, three jobs running at a time. Longer than that, split it — or write to us and we will look at the case.

Questions people ask

Do I have to convert the WEBM first?
No. That is the point of this page: the WEBM goes in as it is. Converting it yourself would only lose quality on the way.
What if the WEBM is a video?
Only the audio track is used. The picture is not looked at, not stored and not needed — you can upload the video as it came out of the camera.
How good does the recording have to be?
Good enough that a person listening would understand it. Room noise and a distant microphone cost accuracy in the transcript, and that mistake then travels into every language, which is why the recording matters more than anything else here.