Simultaneous translation, speech to speech, in real time
AI simultaneous translation: somebody speaks and the audience hears it in their own language a few seconds later, as speech and not only as subtitles. Simultaneous interpretation with no booth, no headsets to hand out and nothing to install.
- Fourteen languages
- You pay for speech, not silence
- No monthly fee
Who this is for
A parish whose Sunday congregation no longer shares one language. A conference that cannot afford a booth and two interpreters per language. A council meeting that has to be followed by people who moved here last year. Anywhere somebody talks to a room and part of the room is missing it.
How it works
- 1
The sound of the room reaches us. A small device on the lectern sends it over SRT, or whatever you already broadcast with sends it over RTMP.
- 2
The engine listens, works out what is being said and translates it while the sentence is still in the air.
- 3
A synthetic voice speaks it in each language, and the audience hears it through a link they opened on their own phone.
- 4
You watch it from your panel: which rooms are live, how long they have run and what that has cost.
What you get
- Fourteen languages at the same time, out of the same audio.
- Under a second from the sentence ending to the voice starting.
- Nothing to install for whoever is listening: a link, or a QR on the wall.
- Your own rooms, created and paused from a panel or from the API.
- Nothing recorded. The audio is translated as it passes and is not kept.
- Credentials per room, so a device that is lost does not open the rest.
Worth knowing: The room needs a usable microphone. Translation can only be as good as what it hears, and a phone at the back of a nave picks up the nave rather than the speaker — the single biggest difference in quality comes from the microphone, not from us.
What it costs
From 0.135 € a minute
You pay for speech, not for the length of the event: an hour with forty minutes of talking spends forty minutes. And each language after the first costs 0.35 of a minute rather than another whole one, because the audio is heard once and every language comes out of that one listen.
See the packsQuestions
- Does the audience need an app?
- No. They open a link in the browser they already have and choose their language. A QR at the entrance is usually enough.
- How long is the delay?
- Under a second from the moment a sentence ends. It is not word by word: the engine waits for the sentence to finish, because translating half a sentence produces the wrong half.
- Is anything recorded?
- Not from a live room. The audio is translated as it goes past and nothing is kept. If you want a recording transcribed, that is a different service and you send us the file.
- What do I need in the room?
- A microphone that hears the speaker and a way to send us the sound: the small device that sends over SRT, or whatever you already broadcast with over RTMP.
- What if the internet drops?
- The room reconnects on its own and carries on. What was said during the gap is lost — there is no way around that — but nobody has to restart anything.
Hear it before you believe it
Fourteenty seconds with your own voice, in a new window. Pick two languages, put headphones on, and talk.
Try real-time speech-to-speech translation →No account, nothing to install. Headphones required.
Where people use it
The same engine in very different rooms. Each of these says how it is set up there, where the audio comes from, and where its own limit is.
And what it is usually compared with
Three honest comparisons, each of them saying what the other option does better.
Try it on something real
The first thirty minutes are on us. That is enough to sit in your own room, with your own microphone, and hear whether it works for you.
Get started