There are two separate language features in Spun and they are easy to confuse. WhatsApp translation happens inside a chat: a message arrives in Portuguese and you read it in English, you type in English and it goes out in Portuguese. A Translation Room is a live call: several people join, each picks the language they want to speak and hear, and everyone hears everyone else in their own.
This page covers both, starting with the one you are most likely to want first.
WhatsApp translation inside a conversation
Every chat has a globe icon in its header, and that is where translation lives. Open it and you get two tabs. The Messages tab has the two settings that matter: translate messages coming in to your language, and translate the messages you send into theirs. Turn both on for one conversation and that conversation becomes bilingual without either side doing anything differently. The customer keeps writing in their language and receiving your replies in it; you never leave English.
The Audio tab does the same job for voice. You can record a voice note in your language and have it sent in theirs, preview the translated audio before it goes, hear their incoming voice notes played back in your language, and have every voice note transcribed to text automatically so you can read instead of listen. You can also tell Spun which language the contact actually speaks, which makes the speech recognition noticeably more accurate.
If you only need one message translated rather than the whole conversation, hovering a message gives you a translate action for that message alone.

Translation settings are per conversation, so a chat that is already in your language stays untouched, and turning translation on for one customer does not turn it on for everyone.
Translation Rooms: a live call in several languages at once
A Translation Room is a video and audio room where translation is the point rather than an add-on. Each participant chooses two things when they join: the language they will speak, and the voice they want to hear other people in. From then on the room does the work. Someone speaks, Spun transcribes what they said, translates it for each listener, and speaks it back in that listener's language with the voice they picked.
Because the translation is spoken rather than only written, a room can hold a real conversation between people who share no common language. Dubbing mode lays the translated audio over the original speech instead of replacing it, which keeps the speaker's tone audible underneath, and the room keeps a running transcript you can read during the call or take away afterwards.
- More than one hundred languages to choose from, per participant rather than per room.
- Push to talk, so a crowded room does not turn into overlapping audio.
- A room code and a guest link, either of which gets someone in.
- A waiting room: guests wait until the host lets them in.
- Host controls: mute one person, mute everyone, silence a participant entirely, lock the room, or remove someone.
- A transcript, plus a video export that composes the participant grid with the subtitles burned in.
Getting an outside participant in
- 1
Create the room
Give it a name, pick your own language and the voice you want to hear.
- 2
Share the link or the code
The guest link works from any browser. The room code is easier to read out over a phone call.
- 3
They pick their language
A guest chooses the language they will speak and the one they want to hear. Nothing is installed.
- 4
Let them in
Guests land in the waiting room. The host admits them, which is what stops a shared link becoming an open door.
- 5
Talk normally
Speak in your language. Everyone else hears it in theirs, within a couple of seconds.
Why a room sometimes changes how it connects
A room with two people in it connects the two browsers directly, which keeps latency low and means the audio and video are encrypted end to end between the participants. As soon as a third person joins, or if someone's connection quality degrades badly, the room switches to a media server that mixes and forwards the streams instead.
That switch is what makes larger rooms and translation-for-everyone possible, and it has a privacy consequence worth stating plainly: in that mode the server necessarily decrypts the audio in order to translate and re-mix it. Everything stays encrypted in transit, but it is no longer end to end. Spun's encryption page lists both room modes explicitly rather than rounding them into one claim.
Costs and limits worth knowing
- Rooms use AI. Every utterance costs speech recognition, translation and speech synthesis, so a long room draws meaningfully on your plan's AI allowance.
- There is a ceiling on how many rooms your workspace can have open at the same time. Closing finished rooms frees the slot.
- A room is built for a working group, not an audience. The participant limit for a room is shown on the room itself.
- Translation is very good, not perfect. Names, slang and heavy background noise are where it slips, and the transcript makes that visible rather than hiding it.
- A recording of a room exists only if you export one. Nothing is captured by default.
Related: What each plan includes