Why whispered interpreting is not "what we do every day", and why we have accepted that for far too long
There is a particular moment every interpreter who has done whispered work will recognize.
You are half a meter behind your delegate. The speaker is at the front, twelve meters away, amplified through a ceiling system that was designed for a wedding reception. Somewhere to your left, a coffee machine. A few meters onyour right, a second interpreter working into another language. Your delegate leans slightly away from you, because nobody enjoys having a voice in their ear, and you lean slightly forward, because you cannot hear the speaker.
And then you begin to talk. And the moment you begin to talk, you can hear the speaker even less.
So you lean closer. You raise your voice a little, not much, just enough to be sure your delegate can follow. And the floor recedes a little further.
You do this perhaps with a partner for ninety minutes. Afterwards you are more tired than a full day in the booth, and you cannot entirely explain why. The content was not difficult. The speaker was not fast. You did the job well enough that nobody complained.
I want to explain why.
The objection I always get first
Whenever I raise this, the answer comes back immediately, and it is always some version of the same thing:
"We speak and listen at the same time all day. In restaurants. In meetings. At home with the television on. The human ear manages perfectly well."
It is a reasonable objection. It is also wrong, and it is wrong in a way that turns out to be measurable.
We do not, in fact, speak and listen at the same time in ordinary life. We take turns. Human conversation is one of the most precisely choreographed turn-taking systems in nature — we hand the floor back and forth with gaps measured in fractions of a second, and genuine overlap is brief, accidental, and usually repaired within a syllable or two. When two people talk over each other for more than a moment, we experience it as a breakdown and one of them stops.
Simultaneous interpreting is the only common human activity in which one person must fully comprehend incoming speech while continuously producing different speech, for twenty or thirty minutes without pause, with no possibility of repair.
That is not an intensified version of everyday listening. It is a different task. And the ear treats it differently, which is where this becomes interesting.

The reflex nobody told us about
Here is something that is not taught in most interpreting programs, and I think it should be.
Inside your middle ear there are two very small muscles. The better studied of the two is the stapedius, and its job is to stiffen the chain of tiny bones that carries sound from your eardrum to your inner ear. When it contracts, less sound energy gets through, and the attenuation can exceed twenty decibels.
Most people know it as a protective response to sudden loud noise. That is real, and it is the mechanism behind much of the discussion of acoustic shock in our profession.
But there is a second trigger, and this is the one that matters for us.
The stapedius contracts when you speak.
Not in response to your voice — before it. Research on themuscle's electrical activity during vocalization found that the signal often begins before any sound is produced, which means the contraction is issued centrally as part of the act of speaking. Your brain tells your ear to turn itself down at the same moment it tells your larynx to switch on.
The threshold sits remarkably low. The muscle begins responding at closeto the quietest sound a person can produce. Whispering is enough.
And at ordinary conversational effort, the muscle is already working at roughly half its maximum.
Why would evolution build such a thing? Because your own voice is, acoustically speaking, an enormous event happening a few centimeters from your inner ear. The reflex mainly attenuates low frequencies, which is where the bulk of your own voice's energy sits. The generally accepted interpretation is that it exists precisely to stop your own speech from drowning out the world, to keep you aware of your surroundings, and to let you hear other people while you are talking.
In other words: evolution already solved the problem of speaking and listening at once.
Which is exactly why the objection feels so convincing.
Why the solution stops working in our job
Read that mechanism again and notice what it does not do.
It does not distinguish between your voice and the speaker's.
The stapedius does not know which sound you want. It stiffens the whole conduction path. When it fires, everything arriving through your eardrum isreduced, your own voice, yes, and the speaker you are being paid to understand.
For everyone else on earth, that trade is free. When you are speaking in a conversation, you are not simultaneously trying to decode critical content from someone else. A slightly attenuated world for a few seconds costs you nothing, and you get your turn to listen properly a moment later.
For us there is no moment later. The attenuation lands squarely on the signal we are working from, continuously, for the entire turn.
The ear's own protective mechanism is free in conversation. Insimultaneous interpreting, it is a tax.

The loop that makes it worse
Now put that together with what we instinctively do when we cannot hear.
We speak up.
It is involuntary. Everyone raises their voice in a noisy environment without deciding to. But look at what raising your voice does in this specific situation.
More vocal effort means more contraction of the muscle, which means more attenuation of the floor. More vocal effort also means more of your own voice reaching your inner ear directly through the bones of your skull, a path the middle ear muscles cannot do much about, because it largely bypasses them. You cannot turn your own voice down. It is inside the building.
So: you cannot hear the speaker, so you speak louder, so you hear the speaker even less, so you strain harder, so you speak louder still.
That loop runs for the length of the assignment. Nothing in it is a mistake on the interpreter's part. It is the correct instinctive response to a noisy room, applied to the one task where it is exactly wrong.

What is actually missing
Now compare a booth.
Floor audio arrives on a dedicated channel. There is a limiter in the chain. There is a volume control, and there is a tone control, and if the floor is hard to follow you turn it up. Your ear is still doing everything described above — the reflex fires in the booth too — but you have a lever to compensate with. When the signal weakens, you strengthen it.
Now compare whispered work. Floor audio arrives through the air, at whatever level the room happens to deliver, filtered by distance, competing with air conditioning, chairs, side conversations and your own voice at close range.
There is no lever. None. Not a quiet one, not an inconvenient one, the control does not exist.
Look around a well-run conference and count the gain stages. The speaker's microphone has one. The mixing desk has several. The amplifier hasone. The delegate's receiver has one. Every single element in that chain can be adjusted to taste except one.
The interpreter, the person whose comprehension the entire event depends on, is the only element with no adjustment at all.
And I want to be clear that this is not an argument against whispered interpreting. Whispering and bidule work exist for excellent reasons. A boothis heavy, expensive, and logistically demanding. It needs a room that will take it, a technician, a delivery van, and a budget. Whispered interpreting is mobile, immediate and affordable. It goes on a factory tour, a site visit, a hospital corridor, a boardroom, a moving bus. It is often the only reason interpretation happens at all.
The problem is not that we work this way. The problem is that we have accepted, for the entire history of the profession, that working this way means working with our hands tied, and we have accepted it so completely that we no longer notice.

What this actually costs
I want to be careful here, because this is a subject where overstatement does real damage.
I am not claiming that whispered interpreting damages your hearing. That is not what the evidence shows, and I am not going to pretend otherwise. The published work on hearing harm in our profession concerns the headphone sound chain: platforms, consoles, compressed audio, sudden surges. AIIC's acoustic shocks study, conducted with an audiologist at Aix-Marseille and drawing on more than a thousand members, found acoustic incidents affecting between roughly half and two thirds of respondents, with symptoms ranging from mild and temporary to severe and permanent, and most cases never formally reported. That is a serious finding about a serious problem, and it is a different problem from this one.
What I am describing is a cost of a different kind, and in day-to-day professional terms it may be the more common one.
Effort. Listening to degraded speech consumes cognitive resources. This is well established, and it is the reason a bad phone line is exhausting in a way a good one is not. In simultaneous interpreting those resources are already fully committed to comprehension, memory, reformulation and production. Whatever the ear spends on extracting the signal is taken directly from the work.
Quality. Everything downstream depends on what you actually heard. There is no interpreting technique that recovers a word that never arrived. Accuracy under a poor input is not a question of skill or effort; it is a question of information that was not there.
Fatigue you cannot account for. This is the part interpreters describe most often and explain least well, coming out of a ninety-minute whispered assignment more depleted than a full booth day, with easy content, and quietly wondering whether the problem is you.
It is not you. You have been decoding speech through a signal path withno gain control while your own vocal apparatus attenuated the input, and doing it continuously, and doing it while producing.
And whispering is not the gentle option. That is the finding I find most striking. Because the reflex threshold sits near the very bottom of the human vocal range, whispering is enough to trigger it. The quietest, most discreet, most considerate way of working, the one we choose precisely because it seems undemanding, still turns the ear down.
What has to change
Not the practice. The conditions.
Stop treating floor audio as something that simply happens to you. For most of our history this was fair enough, because there was nothing to be done about it. That is no longer true, and once it is not true, accepting whatever the room gives you becomes a choice rather than a fact.
Ask for a feed. Where there is a sound system, there is an output. In my experience the technician's answer is very often yes, and the reason we do not get feeds more often is simply that we do not ask. We have been trained to think of a floor feed as something that comes with a booth.
And when there is no feed to ask for, bring your own: clip one of those tiny wireless microphones sold for making TikTok videos onto the speaker, plug its USB receiver into your computer, and you have a live floor feed anywhere — a factory floor, a moving bus, a corridor, a room with no sound system at all.
Treat your input as part of your professional setup, exactly as you treat your voice, your glossary and your preparation. We spend a great deal of time on everything that happens after comprehension, and almost none on whether comprehension had a fair chance.
And say so to clients. Not as a complaint, but as a quality argument. A client who understands that input quality determines output quality is a client who will help you get a feed. A client who has never heard it will assume, quite reasonably, that professionals cope.
One last thought
There is something a little absurd about the position we have been in.
We are a profession that argues, correctly and at length, about acoustic standards for booths. We have ISO standards specifying frequency response forthe audio we receive. We have committees, studies and hard-won technical requirements. All of it built on the entirely sound principle that an interpreter cannot render what an interpreter cannot hear.
And then we walk into a factory, stand behind a delegate, and work with no floor feed at all — and we call it normal.
It is not normal. It is just old.
The ear has been quietly turning itself down every time we open our mouths, for the whole history of this profession, and nobody mentioned it. Now that we know, the question is no longer whether whispered interpreting is harder than it looks.
The question is what we intend to do about it.
----------------------------------------------------------------------------------------------------------------------------------
For anyone who wants to start controlling their own input tomorrow: TerpMate's audio input monitor does exactly this, and it is free, part of the community edition, no subscription, no time limit. Bring the floor into your ears at a level you choose, instead of the level the room chose for you.
I will come back to the output side of this problem — what happens after you have understood, and how interpretation actually reaches your listeners without a console — in a separate piece.