Home  /  Translation under real conditions
Reality check

How well does the translation actually handle noise and people talking over each other?

The showroom is quiet, only one person speaks at a time, and the microphone has an easy job. The real question is what happens once none of those three things are true. Short answer: noise, overlapping speech and accents decide more about the result of any live translation than a spec sheet ever will.

A customer once called us from Zürich HB, standing right under the big departures board, with announcements echoing behind him and a friend talking into his other ear at the same time. He wanted to know whether the Even G2's translation would still hold up there — not just in the quiet corner of the showroom where we'd demoed it to him two weeks earlier. It's the right question, and it deserves more than a shrug.

Short answer: the Even G2's live translation runs on a cloud speech-recognition system, and like almost any such system — from dictation apps to call-centre software — it does its best work under controlled conditions: one speaker, a reasonably quiet room, clear pronunciation. Once background noise, several overlapping voices or strong accents enter the picture, reliability drops noticeably, and that's true of the technology in general, not a quirk of this particular device. We're not going to hand you an invented accuracy percentage here, because we don't have an independently measured figure we'd trust. Instead, here's which condition hurts the most and what you can do about it.

Why "loud" isn't one problem

When people ask "does it work in noise?", they usually mean three quite different situations that are not the same for a speech engine. Steady noise — a fan running, traffic hum, an espresso machine grinding away — is comparatively manageable for a microphone, because it's distinguishable from the actual speech signal and can be partly filtered. Impulse noise — clattering plates, a slamming door, applause — is disruptive but brief. The hardest case is noise that is itself made of human speech: a packed bar, a canteen, an open-plan office with ten conversations happening at once. There, the system has to separate not just speech from non-speech, but one person's speech from everyone else's in the room — and that's the real challenge, not raw volume.

That's also why a claim like "works up to X decibels" wouldn't tell you much even if we had one. A quiet server room with a single low hum can measure louder than a café full of soft chatter, and still be the easier environment for translation.

Condition by condition, honestly

Quiet room, one person speaking

This is the ideal case these systems are built and tuned for — an office, a consultation room, our showroom on Zeltweg 74. Clear standard speech, no overlapping noise, reasonable distance to the microphone. This is where you should expect the system's best result.

Steady background noise — street, train, a restaurant with music

A moderate, constant background is usually still manageable as long as the speaker articulates clearly and isn't too far from the microphone. Once music or traffic noise gets very loud, it gets noticeably harder — true of any ear- or collar-worn microphone, not specific to the G2.

Overlapping voices — several people talking at once

This is the condition that most reliably degrades results. When two people speak simultaneously, the system has to guess which voice is the relevant one, and words from both sentences can bleed together. No live-translation system we know of, from any maker, reliably solves this. The one lever that actually works is the conversation itself: speak one at a time, not over each other.

Strong accents and unfamiliar pronunciation

Speech models are trained on huge amounts of standard speech and comparatively less on strongly regional or foreign-inflected accents. A speaker with a pronounced accent in the source language can therefore trip the system up more often than someone closer to standard pronunciation — regardless of how fluent they actually are. That's not a judgement on the accent, it's a known property of these systems we'd rather not sugar-coat. (Dialects such as Swiss German are a bigger, separate topic we cover on its own.)

Fast speech and short, choppy sentences

Speaking very fast, or breaking off mid-sentence and restarting, gives the system less context to place words correctly. A steadier pace with fuller sentences helps the translation the same way it helps a human listener.

What you can actually control

  • Move to a quieter spot for the part of the conversation that matters — even a few steps from the loudest corner helps.
  • One person speaks, then the other — deliberately avoid talking over each other, especially for the important sentences.
  • Speak at a steady pace with complete sentences rather than choppy fragments or rapid-fire delivery.
  • For a recurring important conversation — a client meeting, an appointment at an office — test the exact setup beforehand rather than relying on it in the moment.
  • If a translation seems unclear, ask the other person to repeat the key point — a safety net no system replaces.

So: buy, skip, or test first?

Buy with realistic expectations if your day-to-day is mostly quieter one-to-one conversations — a consultation, a reception desk, a meeting at a table. That's where translation typically delivers its best result.

Think twice if your main use case is a loud, overlapping environment with many voices at once — a packed Friday-night bar or a busy market stall is a hard test for any live translation, not just this one.

Test before you buy the exact condition that matters to you, not the quiet demo. Bring someone with an accent to the Zürich showroom, try talking over each other on purpose, or ask us to play background music. For more on the feature itself, independent of conditions, see our translation overview page, and for how the Even G2 stacks up overall, the AI glasses database.

Disclosure: AI-Eyewear is the authorised Swiss reseller of Even Realities. We're deliberately not citing an accuracy percentage here — we don't have an independently measured figure we trust, and we'd rather leave it out than sell you false confidence.

Frequently asked

Does live translation work in a loud bar or a busy market stall?

It can, but less reliably than in a quieter setting. Loud, overlapping soundscapes with many simultaneous voices are the hardest condition for practically any live-translation system, not just the Even G2. A slightly quieter spot and speaking one at a time noticeably improve the result.

Why is it a problem if two people speak at the same time?

Because the system then has to guess which voice is the relevant one, and words from both sentences can bleed together. That's a known limitation of speech-recognition systems generally. The most effective fix isn't technical — it's conversational habit: one person speaks, then the next.

Does a strong accent affect translation quality?

Yes, and that's a known property of speech-recognition systems in general: they're trained mostly on large amounts of standard speech, and a pronounced accent can make recognition harder regardless of how fluent the speaker actually is. We don't have a reliable figure for how much any specific accent is affected — the best approach is to test with the voices you'll actually be using day to day.

Does AI-Eyewear quote an accuracy percentage for noise or accents?

No, deliberately not. We don't have an independently measured, trustworthy percentage for specific conditions like noise or accents, and we're not going to invent one just to have a marketing number. Instead we describe honestly which condition tends to hurt more, and recommend testing the exact situation that matters to you at the showroom.

Test translation in exactly the setting that matters to you.