Gunnar Pétur Hauksson, co-founder and chief commercial officer of Iceland's Treble Technologies, argues in a piece published on August 15, 2026 in The Robot Report that humanoids will not gain wide acceptance until they can hear and speak naturally. His thesis: once robots converge on similar vision and movement, communication quality will decide adoption.
Key takeaways
- Author: Gunnar Pétur Hauksson, co-founder and CCO of Treble Technologies (Reykjavik).
- Thesis: hearing and speech are the most underdeveloped area in robotics, despite gains in vision and locomotion.
- Robot audio and voice systems tend to break down in noisy, real-world acoustic conditions.
- In humans, hearing consumes roughly 15–20% of sensory processing energy, second only to vision.
- Treble offers physically accurate acoustic data with a low sim-to-real?sim-to-real: Transferring a model trained in simulation to the real world. The smaller the sim-to-real gap, the better the model performs outside simulation. gap for training at scale.
Why audio fell behind
Hauksson says the industry bet on vision and movement because they deliver visible, impressive progress and have mature foundations — abundant data, ready-made models, and simulators. Training environments such as NVIDIA Isaac Sim are overwhelmingly visual but, for the most part, silent. Sound is harder to gather and label, because perception depends on spatial relationships and room acoustics, not just the content of a recording.
The consequences are practical. As Hauksson writes, current robot audio and voice systems tend to break down in crowded spaces, industrial halls, or homes with ambient noise.
The biology argument and a higher bar
The author reaches for an evolutionary analogy: in humans, hearing accounts for roughly 15–20% of sensory processing energy, second only to vision.
If a biological intelligence needs that much auditory bandwidth just to navigate and survive, silicon intelligence will not succeed without it.
Gunnar Pétur Hauksson, co-founder and CCO of Treble Technologies
Robots are also judged more harshly than people. We forgive a human slip of the tongue, but not a machine's — a robot that mishears or responds out of sync quickly loses the user's trust.
What Treble proposes
The Reykjavik company builds cloud-native, wave-based acoustic simulation?wave-based acoustic simulation: A sound-simulation method that solves the wave equation, capturing physical phenomena like diffraction and interference — more accurate than simplified geometric methods. (a hybrid wave/geometrical engine) and uses it to generate training data for AI models. The key claim is physically accurate acoustics with a very low gap between simulation and reality, allowing training at scale without manually collecting recordings. Treble offers a Python SDK and the open Treble10 set of room impulse responses?room impulse response: A characterization of how a room transforms sound between a source and a receiver — reflections, reverberation, and attenuation. The basis of realistic acoustic simulation., built with Hugging Face, and lists Arup, L-Acoustics, and Jabra among its customers.
| Stage | What happens |
|---|---|
| 1. Wave simulation | A hybrid wave/geometrical engine computes physically accurate room acoustics |
| 2. RIR data | Room impulse responses are produced — the open Treble10 set |
| 3. ML training | AI models learn on data with a low sim-to-real gap, without manually collecting recordings |
| 4. Deployment | The robot hears and speaks reliably in real-world noise |
Why it matters
This is the voice of a player selling a fix for the problem it describes, and it should be read with that in mind. Even so, the diagnosis hits a real gap: in the humanoid race, attention and capital concentrate on grippers, legs, and vision-language models, while the audio layer is often treated as an add-on. If hardware and locomotion truly standardize, advantage will come from what is harder to copy — fluent interaction in noise. Acoustic data quality then becomes a strategic asset rather than a finishing touch.
What's next?
- Signal to watch: whether simulation vendors (like NVIDIA Isaac) extend today's mostly visual training environments with realistic acoustics.
- Treble is expanding its audio data offering for ML with Hugging Face (the Treble10 set) — further datasets will show adoption scale.
- Practical test: whether humanoid makers start reporting voice robustness in real-world noise, not just quiet demos.
Sources
- The Robot Report — Why robots that can't communicate naturally won't be adopted, says Treble
- Treble Technologies — Acoustic simulation platform





