Startup Smallest.ai announced on July 31, 2026 a $13 million Series A to build voice models whose conversation is meant to be indistinguishable from a human's. The round was led by Seligman Ventures. The company targets real-time conversational voice agents rather than general-purpose speech applications.
Key takeaways
- Series A: $13 million, over $21 million raised in total.
- The round was led by Seligman Ventures, joined by Sierra Ventures and 3one4 Capital.
- Two-model architecture: a small real-time voice model plus a larger LLM for complex queries.
- Customers include RingCentral and Truecaller.
- Goal: conversation indistinguishable from a human — breaking the Turing test.
Who is raising and for what
Smallest.ai was founded in late 2024. It builds specialized voice models for real-time conversation, with an emphasis on handling diverse accents, multilingual support and operation in noisy environments. The $13 million Series A was led by Seligman Ventures, with Sierra Ventures and 3one4 Capital participating. The startup's total funding has passed $21 million.
How the model works
According to founder and CEO Sudarshan Kamath, the model mimics human conversation by listening, thinking and speaking at the same time. The approach rests on two models: a small voice model responds in real time with virtually no lag, while a larger LLM is called only for complex queries outside the voice model's knowledge. This differs from classic pipelines, where speech-to-text, processing and speech synthesis run in sequence and stack up latency.
While I'm speaking to you, you're already thinking — and you might interrupt me if I talk for too long.
Sudarshan Kamath, founder and CEO of Smallest.ai.
Market and competition
Customers include RingCentral and Truecaller, and the target market is customer-support companies such as Sierra and Decagon. In the voice-AI market Smallest.ai competes with leader ElevenLabs as well as Cartesia and regional players like Sarvam, focused on local languages. The differentiator is precisely the two-model architecture, cutting latency to a level that makes it hard to tell there is a machine on the other end.
Why it matters
In voice agents, latency decides everything. A conversation stops sounding natural once there is a noticeable pause between question and answer — and classic speech-text-speech pipelines introduce exactly that pause. Splitting the work between a fast voice model and a slower LLM is an attempt to sidestep the trade-off: a cheap, instant model handles simple exchanges, and heavier reasoning kicks in only when needed. If it works, it shifts the boundary of where a voice agent can replace a human — from simple hotlines to longer, contextual conversations. The claim of breaking the Turing test?Turing test: A trial in which a human converses with a machine; the machine "passes" if the interlocutor cannot tell it apart from a person. is marketing, but the direction is real: cost and latency fall enough that mass deployments of voice customer support become economical.
What's next?
- Smallest.ai will use the Series A to develop voice models for real-time agents.
- The company remains in direct competition with ElevenLabs, Cartesia and Sarvam in the voice-AI segment.





