Robots Atlas>ROBOTS ATLAS
Artificial Intelligence

Fish Audio raises $52M seed for AI voice models

Fish Audio raises $52M seed for AI voice models

Fish Audio, a Palo Alto startup building AI voice models, announced a $52 million seed round on July 28, 2026. The money is meant to speed up development of its speech-generation, voice-cloning and speech-to-text models, already used by more than 8 million people.

Key takeaways

  • Seed round: $52 million, led by Coreline Ventures and Capital Today.
  • More than 8 million users of the open-source and hosted versions.
  • Annual recurring revenue (ARR): $21 million.
  • Five models in a year: four for speech generation and one for speech-to-text — three of them open source.
  • The Fish Speech repository on GitHub has over 31,000 stars.

Open models as a funnel, paid API as revenue

Fish Audio was founded by Shijia Liao, a former Nvidia researcher, with Rissa Cao as CEO. The company runs a hybrid model: it open-sources part of its models while selling access through a paid API. Of the five models released in the past year, three speech generators are open, while the strongest — S2.1 Pro — is available only through the paid API. That explains both the mass user base and the $21 million in recurring revenue. The Fish Speech repository has gathered over 31,000 stars on GitHub, and the developer community drives the visibility the company monetizes on the enterprise and creator side. The business rests on monthly subscriptions for creators and teams plus enterprise API and platform offerings.

8M+users of Fish Audio's open-source and hosted versionsTechCrunch

Control over how the voice sounds as the differentiator

Fish Audio's technical selling point is fine-grained control over the generated voice — the company cites more than 15,000 parameters that steer how speech is delivered. The point is to match the sound to a specific use case, from realistic narrators to expressive game-character voices.

Every enterprise has different use cases and different preferences. Companies like HeyGen want realism in voices, and a gaming studio would want expressive voices for their characters.

Rissa Cao, CEO of Fish Audio. That variety of needs is the argument for a flexible API rather than one universal voice.

The market is crowded, though. Fish Audio competes with ElevenLabs — the most recognizable brand in speech synthesis — as well as Cartesia, Speechify, WellSaid, Async and Krisp. Its differentiators remain the open-source strategy, which most of its larger rivals do not follow, and the emphasis on voice steerability.

Why it matters

Speech synthesis has moved in recent years from an add-on feature to a market of its own, where realism, latency and price decide. A $52 million seed round for a company that already has $21 million in recurring revenue signals that investors treat voice as a separate layer of AI infrastructure, not an appendage to language models. Fish Audio's open-source strategy matters here: in a segment dominated by closed APIs, open models lower the barrier for developers and build a community that advertising alone cannot replicate. The risk is that open code makes non-consensual voice cloning easier — a problem that weighs on the whole industry and may draw regulators. For creators and companies, more real competitors to ElevenLabs means pressure on prices and faster progress in voice quality.

What's next?

  • Fish Audio has announced an audio-understanding model and a speech-to-speech model as its next releases.
  • The seed funding is meant to fuel model development and expand the enterprise offering alongside the existing creator subscriptions.
  • The open nature of some models will keep tension between accessibility and the risk of voice-cloning abuse — an area of possible regulation.

Sources

Share this article