TechNewsReel
Live

Fish Audio Raises $52M Seed for Expressive AI Voice Models

The NVIDIA spinout hit $21M ARR and 8 million users in its first year, betting that creator consent and open-source distribution can win a crowded market.

TechNewsReel Newsroom · July 28, 2026

Fish Audio, an AI voice startup founded by former NVIDIA researcher Shijia Liao, has raised $52 million in a seed round led by Coreline Ventures and Capital Today. The company reached $21 million in annual recurring revenue and 8 million users within a year of launch, positioning itself as a rare open-source contender in a market dominated by ElevenLabs, WellSaid, and Cartesia.

Open-Source Roots, Enterprise Ambitions

Fish Audio began as an open-source project addressing what Liao saw as a critical gap: synthetic voices lacked expressiveness and emotional range. The Fish Speech GitHub repository has accumulated over 31,000 stars, signaling strong developer interest. The company now offers five models: four speech generation models (three open-source, one paid S2.1 Pro tier) and one speech-to-text model.

Enterprise customers include HeyGen and Sanas, which integrate Fish Audio's steerable voice APIs into their products. The dual-track strategy—community-driven open-source models alongside paid enterprise services—aims to scale technical capabilities while maintaining grassroots adoption.

Investor Confidence in a Crowded Field

"I think what they've been able to build, state-of-the-art models, with the team they have... is incredible," said Rico Mallozzi, partner at 359 Capital, one of the participating investors. The round also included Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.

Oskue Honda, partner at Coreline Ventures, emphasized the importance of creator trust: "A community-centric approach can only become a durable advantage if creators trust the platform. That means consent, transparency, and attribution must be built into the product rather than treated as afterthoughts."

Consent and Control

Fish Audio has automated its voice take-down process, allowing voices to be removed in less than 3 minutes upon proof of ownership. The feature addresses growing industry concerns about unauthorized voice cloning and intellectual property rights. Competitors in the space include Speechify, Async/Podcastle, and Krisp, alongside the better-funded ElevenLabs.

The $52 million seed round and $21 million ARR indicate strong market demand for expressive, steerable AI voices over generic text-to-speech. However, the company's reliance on user-submitted voice data underscores the tension between rapid model training and creator rights that continues to shape the synthetic media landscape.

Sources

Get a notification when a big story breaks. A few a day at most — no spam.