Palo-Alto based startup Fish Audio, which builds AI models for creators and enterprises has raised $50 million in seed funding. Since launching last year, the startup has more than 8 million people using the open-source or hosted versions of its models, and now generates annual recurring revenue of $21 million.
Fish Audio said that it has raised $50 million in a seed round that was led by Coreline Ventures and Capital Today. The funding also saw participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.
According to TechCrunch, Fish Audio originally began as a small project by former NVIDIA researcher Shijia Liao, who, frustrated by non-expressive synthetic voices available on the market, trained a voice generation model on a single GPU, which he open-sourced. The Fish Speech repository on GitHub now has more than 31,000 stars, and is used by indie developers, video game designers, and creators.
READ: Amazon to raise at least $25 billion for AI infrastructure through bond sale (July 7, 2026)
Fish Audio launched five models last year. This includes four speech generation models and one speech-to-text model. While three of its speech generation models have been open sourced, however its latest S2.1 Pro model is available only through its paid API.
The startup offers paid monthly plans suited for creators and teams that unlock a set number of minutes of generation, as well as voice cloning features. It also offers an enterprise version of its APIs and platform, and says organizations like HeyGen, Sanas and Plaud are already using it.
“Every enterprise has different use cases and different preferences.
For example, companies like HeyGen, which use our voices to power AI avatars, want realism in voices; a gaming studio would want expressive voice for their characters; and voice agent companies like LiveKit want more natural-sounding and low-latency voices that are expressive enough for calls,” Fish Audio co-founder and CEO Risa Cao said.
One way the startup has built its library of voices is by simply asking users to submit their own voices for training its models, and compensating them if their voices are used. This led to some controversy, after some creators alleged that their voices were uploaded to Fish Audio without their consent.
READ: OpenAI raises $122 billion in record funding round, IPO plans expected (April 1, 2026)
Faye Dicker, a U.K.-based voice-over artist mentioned that her voice had been cloned and sold online without her consent. Her voice had been downloaded on Fish Audio over 900 times, according to a BBC report from last month.
Fish Audio has a DMCA content take-down process in place to address such concerns, but the take-downs themselves took a long time. Cao told TechCrunch that the company has now automated the take-down process. Creators can easily submit a short voice sample or a contract to prove that an uploaded voice belongs to them, and their voice will be taken off the startup’s platform in less than 3 minutes, she said.
Cao also said when the startup was only offering its product as an open-source project with plans for creators, it was running efficiently and didn’t need money. But it wanted to develop more advanced models, and also wanted to accommodate enterprises as investor interest was ramping up, which led it to seek capital. According to TechCrunch, the company plans to release an audio understanding model this year. It’s also building a speech-to-speech model.


