{ "title": "Fish Audio raises $52M seed round for AI voice generation", "body": "Fish Audio, an AI voice model startup based in Palo Alto, has raised a $52 million seed round, the company announced Tuesday. The round is co-led by Coreline Ventures and Capital Today, with participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.\n\nThe company, which launched last year, has grown rapidly. More than 8 million people use open source or hosted versions of Fish Audio models, and the startup now generates annual recurring revenue of $21 million.\n\n## From a single GPU to a voice library of 15,000 controls\n\nFish Audio started as a small project by former Nvidia researcher Shijia Liao, who trained a voice-generation model on a single GPU and open-sourced it. That repository, Fish Speech on GitHub, now has more than 31,000 stars and is used by indie developers, video game designers, and creators. The project’s open-source nature helped it gain traction quickly, attracting a community of contributors and users who helped refine the technology.\n\nSince launch, Fish Audio has released five models: four for speech generation and one for speech-to-text. Three of the speech-generation models are open source, while the latest, S2.1 Pro, is available only via a paid API. The company’s library now includes more than 15,000 natural language controls, giving users fine-grained command over voice output. These controls allow users to adjust pitch, speed, emotion, and other vocal characteristics with simple text commands, making the technology accessible to non-technical creators.\n\n“I think what they’ve been able to build, state-of-the-art models, with the team they have, compared to some of these other well-funded AI labs or companies, is incredible. It shows their technical acumen in closing the gap between artificial-sounding and human-like voices,” said Rico Mallozzi, partner at 359 Capital.\n\nFish Audio offers paid monthly plans for creators and teams, with set minutes and voice-cloning features. It also provides an enterprise version of its APIs and platform. Organizations like HeyGen and Sanas use the enterprise API, and voice agent company LiveKit is among the customer types Fish Audio serves. The company’s pricing structure is designed to scale with usage, from individual creators producing short audio clips to large enterprises generating thousands of hours of voice content per month.\n\nThe startup’s growth has been fueled by the increasing demand for synthetic voice technology in content creation, gaming, virtual assistants, and customer service. Fish Audio’s models are used to generate voiceovers for videos, create character voices for games, and power interactive voice response systems. The company’s open-source models have also been adopted by researchers and hobbyists who are experimenting with voice AI in novel ways.\n\n## Trust and consent after voice upload controversy\n\nFish Audio built its voice library partly by asking users to submit voices and compensating them if used. But months ago, some creators alleged that voices were uploaded without their consent. At the time, Fish Audio had a DMCA takedown process, but takedowns took a long time. The controversy highlighted the challenges of building a voice library at scale while respecting creators’ rights.\n\nRissa Cao, CEO and co-founder of Fish Audio, said the takedown process is now automated. Creators can submit a voice sample or contract to prove ownership, and the voice is removed in less than three minutes. However, the automated system does not prevent unauthorized uploads until the artist discovers the violation and files a removal request. This means the burden of monitoring still falls on creators, a point of contention in the voice AI community.\n\nOsuke Honda, partner at Coreline Ventures, emphasized that the community-driven model only works when creators trust the platform. “Consent, transparency, attribution must be built into the product,” he said. “The industry needs verified voice ownership, clear licensing, easy takedown, and revenue-sharing.” Honda’s comments reflect a broader industry push toward ethical AI development, where companies are expected to proactively protect creators’ rights rather than reactively address violations.\n\nFish Audio has also implemented a revenue-sharing program for voice contributors, where creators receive a portion of the revenue generated when their voice is used in commercial applications. The company says it is working on additional tools to verify voice ownership at the point of upload, potentially using blockchain or other cryptographic methods to create an immutable record of consent. These measures are intended to rebuild trust with the creator community and differentiate Fish Audio from competitors that have faced similar controversies.\n\n## Competing in a crowded market\n\nThe speech-generation market is crowded. Competitors include ElevenLabs, WellSaid, Cartesia, Speechify, Async (previously Podcastle), and Krisp. Despite the competition, Fish Audio initially ran efficiently without outside capital, relying on its open source project and creator plans. The startup sought capital to develop more advanced models and accommodate enterprises as investor interest ramped up.\n\n“Fine-grained controls and cost-efficient training help Fish Audio compete with big AI labs,” Mallozzi said. The company’s small team has been able to produce models that rival those from much larger organizations. Fish Audio’s training infrastructure is optimized for efficiency, allowing it to achieve state-of-the-art results with fewer computational resources than competitors. This cost advantage translates into lower prices for customers, making the technology more accessible.\n\nFish Audio plans to release an audio understanding model this year and is building a speech-to-speech model. The audio understanding model will be able to analyze and interpret audio content, enabling applications like automatic transcription, sentiment analysis, and voice-based search. The speech-to-speech model will allow real-time voice conversion, where a user’s voice can be transformed into another voice with different characteristics while preserving natural intonation and emotion.\n\nThe company aims to cater to both creative use cases, which require more expressive voices, and enterprise needs, such as steerable voices for customer support and sales operations. For creative applications, Fish Audio’s models can generate voices that convey a wide range of emotions, from excitement to sadness, making them suitable for storytelling and character development. For enterprise use, the models can be fine-tuned to produce consistent, professional voices for brand representation and automated customer interactions.\n\nCao said the market for AI-generated voice models is massive, and Fish Audio wants to serve all use cases. With $52 million in seed funding, the company is positioned to accelerate development and expand its enterprise footprint. The funding will be used to hire additional engineers, expand the voice library, and build out the sales and marketing team to target larger enterprise customers. Fish Audio also plans to invest in research and development to improve the naturalness and expressiveness of its models, with the goal of making AI-generated voices indistinguishable from human speech.\n\nThe company’s long-term vision includes building a platform where anyone can create, customize, and monetize their own voice models. This would open up new possibilities for content creators, voice actors, and businesses looking to create unique audio experiences. As the technology matures, Fish Audio expects to see adoption in new verticals such as education, healthcare, and accessibility, where synthetic voices can provide personalized learning experiences, assistive communication tools, and more natural human-computer interaction.\n\n## Related on Neura Market\n- AI Voice Generation Market\n- Startup Funding News\n- Open Source AI Models" }
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.

