Five Seconds of Audio Is All Fish Audio Needs to Clone Your Voice

Fish Audio's S2.1 Pro can clone a voice from just 5 seconds of audio. Here's what the model does, who's already using it, and how it stacks up on cost.

Share
Five Seconds of Audio Is All Fish Audio Needs to Clone Your Voice

Fish Audio's flagship model, S2.1 Pro, can clone a voice from just five seconds of reference audio, with word-level control over emotion, intonation, and pacing baked into the output.

The company also announced $52 million in seed funding this week, marking its first anniversary. But the product itself is the more relevant story for anyone doing voiceover, dubbing, or ADR work: a voice cloning tool that's reportedly both faster and cheaper than the two names most people already default to.

What S2.1 Pro Actually Does

The core pitch is speed and control. A five-second sample is enough to generate a working clone, and Fish Audio's own blind listening tests claim S2.1 Pro is preferred by nearly 67% of listeners over competing models. The platform supports native output across dozens of languages, and offers granular tagging for emotion, pacing, and delivery style rather than a single flat "voice" setting.

Fish Audio positions itself against ElevenLabs and Cartesia specifically, claiming roughly 2x the generation speed of Cartesia and about one-sixth the cost of ElevenLabs for comparable output. For anyone doing high-volume voiceover work — narration, explainer videos, dubbing passes — that kind of cost gap is worth testing directly rather than taking at face value.

Who's Already Using It

Fish Audio says its models are already running in production at HeyGen, LiveKit, Retell, Sanas, and OpenArt. HeyGen in particular is a name BRC readers will recognize from AI avatar coverage — a voice engine already embedded in an avatar platform at that scale is a meaningful trust signal, independent of Fish Audio's own claims.

Origin Story Worth Knowing

Fish Audio began as an open-source side project by Shijia Liao, a former NVIDIA researcher frustrated with flat, robotic synthetic voices. The resulting Fish Speech repository passed 31,000 GitHub stars and became one of the more widely used open voice projects before the company formalized into a funded startup. Three of its speech-generation models remain open source; S2.1 Pro is paid-API only for now, though Fish Audio says it's opening free API access to the model at the end of August.

Competitive Context

The AI voice space is crowded — ElevenLabs, Cartesia, WellSaid, Speechify, and Async all compete in overlapping territory. Fish Audio's differentiation claim is specifically on cost and speed rather than a fundamentally different capability set, which makes it worth a side-by-side test rather than an outright switch for anyone already committed to an existing voice pipeline.

The Signal in the Noise

The speed and cost claims are Fish Audio's own numbers, not independently verified — worth treating as a starting point for testing rather than a settled fact. But the specific use case worth paying attention to here is dubbing and multilingual voiceover work, where per-minute cost adds up fast at scale, and a credible lower-cost alternative is genuinely useful even if it doesn't fully match ElevenLabs on ceiling quality.

If you're doing occasional narration work, the cost difference probably won't move the needle. If you're producing high volumes of voiceover or localizing content into multiple languages regularly, it's worth running your own script through both platforms before committing either way.

Specs & Pricing

  • Voice cloning: from 5 seconds of reference audio
  • Claimed speed: roughly 2x faster than Cartesia
  • Claimed cost: roughly 1/6th the cost of ElevenLabs
  • Free API access to S2.1 Pro: opening end of August 2026
  • Enterprise customers: HeyGen, LiveKit, Retell, Sanas, OpenArt
  • Seed funding: $52 million, led by Coreline Ventures and Capital Today

Resources & Reads