This AI Voice Model Topped a Blind Listening Test Against Real Humans
Bland says its new Speech v3 model ranks above ElevenLabs, Grok, Cartesia, and OpenAI on a blind listening benchmark, second only to real human speech. Here's how the test actually works.
Bland launched Speech v3 this week, calling it "the world's first Human Speech Engine." The company says the model now ranks at the top of DesignArena's Audio Realism Benchmark, a blind listening test where vetted human raters compare AI-generated voices directly against real, recorded human speech, ahead of ElevenLabs, Grok, Cartesia, and OpenAI.
Worth being precise about that claim upfront: this comes from Bland's own announcement and product page, not an independently published leaderboard result BRC could verify directly. DesignArena's benchmark methodology itself is real and well-documented, but the specific ranking is currently self-reported by Bland rather than confirmed by a third party.
How the Benchmark Actually Works
Today, we’re launching Bland Speech v3 - The world's first Human Speech Engine.
— Bland (@usebland) August 4, 2026
In @designarena’s Audio Realism benchmark, Speech v3 is the top model, outranking Elevenlabs, Grok, Cartesia, and OpenAI.
Trained on over 100 million real, human conversations: businesses,… pic.twitter.com/uadpeaBP77
DesignArena's Audio Realism Benchmark, built by intelligence.ai, asks one direct question: of two clips reading the same script, which one sounds more human? Vetted native-speaker listeners hear both versions blind, in randomized order, with no identifying information, and simply pick the more convincing one. Real human recordings are seeded into the same pool as a hidden control group, so every model gets measured not just against competitors, but against actual human speech directly.
Results get compiled using the Bradley-Terry model, the same statistical method behind Elo chess ratings, converting thousands of head-to-head picks into a single score per model. According to Bland, in this specific benchmark, only real human recordings ranked higher than Speech v3, meaning the model's own claim is second-best overall, not first overall ahead of humans themselves. That's a meaningfully different, and more credible, claim than "beats humans," and worth stating precisely rather than rounding up.
What Speech v3 Actually Does

The model was trained on more than 100 million real human conversations, according to Bland, and offers two cloning tiers: an instant clone from roughly 10 seconds of audio, and a professional clone fine-tuned on 30 minutes or more of source material for higher fidelity. Pricing runs a flat $0.015 per 1,000 characters regardless of which tier or interface is used, with new accounts starting with enough free credit for roughly two hours of speech.
A notable production detail: the model can perform bracketed direction tags, like [laughs] or [clears throat], actually acting them out in the generated audio rather than reading them aloud as text. Cloning access requires the account holder to confirm they hold rights to the voice being cloned, a real gating step rather than an open, unrestricted process.
The James Story
Watch the full documentary: pic.twitter.com/Ez35hfPH5d
— Bland (@usebland) August 4, 2026
Bland used the launch to showcase a specific, human use case: helping James, a 49-year-old father who recently suffered a stroke, recover his speaking voice using just 5 seconds of old footage as a reference. This falls into a now-familiar category of AI voice application, similar to work ElevenLabs and others have done for people who've lost their voice to illness or injury, using existing voice-cloning technology to restore something genuinely personal rather than purely for content production.
Bland says a full-length documentary covering the story is available alongside the launch. It's worth approaching content like this with real care rather than treating it as a marketing hook, the underlying use case, helping someone regain a voice they've lost, is legitimate and meaningful independent of how a company chooses to promote it.
Competitive Context

Bland's primary business is enterprise AI phone agents, not a general-purpose voice tool in the mold of ElevenLabs or Fish Audio, so Speech v3 launching as a more broadly usable product is a notable expansion beyond its usual focus. ElevenLabs remains the most established name in the AI voice space generally; Fish Audio, covered previously on this beat, has positioned itself specifically on cost efficiency rather than raw realism claims. Bland's entry adds a third, differently positioned competitor into a space that's gotten meaningfully more crowded and more capable over the past year.
The Signal in the Noise

The benchmark methodology behind this claim is legitimate, and "second only to real human recordings" is a genuinely strong result if it holds up under independent scrutiny.
But it's currently a self-reported result from the company being ranked, not something BRC or another outlet has independently confirmed against DesignArena's live leaderboard.
Worth testing directly against a known reference voice before taking the ranking as settled, and worth remembering that "topped a blind listening test" here means topped among AI models, with real human speech still the actual benchmark it fell just short of.