Superunit tested 4 different voice models to see which had the highest resolution rate
“The model that optimizes hardest for sounding like a person on a phone call is also the one that got the most answers out of HR departments.”
Superunit uses AI agents to finish the employment verifications that databases can't. Roughly two thirds of verifications still require someone to call, email, or fax the employer. Superunit's agents place the call, work through the phone menu, get transferred to whoever holds the records, and log every step with timestamps so the result holds up if it's challenged later.
The company has completed more than 200,000 verifications. On employment verification specifically, it closes 70% of files with an average turnaround of 0.82 days.
The voice on the call is not a cosmetic layer. An HR coordinator who finds the caller confusing, odd, or hard to interrupt shares less, transfers less willingly, and hangs up sooner.
The Challenge
Every text-to-speech vendor publishes the same benchmarks. Time to first audio, mean opinion scores, naturalness panels, p95 latency. Superunit had spent months evaluating voice models against those numbers and didn't trust any of them to predict the thing they actually sell, which is a completed verification.
Their working assumption was that the most human-sounding voice would win, and they had a favorite from listening alone. They also knew a listening preference is not a business result. So they designed a test that ignored it.
In their words: "We had strong opinions here. They turned out to be only half right, which is exactly why we didn't score on them."
The Solution
Superunit routed live verification traffic across four voices for five days, August 31 to September 4, 2026, with every arm running at the same time:
- Bland Speech v3
- MiniMax production voice
- Deepgram Aura 2 (aura-2-asteria-en)
- Deepgram Flux (flux-sienna-en)
The voice was chosen at random for every call, not per customer, so each voice saw the same mix of easy and hard employers, morning and afternoon calls, Mondays and Fridays. Traffic split into three even arms, with the Deepgram arm split again between Aura and Flux.
They swapped the voice and nothing else. Prompts, telephony, transfer logic, and logging were identical across arms.
The scoreboard had one column: the share of dialed calls that ended with a completed verification on the first try. They deliberately did not score latency, time to first audio, hangup rate, or how natural the voices sounded to the team.
Bland Speech v3 was built for phone conversations rather than polished narration. It keeps the breaths, stumbles, and pauses that real people make on a call, and treats them as features rather than artifacts to clean up. It streams first audio in about 200 ms, supports voice cloning from about ten seconds of sample audio, and takes performance tags for things like a laugh or a throat-clear.
The Outcome
| Voice | Calls | Completed | Rate |
|---|---|---|---|
| Bland Speech v3 | 6,637 | 278 | 4.19% |
| MiniMax | 6,363 | 212 | 3.33% |
| Deepgram Flux | 3,125 | 103 | 3.30% |
| Deepgram Aura 2 | 3,238 | 104 | 3.21% |
Bland Speech v3 against the other three pooled: 4.19% versus 3.29%, a 27% relative improvement on 19,363 calls with a p-value of roughly 0.0015.
The other three voices finished within 0.12 points of each other. On this metric, Superunit found MiniMax and both Deepgram generations effectively interchangeable.
Superunit was careful about causation, and so are we. The test shows that the voice built to sound like a person on the phone is also the voice that got the most answers out of HR departments. It doesn't prove one caused the other. The mechanism is plausible and the outcome is measured.
Read Superunit's full write-up: We tested four voice models on 19,000 verification calls.
Impact
- 27% more first-call verifications than the other three voices pooled
- Every 1,000 calls now produce roughly 9 more completed files than before
- Bland Speech v3 is the default voice on all Superunit verification calls
The broader lesson is about the scoreboard rather than the winner. The model with the cleanest greetings lost.