Back to blog
Product & company news

How our research team developed Bland TTS v3

How Bland’s NVAlign uses a second model to listen to generated speech, improve non-verbal sounds, and preserve the voice around them.

4 min read

Building speech that sounds human means paying attention to the details of conversation: how a voice pauses, changes pace, or laughs before finishing a sentence. Those details are difficult to teach. A model can pronounce every word correctly and still skip a requested laugh or produce one that sounds disconnected from the speaker.

Our research team developed NVAlign to help speech models produce those sounds more reliably. Below, we’ll explain how we use a second model to listen to generated speech and provide feedback during training, what improved in our tests, and what still needs work.

A tag is easy to write and hard to perform#

A developer can add a tag such as [laughs] or [sighs] in seconds. For the model, following that instruction is much harder.

Take this line:

I thought you were serious. [laughs] You almost had me.

The model has to place the laugh at the right moment, keep the same voice, and return to speech naturally. It still needs to pronounce every word correctly. If the laugh sounds disconnected from the sentence, listeners will notice.

Why examples alone were not enough#

We started by training on recordings paired with transcripts that labeled non-verbal sounds. This teaches the model to associate a tag such as [laughs] with the corresponding audio.

But clean, labeled examples of less common sounds are hard to find. Even laughter varies widely between speakers. After training on those examples, the models could produce many of the sounds, but they still skipped or confused them.

We needed to evaluate the audio the model produced and use that feedback in further training.

We trained a second model to listen#

Our research team trained a separate speech-recognition model to recognize tags and non-verbal sounds. We then kept that model fixed while using it to grade the speech model’s output.

During NVAlign training:

  1. The speech model receives text containing a tag such as [laughs].
  2. It generates the audio.
  3. The recognition model scores how confidently it can detect the requested sound.
  4. We use that feedback to adjust how the speech model plans and produces the audio.

Passing the feedback through every step of audio generation would require a large amount of memory and computation. The team developed a shortcut that uses two selected steps to calculate the training update. Developers don’t need to manage that process when generating speech; it happens during model training.

Qiaolin Wang and the rest of our research team explain the method in the NVAlign paper.

A model can learn the wrong lesson#

A model can improve its tag-following score while making the rest of the audio worse. It may force the sound too loudly, change the speaker’s voice, or distort the sentence around the tag.

We saw this when we removed safeguards during testing. The model could earn a higher score for the requested sound while receiving worse perceptual scores.

We added penalties for changes that damage speaker identity or audio quality, produce excessive peaks or low volume, or move too far from the model’s existing behavior. We wanted it to produce the requested laugh and still sound like the same person afterward.

What listeners heard#

We tested NVAlign on two public speech models and an English production system. Alongside the benchmark, we ran a separate listening test using 100 texts. Three listeners assessed whether each requested sound appeared in the correct position without knowing which system generated the audio. We counted a result as correct only when all three agreed.

For the English production system, accuracy on that separate listening test rose from 39.8% to 54.0%. Listeners rated naturalness at 3.95 out of five, compared with 3.89 before NVAlign training.

There were trade-offs. English word-error rate increased on both evaluation sets, and one automated measure of the vocalization’s perceptual effect declined. We still need to check pronunciation and how each sound fits the sentence, even when the model follows the tag correctly.

For developers, the benefit is more dependable control over delivery through the text they supply. A requested laugh or sigh is more likely to appear where they placed it.

What this means for customers#

  • Developers can direct delivery inside the text instead of editing audio after generation
  • Voice agents can respond with more than words while staying in the flow of a call
  • A single voice can carry changes in pacing, emphasis, and emotion across a conversation
  • Teams can design more specific moments without rebuilding the surrounding voice experience

Where we are still improving#

This work focuses on discrete, tagged sounds. It does not establish that a model knows when laughter is appropriate or can handle every continuous change in emotion and speaking style. The developer still chooses where to request the sound.

Rare sounds remain harder to train because we have fewer good examples. Both the speech model and the recognition model need that variety to improve.

Read the research#

See Bland on your actual call volume.

10 to 15 minutes with the team that ships your first agent. We come prepared with answers, not a pitch deck.

Book a call