Conversational AI

More-inclusive speech recognition with cross-utterance rescoring

In a top-3% paper at ICASSP, Amazon researchers adapt graph-based label propagation to improve speech recognition on underrepresented pronunciations.

June 9, 2023

Automatic-speech-recognition (ASR) models, which convert speech to text in voice agents, typically have two stages. The first stage involves a deep neural network that maps acoustic information representing an utterance to multiple hypotheses about the words spoken. The second stage is a language model that evaluates (rescores) the plausibility of these hypothesized word sequences.

The first stage — the acoustic model — is optimized for average performance on a large set of speakers; consequently, it tends to perform poorly on speech varieties that are underrepresented in the training set, such as pronunciations found in regional accents. Standard rescoring methods cannot correct for this type of majoritarian bias in the first-stage speech recognizer.

Everyone attending #ICASSP2023, please come by our poster presentation, "Cross-utterance ASR Rescoring with Graph-based Label Propagation".
Time: June 9, 8:15-9:45 AM Greece time
Location: Poster Area 4 - Garden (Speech Recognition: Modeling and Context) pic.twitter.com/R9vqCQODNn
— Srinath Tankasala (@srinath_tank) June 6, 2023

At this year’s International Conference on Acoustics, Speech, and Signal Processing (ICASSP), we presented a new approach to rescoring speech recognition hypotheses that can help recover from errors on speech that is underrepresented in, or otherwise mismatched to, the training data.

Our approach builds a graph from speech samples with different speakers but similar hypotheses, and it creates edges between utterances that sound similar. It then boosts the probabilities of the hypotheses that are shared by adjacent nodes in the graph, meaning that similar-sounding utterances cause similar hypotheses to be boosted. This has the effect that pronunciations of words that are unlikely in isolation can support each other if they are consistent across multiple utterances.

Graph construction

We consider the case in which the initial transcription hypotheses are produced by a fully trained, recursive-neural-network-transducer (RNN-T) ASR model. An RNN-T model is an encoder-decoder model, meaning that it has an encoder module that maps inputs to a representational space and a decoder module that uses those mappings — known as embeddings — to generate ASR hypotheses.

To rescore these hypotheses, we adapt the technique of graph-based label propagation to propagate labels from labeled to unlabeled examples. In our case, the graph nodes represent speech embeddings, and the labels are the ASR hypotheses from the first recognition pass.

ASR rescoring.png — An overview of ASR hypothesis rescoring using graph-based label propagation (LP).

The first step in our graph construction method is to select the data for inclusion in the graph. We divide the data into groups of utterances with substantial overlap in their ASR hypotheses, and we construct a separate graph for each such group. A single graph, for instance, might consist largely of similarly phrased queries about the weather.

Label propagation

In the setting of semi-supervised learning, the graphs include some annotated data, whose transcripts are highly accurate, and larger quantities of unannotated data. We use standard graph-based label propagation algorithms to distribute “goodness scores” for different ASR hypotheses across the graph. Essentially, these algorithms are designed to minimize radical discontinuities in label values between connected (i.e., similar) graph nodes.

The idea is that, even if the ASR model has assigned a low confidence score to the correct transcription of an utterance that features nonstandard pronunciations, the embedding of that utterance will share edges with utterances where the correct transcription receives high confidence scores. The correct transcription will then propagate across that region of the graph, and the odds will increase that the utterance with nonstandard pronunciation is transcribed correctly.

About the Author

Venkatesh Ravichandran

Venkatesh Ravichandran is an applied-science manager in the Alexa Speech organization.

More-inclusive speech recognition with cross-utterance rescoring

In a top-3% paper at ICASSP, Amazon researchers adapt graph-based label propagation to improve speech recognition on underrepresented pronunciations.

Graph construction

Label propagation

Related content

Work with us