Attentive contextual carryover for multi-turn end-to-end spoken language understanding

Kai Wei; Thanh Tran; Feng-Ju (Claire) Chang; Kanthashree Mysore Sathyendra; Thejaswi Muniyappa; Jing Liu; Anirudh Raju; Ross McGowan; Nathan Susanj; Ariya Rastrow; Grant Strimel

Publication

Attentive contextual carryover for multi-turn end-to-end spoken language understanding

By Kai Wei, Thanh Tran, Feng-Ju (Claire) Chang, Kanthashree Mysore Sathyendra, Thejaswi Muniyappa, Jing Liu, Anirudh Raju, Ross McGowan, Nathan Susanj, Ariya Rastrow, Grant Strimel

2021

Download Copy BibTeX

Share

Download

Copy BibTeX

Share

Recent years have seen significant advances in end-to-end (E2E) spoken language understanding (SLU) systems, which directly predict intents and slots from spoken audio. While dialogue history has been exploited to improve conventional text-based natural language understanding systems, current E2E SLU approaches have not yet incorporated such critical contextual signals in multi-turn and task-oriented dialogues. In this work, we propose a contextual E2E SLU model architecture that uses a multi-head attention mechanism over encoded previous utterances and dialogue acts (actions taken by the voice assistant) of a multi-turn dialogue. We detail alternative methods to integrate these contexts into the state-of-the-art recurrent and transformer-based models. When applied to a large de-identified dataset of utterances collected by a voice assistant, our method reduces average word and semantic error rates by 10.8% and 12.6%, respectively. We also present results on a publicly available dataset and show that our method significantly improves performance over a non-contextual baseline.

Attentive contextual carryover for multi-turn end-to-end spoken language understanding

Latest news

Work with us