-
AAAI 20182018We describe an intelligent context-aware conversational system that incorporates screen context information to service multimodal user requests. Screen content is used for disambiguation of utterances that refer to screen objects and for enabling the user to act upon screen objects using voice commands. We propose a deep learning architecture that jointly models the user utterance and the screen and incorporates
-
EUSIPCO 20182018Far-field automatic speech recognition (ASR) is a key enabling technology that allows untethered and natural voice interaction between users and Amazon Echo family of products. A key component in realizing far-field ASR on these products is the suite of audio front-end (AFE) algorithms that helps in mitigating acoustic environmental challenges and thereby improving the ASR performance. In this paper, we
-
SLT 20182018This article presents a whisper speech detector in the far-field domain. The proposed system consists of a long short-term memory (LSTM) neural network trained on log-filterbank energy (LFBE) acoustic features. This model is trained and evaluated on recordings of human interactions with voice-controlled, far-field devices in whisper and normal phonation modes. We compare multiple inference approaches for
-
ICSC 20182018We demonstrate the potential for using aligned bilingual word embeddings to create an unsupervised method to evaluate machine translations without a need for a parallel translation corpus or reference corpus. We explain why movie subtitles differ from other text and share our experimental results conducted on them for four target languages (French, German, Portuguese and Spanish) with English-source subtitles
-
ICDM 20182018Machine Learning and NLP (Natural Language Processing) have aided the development of new and improved user experience features in many applications. We address the problem of automatically identifying the “Start Reading Location” (SRL) of eBooks, i.e. the location of the logical beginning or start of main content. This improves eBook reading experience by taking users automatically to the logical start
Related content
-
August 4, 2020Team awarded $500,000 prize for performance of its Emora socialbot.
-
July 24, 2020New position encoding scheme improves state-of-the-art performance on several natural-language-processing tasks.
-
July 22, 2020Amazon scientists are seeing increases in accuracy from an approach that uses a new scalable embedding scheme.
-
July 22, 2020Dialogue simulator and conversations-first modeling architecture provide ability for customers to interact with Alexa in a natural and conversational manner.
-
July 8, 2020New method extends virtual adversarial training to sequence-labeling tasks, which assign different labels to different words of an input sentence.
-
July 7, 2020Watch the replay of the live interview with Alexa evangelist Jeff Blankenburg.