Customer-obsessed science
Research areas
-
August 21, 20269 min readExtendable framework enables testing agents on the full set of capabilities required to successfully complete a procedure, not isolated proxy tasks.
-
July 30, 20268 min read
-
-
July 9, 202610 min read
-
Featured news
-
ICASSP 20202020Wake word (WW) spotting is challenging in far-field because of not only the interference in signal transmission but also the complexity in acoustic environment. Traditional WW model training requires a large amount of in-domain WW-specific data with substantial human annotations. This prevents the model building in the situation of lacking such data. In this paper we present data-efficient solutions to
-
ICASSP 20202020End-to-end approaches for automatic speech recognition (ASR) benefit from directly modeling the probability of the word sequence given the input audio stream in a single neural network. However, compared to conventional ASR systems, these models typically require more data to achieve comparable results. Well-known model adaptation techniques, to account for domain and style adaptation, are not easily applicable
-
ICSC 20202020We present a novel language adaptable spell checking system that detects spelling errors and suggests context-sensitive corrections in real-time. We show that our system can be extended to new languages with minimal language-specific processing. Available literature majorly discusses spell checkers for English but there are no publicly available systems that can be extended to work for other languages out
-
The Web Conference 20202020In this paper, we introduce an Augmented-Lagrangian-based method to incorporate multiple objectives (MO) in a search-ranking algorithm. Optimizing MOs is an essential and realistic requirement for building ranking models in production. The proposed method formulates MO in constrained optimization and solves the problem in the popular Boosting framework — a novel contribution of our work. Furthermore, we
-
ICASSP 20202020Grapheme-to-phoneme (G2P) models convert a written word into its corresponding pronunciation and are essential components in automatic-speech-recognition and text-to-speech systems. Recently, the use of neural encoder-decoder architectures has substantially improved G2P accuracy for monolingual and multilingual cases. However, most multilingual G2P studies focus on sets of languages that share similar graphemes
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all