Customer-obsessed science
Research areas
-
September 25, 202612 min readUsing the Neuron Kernel Interface, a Reactor–AWS collaboration tackled the dynamic shapes, memory access patterns, and cache management that make real-time autoregressive diffusion hard—building techniques that generalize across models.
-
September 21, 202611 min read
-
-
August 21, 20269 min read
-
July 30, 20268 min read
Featured news
-
ICASSP 20202020Wake word (WW) spotting is challenging in far-field because of not only the interference in signal transmission but also the complexity in acoustic environment. Traditional WW model training requires a large amount of in-domain WW-specific data with substantial human annotations. This prevents the model building in the situation of lacking such data. In this paper we present data-efficient solutions to
-
ICASSP 20202020End-to-end approaches for automatic speech recognition (ASR) benefit from directly modeling the probability of the word sequence given the input audio stream in a single neural network. However, compared to conventional ASR systems, these models typically require more data to achieve comparable results. Well-known model adaptation techniques, to account for domain and style adaptation, are not easily applicable
-
ICSC 20202020We present a novel language adaptable spell checking system that detects spelling errors and suggests context-sensitive corrections in real-time. We show that our system can be extended to new languages with minimal language-specific processing. Available literature majorly discusses spell checkers for English but there are no publicly available systems that can be extended to work for other languages out
-
The Web Conference 20202020In this paper, we introduce an Augmented-Lagrangian-based method to incorporate multiple objectives (MO) in a search-ranking algorithm. Optimizing MOs is an essential and realistic requirement for building ranking models in production. The proposed method formulates MO in constrained optimization and solves the problem in the popular Boosting framework — a novel contribution of our work. Furthermore, we
-
ICASSP 20202020Grapheme-to-phoneme (G2P) models convert a written word into its corresponding pronunciation and are essential components in automatic-speech-recognition and text-to-speech systems. Recently, the use of neural encoder-decoder architectures has substantially improved G2P accuracy for monolingual and multilingual cases. However, most multilingual G2P studies focus on sets of languages that share similar graphemes
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all