Customer-obsessed science
Research areas
-
September 25, 202612 min readUsing the Neuron Kernel Interface, a Reactor–AWS collaboration tackled the dynamic shapes, memory access patterns, and cache management that make real-time autoregressive diffusion hard—building techniques that generalize across models.
-
September 21, 202611 min read
-
-
August 21, 20269 min read
-
July 30, 20268 min read
Featured news
-
ICASSP 20202020In this paper, we present an end-to-end deep convolutional neural network operating on multi-channel raw audio data to localize multiple simultaneously active acoustic sources in space. Previously reported deep-learning-based approaches work well in localizing a single source directly from multi-channel raw audio but are not easily extendable to localize multiple sources due to the well-known permutation
-
ICASSP 20202020We study few-shot acoustic event detection (AED) in this paper. Few-shot learning enables detection of new events with very limited labeled data. Compared to other research areas like computer vision, few-shot learning for audio recognition has been understudied. We formulate the few-shot AED problem and explore different ways of utilizing traditional supervised methods for this setting as well as a variety
-
ICASSP 20202020We propose an approach for pre-training speech representations via a masked reconstruction loss. Our pre-trained encoder networks are bidirectional and can therefore be used directly in typical bidirectional speech recognition models. The pre-trained networks can then be fine-tuned on a smaller amount of supervised data for speech recognition. Experiments with this approach on the LibriSpeech and Wall Street
-
ICASSP 20202020We propose a Text-to-Speech method to create an unseen expressive style using one utterance of expressive speech of around one second. Specifically, we enhance the disentanglement capabilities of a state-of-the-art sequence-to-sequence based system with a Variational Auto Encoder (VAE) and a Householder Flow. The proposed system provides a 22% KL-divergence reduction while jointly improving perceptual metrics
-
ICASSP 20202020In large-scale domain classification, an utterance can be handled by multiple domains with overlapped capabilities. However, only a limited number of ground-truth domains are provided for each training utterance in practice, while knowing as many correct target labels as possible is helpful for improving the model performance. In this paper, given one ground-truth domain for each training utterance, we
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all