Customer-obsessed science
Research areas
-
July 30, 20268 min readInstead of compromising among parameter updates dictated by different training objectives, ControlG allocates computational capacity to objectives sequentially and dynamically.
-
-
July 9, 202610 min read
-
Featured news
-
Interspeech 20212021The presence of multiple talkers in the surrounding environment poses a difficult challenge for real-time speech communication systems considering the constraints on network size and complexity. In this paper, we present Personalized PercepNet, a real-time speech enhancement model that separates a target speaker from a noisy multi-talker mixture without compromising on complexity of the recently proposed
-
Interspeech 20212021This paper describes two novel complementary techniques that improve the detection of lexical stress errors in non-native (L2) English speech: attention-based feature extraction and data augmentation based on Neural Text-To-Speech (TTS). In a classical approach, audio features are usually extracted from fixed regions of speech such as the syllable nucleus. We propose an attention-based deep learning model
-
Interspeech 20212021Masked language models have revolutionized natural language processing systems in the past few years. A recently introduced generalization of masked language models called warped language models are trained to be more robust to the the types of errors that appear in automatic or manual transcriptions of spoken language by exposing the language model to the same types of errors during the training of language
-
ICASSP 20212021Prosody is an integral part of communication, but remains an open problem in state-of-the-art speech synthesis. There are two major issues faced when modelling prosody: (1) prosody varies at a slower rate compared with other content in the acoustic signal (e.g. segmental information and background noise); (2) determining appropriate prosody without sufficient context is an ill-posed problem. In this paper
-
ICASSP 20212021In this paper, we introduce Kathaka, a model trained with a novel two-stage training process for neural speech synthesis with contextually appropriate prosody. In Stage I, we learn a prosodic distribution at the sentence level from melspectrograms available during training. In Stage II, we propose a novel method to sample from this learnt prosodic distribution using the contextual information available
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all