Customer-obsessed science
Research areas
-
August 26, 20265 min readDiscounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.
-
August 21, 20269 min read
-
July 30, 20268 min read
-
-
July 9, 202610 min read
Featured news
-
IEEE ICMLA 20222022In this paper, we present the design of a robust deep neural network based speech enhancement (DNNSE) solution for joint noise reduction and dereverberation under real-world acoustic conditions. This makes our proposed solution suitable for smart-speaker products that encounter a wide variety of acoustic challenges during their real-world deployment. We provide a systematic introduction to the acoustic
-
SLT 20222022In expressive speech synthesis it is widely adopted to use latent prosody representations to deal with variability of the data during training. Same text may correspond to various acoustic realizations, which is known as a one-to-many mapping problem in text-to-speech. Utterance, word, or phoneme-level representations are extracted from target signal in an auto-encoding setup, to complement phonetic input
-
SLT 20222022Differential privacy (DP) is one data protection avenue to safeguard user information used for training deep models by imposing noisy distortion on privacy data. Such a noise perturbation often results in a severe performance degradation in automatic speech recognition (ASR) in order to meet a privacy budget ε. Private aggregation of teacher ensemble (PATE) utilizes ensemble probabilities to improve ASR
-
NeurIPS 20222022Pretraining on large unlabeled datasets has been proven to improve the down stream task performance on many computer vision tasks, such as 2D object detection and video classification. However, for large scale 3D scenes, such as outdoor LiDAR point clouds, pretraining is not widely used. Due to the special data characteristics of large 3D point clouds, approaches for 2D pretraining frameworks tend to not
-
SLT 20222022End-to-end speech recognition models trained using joint Connectionist Temporal Classification (CTC)-Attention loss have gained popularity recently. In these models, a non-autoregressive CTC decoder is often used at inference time due to its speed and simplicity. However, such models are hard to personalize because of their conditional independence assumption that prevents output tokens from previous time
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all