Customer-obsessed science
Research areas
-
July 10, 20265 min readHydroShear, a new physics-based simulator, teaches robots how to use their sense of touch to perform complex manipulation tasks, in a way that transfers seamlessly to the real world.
-
July 9, 202610 min read
-
-
Featured news
-
ICML 20222022We propose simple active sampling and reweighting strategies for optimizing min-max fairness that can be applied to any classification or regression model learned via loss minimization. The key intuition behind our approach is to use at each timestep a data point from the group that is worst off under the current model for updating the model. The ease of implementation and the generality of our robust formulation
-
Interspeech 20222022Recurrent Neural Network Transducers (RNN-T) — a streaming variant of end-to-end models — became very popular in recent years. Since RNN-T networks condition the future output label sequence on all previous labels, the natural search space is represented by a tree. In contrast, hybrid systems employ limited-context language models, where the natural search space is a network, i.e. a lattice. While this
-
Interspeech 20222022Many of the recent advances in speech separation are primarily aimed at synthetic mixtures of short audio utterances with high degrees of overlap. Most of these approaches need an additional stitching step to stitch the separated speech chunks for long form audio. Since most of the approaches involve Permutation Invariant training (PIT), the order of separated speech chunks is nondeterministic and leads
-
Interspeech 20222022Automatic Speech Recognition (ASR) systems have found their use in numerous industrial applications in very diverse domains creating a need to adapt to new domains with small memory and deployment overhead. In this work, we introduce domain-prompts, a methodology that involves training a small number of domain embedding parameters to prime a Transformer-based Language Model (LM) to a particular domain.
-
Interspeech 20222022The availability of data in expressive styles across languages is limited, and recording sessions are costly and time consuming. To overcome these issues, we demonstrate how to build low-resource, neural text-to-speech (TTS) voices with only 1 hour of conversational speech, when no other conversational data are available in the same language. Assuming the availability of non-expressive speech data in that
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all