Customer-obsessed science
Research areas
-
July 30, 20268 min readInstead of compromising among parameter updates dictated by different training objectives, ControlG allocates computational capacity to objectives sequentially and dynamically.
-
-
July 9, 202610 min read
-
Featured news
-
NeurIPS 2021 Workshop on Databases and AI (DBAI)2021While transformers demonstrate impressive performance on many knowledge intensive (KI) tasks, their ability to serve as implicit knowledge bases (KBs) remains limited, as shown on several slot-filling, question-answering (QA), fact verification, and entity-linking tasks. In this paper, we implement an efficient, data-programming technique that enriches training data with KB-derived context and improves
-
NeurIPS 2021 Workshop on Efficient Natural Language and Speech Processing2021Pretraining and then finetuning of large language models is one of the commonly used approaches to achieve good performance in natural language processing (NLP) tasks. However most pre-trained models have large memory footprint and low inference speed. Deploying such large models to applications with latency constraint is challenging. In this work, we focus on accelerating the inference via conditional
-
ICCV 20212021We introduce Video Transformer (VidTr) with separable attention for video classification. Comparing with commonly used 3D networks, VidTr is able to aggregate spatiotemporal information via stacked attentions and provide better performance with higher efficiency. We first introduce the vanilla video transformer and show that transformer module is able to perform spatio-temporal modeling from raw pixels,
-
ICCV 20212021Most action recognition solutions rely on dense sampling to precisely cover the informative temporal clip. Extensively searching the temporal region is expensive for a real-world application. In this work, we focus on improving the inference efficiency of current action recognition backbones on trimmed videos and illustrate that an action model can accurately classify an action with a single pass over the
-
ASRU 20212021Speaker identification typically involves three stages. First, a front-end speaker embedding model is trained to embed utterance and speaker profiles. Second, a scoring function is applied between a runtime utterance and each speaker profile. Finally, the speaker is identified using nearest neighbor according to the scoring metric. To better distinguish speakers sharing a device within the same household
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all