Customer-obsessed science
Research areas
-
September 25, 202612 min readUsing the Neuron Kernel Interface, a Reactor–AWS collaboration tackled the dynamic shapes, memory access patterns, and cache management that make real-time autoregressive diffusion hard—building techniques that generalize across models.
-
September 21, 202611 min read
-
-
August 21, 20269 min read
-
July 30, 20268 min read
Featured news
-
EMNLP 20192019The need for high-quality, large-scale, goal-oriented dialogue datasets continues to grow as virtual assistants become increasingly widespread. However, existing publicly available datasets useful for this area are limited either in their size, linguistic diversity, domain coverage, or annotation granularity. We introduce the MultiDoGO dataset to overcome these limitations. With a total of over 65,000 dialogues
-
EMNLP 2019 Workshop on Machine Reading for Question Answering2019Although advances in neural architectures for NLP problems as well as unsupervised pretraining have led to substantial improvements on question answering and natural language inference, understanding of and reasoning over long texts still poses a substantial challenge. Here, we consider the task of question answering from full narratives (e.g., books or movie scripts), or their summaries, tackling the NarrativeQA
-
IDCM 20192019How can we early warn against an impending student drop out or an adverse health condition in near real-time? How can we leverage recent interventions such as tutoring or medicines to early warn more accurately? More challengingly, how do we learn to early warn from data that is peppered with such interventions? Early warnings are pivotal for avoiding long-term problems in healthcare, education, mechanical
-
IDCM 20192019A key obstacle in automated analytics and metalearning is the inability to recognize when different datasets contain measurements of the same variable. Because provided attribute labels are often uninformative in practice, this task may be more robustly addressed by leveraging the data values themselves rather than just relying on their arbitrarily selected variable names. Here, we present a computationally
-
SIGMOD 20192019Modern companies and institutions rely on data to guide every single decision. Missing or incorrect information seriously compromises any decision process. We demonstrate Deequ, an Apache Spark-based library for automating the verification of data quality at scale. This library provides a declarative API, which combines common quality constraints with user-defined validation code, and thereby enables unit
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all