Customer-obsessed science
Research areas
-
August 26, 20265 min readDiscounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.
-
August 21, 20269 min read
-
July 30, 20268 min read
-
-
July 9, 202610 min read
Featured news
-
2026Large language model (LLM) agents increasingly operate in streaming case-based reasoning (CBR) settings, where continuous improvement from past experience is crucial. Existing methods achieve this by storing past cases and retrieving similar ones as few-shot examples. This strategy fails near decision boundaries, where highly similar cases have conflicting outcomes and the discriminative factors are not
-
ICML 2026 Workshop on Deep Learning for Code (DL4C)2026Recent agentic approaches to LLM-based kernel generation have achieved strong results on CUDA, yet emerging AI accelerators such as AWS Trainium and Inferentia remain unaddressed. Writing kernels for these chips via the Neuron Kernel Interface (NKI) is particularly challenging due to a multi-engine architecture, tile-based programming with a fixed 128-element partition dimension, and explicit memory management
-
2026Large language models can memorize information that must be removed–ranging from copyright-sensitive content (e.g., book chapters) to personally identifiable information (e.g., income)–to ensure responsible and compliant behavior. Unlearning has emerged as an efficient alternative to full retraining, aiming to remove specific knowledge. However, users may still expect model to leverage the removed information
-
2026E-commerce assistants must go beyond product search to support idea inspiration, criteria formation, comparison, and tool-grounded fact-checking over non-linear shopping journeys. Teaching these behaviors into deployable latency-constrained models is bottlenecked by post-training data: trajectories must cover the full agentic workflow with diversity and fidelity, yet desired outputs are open-ended (often
-
2026Large-scale AI evaluation increasingly relies on aggregating binary judgments from K annotators, including LLMs used as judges. Most classical methods, e.g., Dawid-Skene or (weighted) majority voting, assume annotators are conditionally independent given the true label Y ∈ {0, 1}, an assumption often violated by LLM judges due to shared data, architectures, prompts, and failure modes. Ignoring such dependencies
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all