Customer-obsessed science
Research areas
-
July 30, 20268 min readInstead of compromising among parameter updates dictated by different training objectives, ControlG allocates computational capacity to objectives sequentially and dynamically.
-
-
July 9, 202610 min read
-
Featured news
-
arXiv2026Container image pulling accounts for the majority of pod startup time in Kubernetes environments. Standard pull down loads the entire image before the container can start, even when the application accesses only a fraction of the image content at startup. We present SOCI (Seekable OCI), a lazy-loading architecture that enables containers to start without downloading the full image. SOCI builds an external
-
KDD 2026 Workshop on Evaluation and Trustworthiness of Agentic AI2026Evaluating rule compliance in industry requires assessing products against complex regulatory standards using multimodal data sources—a task where both correctness and trustworthiness of automated judgments are critical. Existing approaches either rely on costly human audits, supervised classifiers that demand large-scale labeled training data, or monolithic multimodal models that apply uniform reasoning
-
2026Recent advancement in vision-language models have enabled multi-modal person re-identification (Re-ID), where the system takes both an image and a text query to identify matching individuals. While previous state-of-the-art methods perform well with detailed, sentence-level descriptions, we found that their Recall@1 drops by half when using short, keyword-based queries due to ambiguity, training biases,
-
KDD 2026 Workshop on Evaluation and Trustworthiness of Agentic AI2026Evaluating multi-step diagnostic reasoning in LLM agents remains an open problem. When cause labels are extracted from resolved operational cases (customer-service tickets, incident reports, clinical notes), the resulting gold standards exhibit extreme vocabulary explosion—5,076 unique cause strings from 2,196 tickets on a single symptom, 92% appearing only once—making LLM-as-judge protocols variance-prone
-
IEEE/ACM 2026 International Conference on Automated Software Engineering (ASE 2026), FSE 20262026At Amazon Prime Video, we face the critical operational challenge of managing code deployments during live events and rapid feature releases without causing service outages. Current change control approaches use blanket deployment freezes that block all changes regardless of risk, creating significant developer toil. While prior re-search has explored risky change predictors, these rely on developer-specific
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all