Customer-obsessed science
Research areas
-
August 26, 20265 min readDiscounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.
-
August 21, 20269 min read
-
July 30, 20268 min read
-
-
July 9, 202610 min read
Featured news
-
2026A/B testing remains the standard for rolling out new features in the technology industry. Each experiment, however, consumes real traffic, engineering effort, and weeks of wall-clock time. Can AI agents—conditioned on behavioral profiles and contextual descriptions of the intervention—simulate outcomes accurately enough to vet candidate treatments before committing live traffic? We formalize this question
-
2026Text-to-image models have made significant strides, producing impressive results in generating images from textual descriptions. However, creating a scalable pipeline for deploying these models in production remains a challenge. Achieving the right balance between automation and human feedback is critical to maintain both scale and quality. While automation can handle large volumes, human oversight is still
-
ACM SIGSPATIAL 20252025In this paper we propose a novel optimization framework for store shelf space planning, specifically tailored for physical stores. The framework leverages machine learning and data-driven techniques to optimize shelf space allocation, aiming to maximize both short-term and long-term business metrics such as sales and profit. Our approach consists of three key components: a geospatial space elasticity model
-
RecSys 20262026Traditional online A/B experimentation limits the number of treatments that can be evaluated concurrently. Bandit-based adaptive experimentation algorithms address this by dynamically reallocating traffic across an order of magnitude more treatments. Yet existing methods involve a fundamental trade-off: single-metric Thompson Sampling-based methods are robust to novelty effects but cannot accommodate multiple
-
arXiv2026We introduce a vocabulary for automated research systems built from one or more agents to make their design choices easier to describe and compare. The vocabulary specifies who the agents are, which operations the harness exposes, who may invoke them, how agents communicate, what state is visible within and across runs, how the next action is chosen, how a run begins, and how outputs are scored. A trajectory
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all