Customer-obsessed science
Research areas
-
September 21, 202611 min readThree new papers from Amazon Bio Discovery address bottlenecks in AI-driven antibody engineering, from benchmarking binding predictors to experimentally validating de novo design.
-
-
August 21, 20269 min read
-
July 30, 20268 min read
-
Featured news
-
EMNLP 20262026Production tool calling for agentic systems must satisfy four operational requirements: lower per-invocation cost, low latency, adaptability to evolving tool catalogs, and the ability to catch errors before execution. The ReAct paradigm keeps the large language model (LLM) in the loop, but at production scale, reprocessing intermediate results inflates cost and latency and degrades answer quality. Programmatic
-
EMNLP 20262026Root Cause Analysis (RCA) for defects in large e-commerce enterprises is manual and slow: a single defect spanning thousands of microservices and SOPs takes experts one to two weeks to diagnose. Generic agentic RCA frameworks fail on this workload because they stop at proximate causes, cannot verify numeric claims, ignore past reviewer decisions, and return enumerated hypotheses rather than the quantified
-
PSC 2026 (International Conference on Photonics in Switching and Computing)2026We analyze Optical Add/Drop Multiplexing architecture trade-offs for Cloud Content Providers focusing on G-SNR, operations, reliability, and cost/power. We show that as coherent baud-rates increase, directly attaching transponders to the WSS satisfies all key metrics.
-
COLM 2026— Workshop on Context Beyond the Window2026Deploying LLMs for enterprise Text-to-SQL is bottlenecked less by the model than by what context reaches it: business logic spans thousands of tables, and no model can ingest a full catalog at once. We argue that the most effective place to intervene is therefore the knowledge-base context the model consumes, and that this context should be constructed from historical usage rather than tuned for as a fixed
-
2026Single-metric task-success scores systematically misrank LLM agents in domain-specific deployments: two configurations can achieve identical end-to-end accuracy while exhibiting structurally different production behaviors — silent argument hallucination, infinite tool loops, redundant retries, semantically-equivalent-but-syntactically-divergent queries. We propose a behavioral evaluation protocol that decomposes
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all