Customer-obsessed science
Research areas
-
August 26, 20265 min readDiscounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.
-
August 21, 20269 min read
-
July 30, 20268 min read
-
July 29, 20266 min read
-
Featured news
-
2026Domain-specific finetuning is essential for dense retrievers, yet not all data pairs contribute equally to the learning process. We introduce OPERA1 , a data pruning framework that exploits this heterogeneity to improve both the effectiveness and efficiency of retrieval model adaptation. We first investigate static pruning (SP), which retains only high-similarity query document pairs, revealing an intrinsic
-
EMNLP 20262026Evolutionary prompt optimizers such as GEPA suffer from prompt bloat: each iteration appends rules and caveats, producing prompts up to 3×longer yet no more accurate. We trace this to three deficiencies - incomplete error observation, limited search diversity, and unreliable selection - and propose ESPO (Error-Structured Prompt Optimization), which decomposes prompt optimization into three phases: Diagnose
-
EMNLP 20262026Fine-tuning large language models (LLMs) for e-commerce attribute extraction requires labeled data representative across thousands of product types, attributes, and multiple languages. This combinatorial scale translates to millions of annotations, rendering human labeling prohibitively costly. While recent work has demonstrated synthetic label generation using LLMs (Negri et al., 2025), deploying such
-
2026Agent evaluation today depends on per-trace LLM-judge inference or human review, too expensive to run on every trace; production systems fall back to sampling a fraction. We find that failing agents leave a detectable behavioral signature in standard observability telemetry: disproportionate effort relative to outcome. We formalize four behavioral failure signatures from agent telemetry, validated on 4,671
-
2026Reinforcement-learning (RL) post-training of tool-using LLM agents leaves expensive rollout GPUs idle while each trajectory waits on CPU-side environment work: sandbox cold-start, code execution, retrieval, and external API calls. We term this idle the environment bubble and measure it directly on 8×A100 hardware using a 20 Hz NVML profiler that avoids the kernel-residency pitfall in the utilization counter
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all