Customer-obsessed science
Research areas
-
August 26, 20265 min readDiscounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.
-
August 21, 20269 min read
-
July 30, 20268 min read
-
-
July 9, 202610 min read
Featured news
-
RecSys 20262026Off-policy evaluation ( OPE) estimates the performance of new recommendation policies using logged data, thus enabling fast, safe and inexpensive iteration prior to costly A/B tests. To evaluate ranking policies, existing OPE estimators all make structural assumptions about user behavior, leading to a spectrum of trade-offs between bias and variance. The recently proposed INTERPOL estimator navigates these
-
EMNLP 20262026RAG systems are increasingly used to summarize what large collections of documents say. A user asks “What do people think about X?” and receives an answer that reads as consensus. But standard top-k retrieval ranks documents by query similarity, not by how faithfully they represent the population, so minority views quietly disappear. Existing fixes fall short. Diversity re-rankers like MMR and DPP spread
-
The 11th Workshop on Financial Technology and Natural Language Processing2026Long-context financial document understanding and reasoning pose significant challenges for small language models (SLMs). In this paper, we scale the financial long-context reasoning capability of SLMs through reinforcement learning. Specifically, we propose an efficient curriculum reinforcement learning recipe that features staged training across context lengths and difficulty-aware data sampling. To support
-
2026Optical Character Recognition (OCR) is a fundamental task for digitizing information, serving as a critical bridge between visual data and textual understanding. While modern Vision-Language Models (VLM) have achieved high accuracy in this domain, they predominantly rely on autoregressive decoding, which becomes computationally expensive and slow for long documents as it requires a sequential forward pass
-
2026Creating authentic agentic twins of mobile users through realistic user behavior simulation is critical to truly understand and anticipate customer needs at scale. Toward this, we introduce Agentic Twins of Mobile Users (AgenTwin), an end-to-end framework for high-fidelity mobile user simulation that addresses four key limitations in existing approaches: subjective decision-making diversity, scalable experience
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all