Customer-obsessed science
Research areas
-
August 26, 20265 min readDiscounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.
-
August 21, 20269 min read
-
July 30, 20268 min read
-
-
July 9, 202610 min read
Featured news
-
2026Compound retrieval-augmented question-answering (QA) systems present a fundamental evaluation challenge: manual annotation does not scale, yet automated evaluation lacks the ground truth necessary for calibration. We introduce a self-improving evaluation architecture that addresses this circular dependency through three contributions. First, iterative consensus synthesis: an algorithm that treats LLM-human
-
RecSys 20262026Industrial recommender systems typically operate in two stages: retrieving a candidate set from a large catalog, then ranking those candidates using contextual information. The ranking stage relies on features that summarize a user's prior interactions with the system. These features are often carefully hand-crafted, and designing and maintaining them is time-consuming and computationally expensive. In
-
2026Large language models (LLMs) are increasingly used as judges to evaluate, rank, and supervise other models, yet their reliability in judging LLMs' reasoning process under long-context settings remains underexplored. Existing benchmarks either overly rely on human annotators, who may miss subtle flaws in lengthy reasoning chains, or focus solely on final responses while ignoring the underlying context and
-
2026This paper presents SOMO, a scalable framework for cross-embodiment grasp synthesis that transfers to novel robot hands using only their hand description (i.e., a Unified Robot Description Format (URDF) file), without requiring any hand–object interaction annotations. Unlike prior approaches that rely on hand-specific models or annotated grasp data for each embodiment, SOMO introduces a shared Morphology-Prior
-
2026Answering questions accurately and efficiently in embodied scenarios presents significant challenges due to limited computational and memory resources for Vision Language Model (VLM) inference. Existing methods adopt visual search key frame retrieval method to select critical question-related key frames for VLM input. However, visual search methods are inefficient because they require visual search among
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all