Customer-obsessed science
Research areas
-
July 10, 20265 min readHydroShear, a new physics-based simulator, teaches robots how to use their sense of touch to perform complex manipulation tasks, in a way that transfers seamlessly to the real world.
-
July 9, 202610 min read
-
-
Featured news
-
2026Test-time scaling via sequential revision has emerged as a powerful paradigm for enhancing Large Language Model (LLM) reasoning. However, standard post-training methods primarily optimize single-shot objectives, creating a fundamental misalignment with multi-step inference dynamics. While recent work treats this as multi-turn reinforcement learning (RL), conventional approaches optimize over the multi-step
-
2026Deploying LLM-based analytics agents in enterprise settings requires evaluation frameworks that can reliably detect failures across complex, multi-tool workflows. We present a three-phase comparative study of three evaluation frameworks, each representing a distinct evaluation paradigm (trace-based LLM judging, text-only LLM judging with red-teaming, and deterministic heuristics), applied to two analytics
-
ECML-PKDD 20262026Next-basket recommendation (NBR) in online grocery must capture both habitual repeat purchases and explore behavior. We propose BasketFormer, a Transformer encoder trained with a contrastive masked language modeling (C-MLM) objective that unifies three innovations: (1) an InfoNCE-based MLM loss replacing the full-vocabulary softmax with in-batch contrastive scoring; (2) a bit-level temporal encoding that
-
ICML 2026 Workshop on Resource-Adaptive Foundation Model Inference (AdaptFM), ICML 20262026Language models struggle to generalize beyond pretraining context lengths, limiting long-horizon reasoning and retrieval. Continued long-context training is effective but expensive due to the quadratic cost of Attention. We observe that most tokens do not require (Global) Attention over the entire sequence and can rely on local context. Based on this, we propose L2A (Learning To Attend), a layer that enables
-
2026Hallucination, broadly referring to unfaithful, fabricated, or inconsistent content generated by LLMs, has wide-ranging implications. Therefore, a large body of effort has been devoted to detecting LLM hallucinations, as well as designing benchmark datasets for evaluating these detectors. In this work, we first establish a desiderata of properties for hallucination detection benchmarks (HDBs) to exhibit
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all