Customer-obsessed science
Research areas
-
August 21, 20269 min readExtendable framework enables testing agents on the full set of capabilities required to successfully complete a procedure, not isolated proxy tasks.
-
July 30, 20268 min read
-
-
July 9, 202610 min read
-
Featured news
-
Code@MIT 20252025Organizations increasingly face the challenge of evaluating numerous potential user experiences, particularly in digital contexts where content can be easily modified through page layouts, personalized recommendations, and automated content generation. While creating these experiences has become increasingly straightforward, especially with the advent of generative AI, determining user preferences remains
-
IDETC-CIE 20252025Deformable packages are becoming increasingly prevalent in the logistics and warehouse industry, demanding robotic manipulation strategies that are robust and adaptive. Unlike rigid objects, these packages undergo significant shape changes under external forces, making their handling more complex. These deformable packages are often manipulated using suction cups and contain internal objects that shift
-
Transportation Research Part E: Logistics and Transportation Review2025The 21st century workforce is increasingly characterized by more flexible labor models, particularly in e-commerce and supply chain operations. While previous research has focused mostly on last-mile, on-the-road settings, we focus on under-the-roof (UTR) environments, which present unique challenges due to their complex, varied tasks requiring training and experience. Our study addresses the need to better
-
2025Large language models (LLMs) have exhibited extraordinary performance in a variety of tasks while it remains challenging for them to solve complex multi-step tasks as agents. In practice, agents are sensitive to the outcome of certain key steps which makes them likely to fail the task because of a subtle mistake in the planning trajectory. Recent approaches resort to calibrating the reasoning process through
-
ICLR 2025 Workshop on Resource-Adaptive Foundation Model Inference (AdaptFM), ICML 2026 Workshop on Resource-Adaptive Foundation Model Inference (AdaptFM)2025Multi-model inference systems—whether based on routing, cascading, or unified strategies—often rely on confidence signals to decide when a small language model (SLM) output should be accepted or deferred. While such signals are commonly used in classification and short-form generation, their reliability in structured generation settings remains poorly understood. In this work, we study log-probability confidence
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all