Customer-obsessed science
Research areas
-
August 26, 20265 min readDiscounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.
-
August 21, 20269 min read
-
July 30, 20268 min read
-
-
July 9, 202610 min read
Featured news
-
ICML 2026 Workshop on AI for Math2026Large language models (LLMs) are good at generating code, but remain brittle for formal verification in systems like LEAN4. A core scalability challenge is that verified synthesis requires consistent outputs across multiple artifacts: executable code, precise specifications, theorem statements, and ultimately proofs. Existing approaches rarely treat these as a unified pipeline. We present BRIDGE, a structured
-
ACL 2026 Workshop on Generation, Evaluation & Metrics (GEM)2026As advanced RAG variants like GraphRAG and Agentic RAG emerge, one leading question is when and how to use them. Here, we introduce a framework for different RAG scenarios evaluation and comparison on semi-structured knowledge bases, including regular RAG, GraphRAG, Modular RAG and Agentic RAG. We provide implementation for 9 standardized RAG scenarios, and conduct experiments for a comprehensive comparison
-
Winter Simulation Conference 20262026Simulation models are widely used for decision support in manufacturing systems, yet their accuracy degrades over time as real-world operations evolve. Existing approaches to model maintenance rely heavily on manual diagnosis by domain experts, creating a bottleneck in sustaining model fidelity. This paper proposes a methodology for root cause attribution of simulation model drift using system logs and
-
2026The ability to push large objects in a goal-directed manner using onboard egocentric perception is an essential skill for humanoid robots to perform complex tasks such as material handling in warehouses. To robustly manipulate heavy objects to arbitrary goal configurations, the robot must cope with unknown object mass and ground friction, noisy onboard perception, and actuation errors; all in a real-time
-
IEEE Security & Privacy2026The Rust programming language provides stronger guarantees about memory safety than C. Therefore, translating to the Rust programming language is one way to reduce security vulnerabilities. One common translation strategy is using LLM queries. We present an alternate agentic approach that gives an LLM freedom and guardrails. When programming in C, a programmer must manually manage memory and other resources
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all