Customer-obsessed science
Research areas
-
August 21, 20269 min readExtendable framework enables testing agents on the full set of capabilities required to successfully complete a procedure, not isolated proxy tasks.
-
July 30, 20268 min read
-
-
July 9, 202610 min read
-
Featured news
-
ICML 2026 Workshop on Combining Theory and Benchmark2026Generative text classifiers, which assign labels by modeling the joint distribution over inputs and labels, have recently regained attention due to strong low-sample performance and a growing perception that they are less prone to shortcut learning than discriminative classifiers. However, existing evidence for shortcut avoidance is often indirect, frequently conflates classifier formulation with architectural
-
2026The rapid evolution of Large Language Models (LLMs) has driven a growing demand for automated, high-performance system kernels to accelerate machine learning workloads. We introduce TRITON RL, a domain-specialized 8B-scale LLM for Triton programming, trained via a novel reinforcement learning (RL) framework. While Triton synthesis faces unique challenges, including data scarcity and a high susceptibility
-
DARS 20262026We propose a novel algorithm for forming arbitrarily shaped assemblies using decentralized robots. By relying on local interactions, the algorithm ensures there are no unreachable states or gaps in the assembly, which are global properties. The in-assembly robots attract passing-by robots into expanding the assembly via a simple implementation of signaling and alignment. Our approach is minimalistic, requiring
-
2026Large language model (LLM)–based multi-agent systems (MAS) show strong promise for complex reasoning, planning, and tool-augmented tasks, but designing effective MAS architectures remains labor-intensive, brittle, and hard to generalize. Existing automatic MAS generation methods either rely on code generation, which often leads to executability and robustness failures, or impose rigid architectural templates
-
2026LLMs utilizing chain-of-thought reasoning often waste substantial compute by producing long, incorrect responses. Abstention can mitigate this by withholding outputs unlikely to be correct. While most abstention methods decide to withhold outputs before or after generation, dynamic mid-generation abstention considers early termination of unpromising reasoning traces at each token position. Prior work has
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all