Customer-obsessed science
Research areas
-
August 21, 20269 min readExtendable framework enables testing agents on the full set of capabilities required to successfully complete a procedure, not isolated proxy tasks.
-
July 30, 20268 min read
-
-
July 9, 202610 min read
-
Featured news
-
NeurIPS 2025 Workshop on Multimodal Algorithmic Reasoning2025Large Language Models (LLMs) perform well on short-horizon tasks but struggle with long-horizon, multimodal scenarios that require multi-step reasoning, perception, and adaptive planning. We identify two key challenges in these settings: the difficulty of long-term coordination between planning and execution within single-agent architectures and the inefficiency of indiscriminate visual grounding. To address
-
IEEE Symposium on Foundations of Computer Science (FOCS)2025We present a protocol for fault-tolerantly implementing the logical quantum random access memory (QRAM) operation, given access to a specialized, noisy QRAM device. For coherently accessing classical memories of size 2^n, our protocol consumes only poly(n) fault-tolerant quantum resources (logical gates, logical qubits, quantum error correction cycles, etc.), avoiding the need to perform active error correction
-
2025This paper investigates synthetic data generation strategies in developing generative retrieval models for domain-specific corpora, thereby addressing the scalability challenges inherent in manually annotating in-domain queries. We study the data strategies for a two-stage training framework: in the first stage, which focuses on learning to decode document identifiers from queries, we investigate LLM-generated
-
2025The additional modality (such as speech) in multimodal large language models (LLM) increases their vulnerability to adversarial jailbreak attacks. Adversarial training (AT) techniques have shown great promise as defenses in traditional adversarial robustness literature. But they are less explored as countermeasures in speech-enabled LLMs due to the limited availability of training data and computational
-
2025Composed Image Retrieval (CIR) targets the retrieval of images conditioned on a reference image and a textual modification, but constructing labeled triplets (reference image, textual modification, target image) is inherently challenging. Existing Zero-Shot CIR (ZS-CIR) approaches often rely on well-aligned vision-language models (VLMs) to combine visual and textual inputs, or use large language models
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all