Customer-obsessed science
Research areas
-
October 1, 202610 min readAugmenting a network graph with agentic AI produces a “digital twin” that can help isolate network failures.
-
-
August 21, 20269 min read
-
July 30, 20268 min read
-
July 29, 20266 min read
Featured news
-
2026Memory makes LLM-based web agents personalized, powerful, yet exploitable. By storing past interactions to personalize future tasks, agents inadvertently create a persistent attack surface that spans websites and sessions. While existing security research on memory assumes attackers can directly inject into memory storage or exploit shared memory across users, we present a more realistic threat model: contamination
-
2026LLM judges are increasingly used to evaluate open-ended responses, but their scores depend strongly on the rubrics that condition them. A vague rubric asking for a response to be “helpful and factual” can reward polished answers that invent facts or violate user intent. We treat reusable rubrics as measurement specifications: changing the rubric changes the response quality measurement induced by a fixed
-
2026Music recommendations rely on robust artist embeddings that capture stylistic and behavioral similarity. A common approach is to learn track-level embeddings from customer co-listening data and aggregate them per artist, but simple heuristic aggregation (e.g., averaging) treats all tracks equally and may fail to capture which tracks best define an artist’s identity. We introduce MusicContainerNet (MCN),
-
2026Music recommendation surfaces need human-readable explanations, but LLM-quality generation does not scale on the serving path. Using MyMix, an algorithmically-generated personalized playlist surface, as a testbed, we start from a whole-playlist 4B baseline and identify two limitations: feasibility (per-request inference does not scale) and input granularity (the model sees only coarse aggregated tags, not
-
2026Organizing unstructured feedback text into hierarchical taxonomy is a fundamental challenge in NLP, particularly in domains where feedback arrives at massive scale in varied forms such as reviews, transcripts, and surveys. Existing approaches either produce shallow hierarchies, neglect long-tail topics, or lack rigorous evaluation frameworks. We present TaxCE, a fully automated framework that constructs
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all