Customer-obsessed science
Research areas
-
August 26, 20265 min readDiscounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.
-
August 21, 20269 min read
-
July 30, 20268 min read
-
-
July 9, 202610 min read
Featured news
-
ACM FAccT 20262026Popularity bias is a pervasive problem in recommender systems, where recommendations disproportionately favor popular items. This not only results in 'rich-get-richer' dynamics and a homogenization of visible content, but can also lead to misalignment of recommendations with individual users' preferences for popular or niche content. This work studies popularity bias through the lens of user-recommender
-
2026Building on recent formalizations of root cause analysis for rare events (“outliers”) in structural equation models, we propose a formal definition of a causal pathway and discuss its testable implications. We identify conditions under which these implications depend only on a causal abstraction defined by the pathway of rare events, rather than on the full causal graph of the underlying system. Accordingly
-
2026In many real-world systems, causal ground truth is hard to obtain, making it challenging to evaluate expert statements about causal effects. In this paper, we propose new methods for evaluating lists of bivariate causal statements in the con-text of linear structural causal models and causal graphs. Our approach is based on a mathematical formalization of mutual compatibility, measuring the extent to which
-
2026Vision language models (VLMs) often generate hallucination, i.e., content that cannot be substantiated by either textual or visual inputs. Prior work primarily attributes this to over-reliance on linguistic prior knowledge rather than visual inputs.Some methods attempt to mitigate hallucination by amplifying visual token attention proportion-ally to their attention scores. However, these methods overlook
-
AAAI 2026 Workshop on Trustworthy Agentic AI2026Large Language Models (LLMs) have shown remarkable capabilities in tool calling and tool usage, but suffer from hallucinations where they choose incorrect tools, provide malformed parameters and exhibit 'tool bypass' behavior by performing simulations and generating outputs instead of invoking specialized tools or external systems. This undermines the reliability of LLM based agents in production systems
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all