Customer-obsessed science
Research areas
-
August 26, 20265 min readDiscounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.
-
August 21, 20269 min read
-
July 30, 20268 min read
-
-
July 9, 202610 min read
Featured news
-
arXiv2026We show that the standard basis of transformer hidden states already provides a training-free, architecture-general feature basis. Individual dimensions encode semantic content via their signs (±1) and confidence via their magnitudes, functioning as independent binary registers. A feature is simply a subset of dimensions with a consistent sign pattern, readable by counting sign agreements with no learned
-
CVPR IEEE 2026 Workshop on Computer Vision in Sports2026Multi-agent trajectory generation in team sports requires models that capture both the diversity of possible plays and realistic spatial coordination between players on plays. Standard generative approaches such as Conditional Variational Autoencoders (CVAE) and diffusion models struggle with this task, exhibiting posterior collapse or convergence to the dataset mean. Moreover, most trajectory prediction
-
ICML 2026 Workshop on Statistical Frameworks for Uncertainty in Agentic Systems2026The LLM Jury, a Panel of LLM Evaluators (POLL) (Verga et al., 2024) reporting consensus scores, has become a practical alternative to single judge LLM evaluation, yet its statistical behavior remains poorly understood. Formalizing the setup under the Huber contamination model, we show that POLL incurs unbounded bias under any positive contamination, regardless of jury size, whenever a single judge fails
-
ECML-PKDD 20262026When fusing heterogeneous modalities for classification, a central challenge is cardinality heterogeneity: modalities often produce token sequences of vastly different lengths, yet standard symmetric fusion wastes attention capacity under this asymmetry. We present CRAFT, a modality-agnostic fusion framework that selects a high-density attention backbone using token cardinality and standalone task relevance
-
arXiv2026Large language model (LLM) agents deployed in healthcare and life sciences (HCLS) routinely receive queries that are semantically ambiguous—the same terms carry different meanings across clinical, regulatory, pharmacovigilance, data-standards, and research domains. Existing approaches address ambiguity post-hoc through output filtering or retrieval augmentation, but do not quantify it before the model responds
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all