Customer-obsessed science
Research areas
-
August 26, 20265 min readDiscounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.
-
August 21, 20269 min read
-
July 30, 20268 min read
-
-
July 9, 202610 min read
Featured news
-
ICML 2025 Workshop on Foundation Models for Structured Data2026Contextual bandits offer a natural framework for sample-efficient personalization, but practical deployment remains difficult under sparse, biased interaction data, unreliable uncertainty estimates,and severe cold starts. We study whether pretrained tabular foundation models with in-context learning can be turned into randomized policies for online decision making. We propose BC-ICL (Bootstrap-conditioned
-
ICML 2026 Workshop on Failure Modes in Agentic AI2026Self-evolving skill libraries face a silent failure mode we term library drift: unbounded skill accumulation without outcome-driven lifecycle management causes retrieval degradation, false-positive injections, and performance stagnation. Recent evaluation confirms the symptom (LLM-authored skills deliver +0.0pp gain while human-curated ones deliver +16.2pp; SkillsBench (Li et al., 2026)), yet the underlying
-
Transactions on Machine Learning Research2026In many time series forecasting settings, the target time series is accompanied by exogenous covariates, such as promotions and prices in retail demand; temperature in energy load; calendar and holiday indicators for traffic or sales; and grid load or fuel costs in electricity pricing. Ignoring these exogenous signals can substantially degrade forecasting accuracy, particularly when they drive spikes, discontinuities
-
NBER-NSF 20262026Despite their strong zero-shot forecasting capabilities, Time Series Foundation Models (TSFMs) lack mechanisms for incorporating the structured domain knowledge that practitioners need for interpretability and forecast control. Dynamic Factor Models (DFMs) provide this structure, decomposing series into interpretable, adjustable factors like trend and seasonality while capturing shared dynamics across related
-
ICML 2026 Workshop on Mechanistic Interpretability, ICML 2026 Workshop on Combining Theory and Benchmark2026Video-Language Models (VidLMs) achieve strong benchmark scores, yet these scores often hide whether models use the video at all. We show that VidLM failures follow two pathways: some visual signals are never reliably encoded, while others are encoded but overridden by model priors. We introduce REVEAL, a diagnostic stress-test benchmark for quantifying when and why VidLMs under-use visual evidence. REVEAL
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all