Customer-obsessed science
Research areas
-
August 26, 20265 min readDiscounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.
-
August 21, 20269 min read
-
July 30, 20268 min read
-
-
July 9, 202610 min read
Featured news
-
ICML 2026 Workshop on Failure Modes in Agentic AI2026Self-evolving skill libraries face a silent failure mode we term library drift: unbounded skill accumulation without outcome-driven lifecycle management causes retrieval degradation, false-positive injections, and performance stagnation. Recent evaluation confirms the symptom (LLM-authored skills deliver +0.0pp gain while human-curated ones deliver +16.2pp; SkillsBench (Li et al., 2026)), yet the underlying
-
Transactions on Machine Learning Research2026In many time series forecasting settings, the target time series is accompanied by exogenous covariates, such as promotions and prices in retail demand; temperature in energy load; calendar and holiday indicators for traffic or sales; and grid load or fuel costs in electricity pricing. Ignoring these exogenous signals can substantially degrade forecasting accuracy, particularly when they drive spikes, discontinuities
-
NBER-NSF 20262026Despite their strong zero-shot forecasting capabilities, Time Series Foundation Models (TSFMs) lack mechanisms for incorporating the structured domain knowledge that practitioners need for interpretability and forecast control. Dynamic Factor Models (DFMs) provide this structure, decomposing series into interpretable, adjustable factors like trend and seasonality while capturing shared dynamics across related
-
ICML 2026 Workshop on Mechanistic Interpretability, ICML 2026 Workshop on Combining Theory and Benchmark2026Video-Language Models (VidLMs) achieve strong benchmark scores, yet these scores often hide whether models use the video at all. We show that VidLM failures follow two pathways: some visual signals are never reliably encoded, while others are encoded but overridden by model priors. We introduce REVEAL, a diagnostic stress-test benchmark for quantifying when and why VidLMs under-use visual evidence. REVEAL
-
arXiv2026Container image pulling accounts for the majority of pod startup time in Kubernetes environments. Standard pull down loads the entire image before the container can start, even when the application accesses only a fraction of the image content at startup. We present SOCI (Seekable OCI), a lazy-loading architecture that enables containers to start without downloading the full image. SOCI builds an external
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all