Customer-obsessed science
Research areas
-
August 26, 20265 min readDiscounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.
-
August 21, 20269 min read
-
July 30, 20268 min read
-
-
July 9, 202610 min read
Featured news
-
VLDB 20262026Compilation-based query execution produces optimized machine code per query but introduces a cold-start problem: when the compiled code is not cached, the query stalls during compilation, delaying data processing by up to orders of magnitude relative to the query’s execution time. This overhead dominates short-running queries and creates latency variability for both interactive analytics and ETL pipelines
-
EACL 2026 Industry Track2026Personalized shopping agents must adapt their decisions to different user personas, balancing efficiency, preference alignment, and goal success. Building upon the WebShop dataset and τ2-Bench environment, ShopperBench introduces a persona-guided benchmark for evaluating such adaptive behaviors. ShopperBench augments shopping trajectories with persona-conditioned goals, reasoning rationales, and preference
-
2026We introduce SWAN (Semantic Watermarking with Abstract Meaning Representation)1 , a novel framework that embeds watermark signatures into the semantic structure of a sentence using Abstract Meaning Representation (AMR). In contrast to existing watermarking methods, which typically encode signatures by adjusting token selection preferences during text generation, SWAN embeds the signature directly in the
-
2026Individual treatment effect (ITE) estimation from observational data becomes unreliable when three challenges co-occur: extreme class imbalance (0.4% treatment rate), outcome sparsity (97.6% zeros), and pervasive cold-start (99.2% incomplete profiles). These conditions violate identifying assumptions—propensity scores collapse toward boundary values, and outcome predictions degrade for subjects with sparse
-
2026Structured texts refer to texts containing structured elements beyond plain texts, such as code snippets and placeholders. Such structured texts increasingly require segmentation into semantically meaningful components, which cannot be effectively handled by conventional sentence-level segmentation methods. To address this, we propose BoundRL, a novel approach that jointly performs efficient token-level
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all