Building commonsense knowledge graphs to aid product recommendation

Using large language models to discern commonsense relationships can improve performance on downstream tasks by as much as 60%.

May 10, 2024

At the Amazon Store, we strive to deliver the product recommendations most relevant to customers’ queries. Often, that can require commonsense reasoning. If a customer, for instance, submits a query for “shoes for pregnant women”, the recommendation engine should be able to deduce that pregnant women might want slip-resistant shoes.

At left is a flow chart that begins with the query "Want shoes for pregnant women", with an arrow connecting it to the action "Bought a slip-resistant shoe", which is in turn connected to the commonsense triple <pregnant, require, slip-resistant>. At right are a selection of product pages for slip-resistant shoes. — Mining implicit commonsense knowledge from customer behavior.

To help Amazon’s recommendation engine make these types of commonsense inferences, we’re building a knowledge graph that encodes relationships between products in the Amazon Store and the human contexts in which they play a role — their functions, their audiences, the locations in which they’re used, and the like. For instance, the knowledge graph might use the used_for_audience relationship to link slip-resistant shoes and pregnant women.

In a paper we’re presenting at the Association for Computing Machinery’s annual Conference on Management of Data (SIGMOD) in June 2024, we describe COSMO, a framework that uses large language models (LLMs) to discern the commonsense relationships implicit in customer interaction data from the Amazon Store.

COSMO involves a recursive procedure in which an LLM generates hypotheses about the commonsense implications of query-purchase and co-purchase data; a combination of human annotation and machine learning models filters out the low-quality hypotheses; human reviewers extract guiding principles from the high-quality hypotheses; and instructions based on those principles are used to prompt the LLM.

A cyclical flow chart that begins in the upper left with "user behavior", featuring icons that represent search, product views, ratings, and purchases. A right arrow labeled "prompt" connects the user behavior to a neural-network icon labeled "LLMs". A right arrow labeled "generate" connects the LLMs to a stacked-papers icon representing "knowledge". A downward arrow labeled "filter" connects "knowledge" to a box containing the words "rule-based filtering" and "similarity filtering". A left arrow labeled "annotate" connects the filtering box to a box labeled "Human feedback". A final left arrow connects "human feedback" to a box labeled "Instructions", which contains an example instructing the LLM to use the "capableOf" relation to explain the connection between the query "winter coat" and the product "long-sleeve puffer coat". The LLM's output is "Provide high-level warmth". — The COSMO framework.

To evaluate COSMO, we used the Shopping Queries Data Set we created for KDD Cup 2022, a competition held at the 2022 Conference on Knowledge Discovery and Data Mining (KDD). The dataset consists of queries and product listings, with the products rated according to their relevance to each query.

In our experiments, three models — a bi-encoder, or two-tower model; a cross-encoder, or unified model; and a cross-encoder enhanced with relationship information from the COSMO knowledge graph — were tasked with finding the products most relevant to each query. We measured performance using two different F1 scores: macro F1 is an average of F1 scores in different categories, and micro F1 is the overall F1 score, regardless of categories.

When the models’ encoders were fixed — so the only difference between the cross-encoders was that one included COSMO relationships as inputs and the other didn’t — the COSMO-based model dramatically outperformed the best-performing baseline, achieving a 60% increase in macro F1 score. When the encoders were fine-tuned on a subset of the test dataset, the performance of all three models improved significantly, but the COSMO-based model still held a 28% edge in macro F1 and a 22% edge in micro F1 over the best-performing baseline.

COSMO

COSMO’s knowledge graph construction procedure begins with two types of data: query-purchase pairs, which combine queries with purchases made within a fixed span of time or a fixed number of clicks, and co-purchase pairs, which combine purchases made during the same shopping session. We do some initial pruning of the dataset to mitigate noise — for instance, removing co-purchase pairs in which the product categories of the purchased products are too far apart in the Amazon product graph.

Evaluation and application

The bi-encoder model we used in our experiments had two separate encoders, one for a customer query and one for a product. The outputs of the two encoders were concatenated and fed to a neural-network module that produced a relevance score.

In the cross-encoder, all the relevant features of both the query and the product description pass to the same encoder. In general, cross-encoders work better than bi-encoders, so that’s the architecture we used to test the efficacy of COSMO data.

Building commonsense knowledge graphs to aid product recommendation

Using large language models to discern commonsense relationships can improve performance on downstream tasks by as much as 60%.

COSMO

Evaluation and application

Related content

Work with us