IP Library › Granted Patent US 12,657,618
Granted Patent B2
US 12,657,618 · App. 18/182,944 · Granted Jun 16, 2026

Systems and methods for universal item learning in item recommendation

Inventors: Ziwei Fan (Chicago, IL); Yongjun Chen (Palo Alto, CA); Zhiwei Liu (Palo Alto, CA); Huan Wang (Palo Alto, CA)
Assignee: Salesforce, Inc.
G06Q30/0631G06Q30/0201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,657,618
App. No.
18/182,944
Granted
Jun 16, 2026
Kind
B2
Abstract

Embodiments described herein provide a universal item learning framework that generates universal item embeddings for zero-shot items. Specifically, the universal item learning framework performs generic features extraction of items and product knowledge characterization based on a product knowledge graph (PKG) to generate embeddings of input items. A pretrained language model (PLM) may be adopted to extract features from generic item side information, such as titles, descriptions, etc., of an item. A PKG may be constructed to represent recommendation-oriented knowledge, which comprise a plurality of nodes representing items and a plurality of edges connecting nodes represent different relations between items. As those relations in PKG are usually retrieved from user-item interactions, the PKG adapts the universal representation for recommendation with knowledge of user-item interactions.

Claims (87)

1 . A system for pretraining a multi-task model to generate universal item embeddings, the system comprising:

a data interface that receives information relating to a plurality of items and user-item interactions;

a memory storing a product knowledge graph representing item-item relations derived from the user-item interactions, and a plurality of processor-executable instructions; and

one or more processors executing the instructions to perform operations including:

encoding, by a graph encoder, at least a portion of the product knowledge graph corresponding to the plurality of items into a plurality of item relational embeddings;

generating, by a respective task-oriented adaptation layer, a respective pretraining output based on the plurality of item relational embeddings;

computing a respective pretraining objective based on the respective pretraining output and the at least portion of the product knowledge graph, wherein the respective pretraining objective is computed by computing a knowledge reconstruction score based on item-relational embeddings corresponding to a triplet of a first item, a second item and a specific relation between the first item and the second item and computing a cross-entropy loss based on knowledge reconstruction scores computed from positive triplets and negative triplets;

updating at least the graph encoder based on multiple pretraining objectives via backpropagation;

generating, by the updated graph encoder and at least one task-oriented adaptation layer, predicted ranking scores between the plurality of items and a set of users based on the product knowledge graph;

computing a Bayesian ranking loss based on the ranking scores; and

finetuning the at least one task-oriented adaptation layer based on the Bayesian ranking loss while keeping the updated graph encoder frozen.

2 . The system of claim 1 , wherein the product knowledge graph comprises a plurality of nodes representing the plurality of items, and a plurality of edges connecting the plurality of nodes and representing item-item connections among the plurality of items, and

wherein the product knowledge graph is constructed based on the information relating to the plurality of items and user-item interactions by:

extracting, by a pretrained language model, item feature embeddings from the information relating to the plurality of items; and

deriving, by the pretrained language model, the item-item connections from collected feedback from the user-item interactions relating to the plurality of items.

3 . The system of claim 1 , wherein the operation of generating, by the respective task-oriented adaptation layer, the respective pretraining output comprises:

adapting the plurality of item relational embeddings to a respective task; and

fusing the adapted plurality of item relational embeddings.

4 . The system of claim 1 , wherein the respective pretraining objective is computed by:

concatenating, for each item, item-relational embeddings corresponding to the respective item, into a respective item embedding;

computing a neighbor reconstruction score between a first item and a second item based on a first item embedding and a second item embedding; and

computing a cross-entropy loss based on neighbor reconstruction scores between pairs of items that are within a pre-defined number hops from each other.

5 . The system of claim 1 , wherein the respective pretraining objective is computed by:

concatenating the plurality of item-relational embeddings into a concatenated relational embedding;

generating, by a decoder, a decoded feature from the concatenated relational embedding; and

computing a feature reconstruction loss based on a distance between the decoded feature and original encoded item features.

6 . The system of claim 1 , wherein the respective pretraining objective is computed by:

computing, for each respective item, a weighted sum of the plurality of item-relational embeddings corresponding to the respective item into a respective item embedding;

computing a prediction score between a first item embedding corresponding to a first item and a second item embedding corresponding to a second item; and

computing a cross-entropy loss based on first prediction scores between pairs of items that are connected according to a specific relation and second prediction scores between pairs of items that are not connected according to the specific relation.

7 . The system of claim 1 , wherein the graph encoder is updated based on a weighted sum of the multiple pretraining objectives.

8 . The system of claim 1 , wherein the predicted ranking scores are generated by:

computing, for a first item, a weighted sum of the plurality of item-relational embeddings corresponding to the first item into a first item embedding;

computing, for a first user, a mean aggregation of all iterated items by averaging item embeddings corresponding to items that the first user has interacted with; and

computing the predicted ranking score between the first user and the first item based on a similarity between the mean aggregation and the first item embedding.

9 . The system of claim 1 , wherein the Bayesian ranking loss is computed based a first predicted ranking score corresponding to a first user and a first item that interacted with each other, and a second predicted ranking score corresponding to a first user and a second item that do not interact with each other.

10 . The system of claim 1 , wherein the operations further comprise:

updating the product knowledge graph with a new item and a set of relations between the new item and the plurality of items;

generating, by the updated graph encoder and the finetuned at least one task-oriented adaptation layer, an item embedding for the new item,

wherein the item embedding comprises knowledge for a recommendation task deciding whether to recommend the new item for a specific user.

11 . A method for pretraining a multi-task model to generate universal item embeddings, the method comprising:

receiving, via a data interface, information relating to a plurality of items and user-item interactions;

obtaining a product knowledge graph representing item-item relations derived from the user-item interactions;

encoding, by a graph encoder, at least a portion of the product knowledge graph corresponding to the plurality of items into a plurality of item relational embeddings;

generating, by a respective task-oriented adaptation layer, a respective pretraining output based on the plurality of item relational embeddings;

computing a respective pretraining objective based on the respective pretraining output and the at least portion of the product knowledge graph, wherein the respective pretraining objective is computed by computing a knowledge reconstruction score based on item-relational embeddings corresponding to a triplet of a first item, a second item and a specific relation between the first item and the second item and computing a cross-entropy loss based on knowledge reconstruction scores computed from positive triplets and negative triplets;

updating at least the graph encoder based on multiple pretraining objectives via backpropagation;

generating, by the updated graph encoder and at least one task-oriented adaptation layer, predicted ranking scores between the plurality of items and a set of users based on the product knowledge graph;

computing a Bayesian ranking loss based on the ranking scores; and

finetuning the at least one task-oriented adaptation layer based on the Bayesian ranking loss while keeping the updated graph encoder frozen.

12 . The method of claim 11 , wherein the product knowledge graph comprises a plurality of nodes representing the plurality of items, and a plurality of edges connecting the plurality of nodes and representing item-item connections among the plurality of items, and

wherein the product knowledge graph is constructed based on the information relating to the plurality of items and user-item interactions by:

extracting, by a pretrained language model, item feature embeddings from the information relating to the plurality of items; and

deriving, by the pretrained language model, the item-item connections from collected feedback from the user-item interactions relating to the plurality of items.

13 . The method of claim 11 , wherein the operation of generating, by the respective task-oriented adaptation layer, the respective pretraining output comprises:

adapting the plurality of item relational embeddings to a respective task; and

fusing the adapted plurality of item relational embeddings.

14 . The method of claim 11 , wherein the respective pretraining objective is computed by:

concatenating, for each item, item-relational embeddings corresponding to the respective item, into a respective item embedding;

computing a neighbor reconstruction score between a first item and a second item based on a first item embedding and a second item embedding; and

computing a cross-entropy loss based on neighbor reconstruction scores between pairs of items that are within a pre-defined number hops from each other.

15 . The method of claim 11 , wherein the respective pretraining objective is computed by:

concatenating the plurality of item-relational embeddings into a concatenated relational embedding;

generating, by a decoder, a decoded feature from the concatenated relational embedding; and

computing a feature reconstruction loss based on a distance between the decoded feature and original encoded item features.

16 . The method of claim 11 , wherein the respective pretraining objective is computed by:

computing, for each respective item, a weighted sum of the plurality of item-relational embeddings corresponding to the respective item into a respective item embedding;

computing a prediction score between a first item embedding corresponding to a first item and a second item embedding corresponding to a second item; and

computing a cross-entropy loss based on first prediction scores between pairs of items that are connected according to a specific relation and second prediction scores between pairs of items that are not connected according to the specific relation.

17 . The method of claim 11 , wherein the graph encoder is updated based on a weighted sum of the multiple pretraining objectives.

18 . A non-transitory machine-readable medium comprising a plurality of machine-executable instructions which, when executed by one or more processors, are adapted to cause the one or more processors to perform operations comprising:

receiving, via a data interface, information relating to a plurality of items and user-item interactions;

obtaining a product knowledge graph representing item-item relations derived from the user-item interactions;

encoding, by a graph encoder, at least a portion of the product knowledge graph corresponding to the plurality of items into a plurality of item relational embeddings;

generating, by a respective task-oriented adaptation layer, a respective pretraining output based on the plurality of item relational embeddings;

computing a respective pretraining objective based on the respective pretraining output and the at least portion of the product knowledge graph, wherein the respective pretraining objective is computed by computing a knowledge reconstruction score based on item-relational embeddings corresponding to a triplet of a first item, a second item and a specific relation between the first item and the second item and computing a cross-entropy loss based on knowledge reconstruction scores computed from positive triplets and negative triplets;

updating at least the graph encoder based on multiple pretraining objectives via backpropagation;

generating, by the updated graph encoder and at least one task-oriented adaptation layer, predicted ranking scores between the plurality of items and a set of users based on the product knowledge graph;

computing a Bayesian ranking loss based on the ranking scores; and

finetuning the at least one task-oriented adaptation layer based on the Bayesian ranking loss while keeping the updated graph encoder frozen.

19 . The non-transitory machine-readable medium of claim 18 , wherein the product knowledge graph comprises a plurality of nodes representing the plurality of items, and a plurality of edges connecting the plurality of nodes and representing item-item connections among the plurality of items, and

wherein the product knowledge graph is constructed based on the information relating to the plurality of items and user-item interactions by:

extracting, by a pretrained language model, item feature embeddings from the information relating to the plurality of items; and

deriving, by the pretrained language model, the item-item connections from collected feedback from the user-item interactions relating to the plurality of items.

20 . The non-transitory machine-readable medium of claim 18 , wherein the operation of generating, by the respective task-oriented adaptation layer, the respective pretraining output comprises:

adapting the plurality of item relational embeddings to a respective task; and

fusing the adapted plurality of item relational embeddings.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 17, 2023
From: FAN, ZIWEI; CHEN, YONGJUN; LIU, ZHIWEI; WANG, HUAN
To: SALESFORCE, INC.
Reel/Frame 063347/0178 →
Continuity (3)
Provisional Application 63481372 · Jan 24, 2023
Provisional Application 63395709 · Aug 5, 2022
Related Publication 20240046330A1 · Feb 8, 2024
References Cited (8)
US 20210027178A1 · Ding · 2021 [cited by examiner]
US 20210233124A1 · Rahman · 2021 [cited by examiner]
US 20210241343A1 · Arora · 2021 [cited by examiner]
US 20220207587A1 · Yang · 2022 [cited by examiner]
US 20230206076A1 · Xu · 2023 [cited by examiner]
US 20230229859A1 · Khandelwal · 2023 [cited by examiner]
US 20230267317A1 · Shin · 2023 [cited by examiner]
US 20240289823A1 · Wu · 2024 [cited by examiner]