IP Library › Granted Patent US 12,450,314
Granted Patent B2
US 12,450,314 · App. 17/586,451 · Granted Oct 21, 2025

Systems and methods for sequential recommendation

Inventors: Yongjun Chen (Palo Alto, CA); Zhiwei Liu (Chicago, IL); Jia Li (Mountain View, CA); Caiming Xiong (Menlo Park, CA)
Assignee: Salesforce, Inc.
G06F18/24137G06F18/23213G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,314
App. No.
17/586,451
Granted
Oct 21, 2025
Kind
B2
Abstract

Embodiments described herein provides an intent prototypical contrastive learning framework that leverages intent similarities between users with different behavior sequences. Specifically, user behavior sequences are encoded into a plurality of user interest representations. The user interest representations are clustered into a plurality of clusters based on mutual distances among the user interest representations in a representation space. Intention prototypes are determined based on centroids of the clusters. A set of augmented views for user behavior sequences are created and encoded into a set of view representations. A contrastive loss is determined based on the set of augmented views and the plurality of intention prototypes. Model parameters are updated based at least in part on the contrastive loss.

Claims (53)

1. A method for sequential recommendation based on user intent modeling, the method comprising:

receiving a plurality of user behavior sequences;

encoding, via an encoder, the plurality of user behavior sequences into a plurality of user interest representations;

clustering the plurality of user interest representations into a plurality of clusters based on mutual distances among the user interest representations in a representation space;

determining a plurality of intention prototypes based on centroids of the plurality of clusters;

constructing a set of augmented views for a first user behavior sequence from the plurality of user behavior sequences;

encoding, via the encoder, the set of augmented views into a set of view representations;

computing a contrastive loss based on a summation of user contrastive losses corresponding to a number of users, wherein each user contrastive loss is computed based on a first similarity between a first positive view representation and an intention prototype of the plurality of intention prototypes corresponding to a respective user, and a plurality of similarities between the first positive view representation and a set of intention prototypes of the plurality of intention prototypes that do not correspond to the respective user; and

iteratively updating the encoder to minimize the contrastive loss alone or in weighted combination with an additional loss component.

2. The method of claim 1 , wherein the clustering the plurality of user interest representations is performed based on K-means clustering.

3. The method of claim 1 , wherein the set of intention prototypes are different from the intention prototype corresponding to the respective user.

4. The method of claim 1 , further comprising:

computing a next item loss based on a summation of user next item losses corresponding to a number of users,

wherein each user next item loss is computed based on a similarity between a first positive view representation and an embedding of a target item, and a plurality of similarities between the first positive representation view and a set of embeddings of user behavior sequences that do not correspond to the respective user.

5. The method of claim 1 , further comprising:

computing a sequential contrastive loss based on a summation of user sequential contrastive losses corresponding to a number of users,

wherein each user sequential contrastive loss is computed based on a similarity between a first positive view representation and a second positive view representation, and a plurality of similarities between the first positive view representation and a set of negative view representations that do not correspond to the respective user.

6. The method of claim 1 , further comprising:

computing a weighted sum of the contrastive loss, a next item loss, and a sequential contrastive loss; and

jointly updating the encoder based on the weighted sum.

7. The method of claim 1 , wherein at least one user behavior sequence from the plurality of user behavior sequences includes information about a sequence of items that a user has interacted with.

8. A system for sequential recommendation based on user intent modeling, the system comprising:

a memory that stores a sequential recommendation model;

a communication interface that receives a plurality of user behavior sequences; and

one or more hardware processors that:

encodes, via an encoder, the plurality of user behavior sequences into a plurality of user interest representations;

clusters the plurality of user interest representations into a plurality of clusters based on mutual distances among the user interest representations in a representation space;

determines a plurality of intention prototypes based on centroids of the plurality of clusters;

constructs a set of augmented views for a first user behavior sequence from the plurality of user behavior sequences;

encodes, via the encoder, the set of augmented views into a set of view representations;

computes a contrastive loss based on a summation of user contrastive losses corresponding to a number of users, wherein each user contrastive loss is computed based on a first similarity between a first positive view representation and an intention prototype of the plurality of intention prototypes corresponding to a respective user, and a plurality of similarities between the first positive view representation and a set of intention prototypes of the plurality of intention prototypes that do not correspond to the respective user; and

iteratively updates the encoder to minimize the contrastive loss alone or in weighted combination with an additional loss component.

9. The system of claim 8 , wherein the clustering the plurality of user interest representations is performed based on K-means clustering.

10. The system of claim 8 , wherein the set of intention prototypes are different from the intention prototype corresponding to the respective user.

11. The system of claim 8 , wherein the one or more hardware processors further:

computes a next item loss based on a summation of user next item losses corresponding to a number of users,

wherein each user next item loss is computed based on a similarity between a first positive view representation and an embedding of a target item, and a plurality of similarities between the first positive view representation and a set of embeddings of user behavior sequences that do not correspond to the respective user.

12. The system of claim 8 , wherein the one or more hardware processors further:

computes a sequential contrastive loss based on a summation of user sequential contrastive losses corresponding to a number of users,

wherein each user sequential contrastive loss is computed based on a similarity between a first positive view representation and a second positive view representation, and a plurality of similarities between the first positive view representation and a set of negative view representations that do not correspond to the respective user.

13. The system of claim 8 , wherein the one or more hardware processors further:

computes a weighted sum of the contrastive loss, a next item loss, and a sequential contrastive loss; and

jointly updates the encoder based on the weighted sum.

14. The system of claim 8 , wherein at least one user behavior sequence from the plurality of user behavior sequences includes information about a sequence of items that a user has interacted with.

15. A processor-readable non-transitory storage medium storing a plurality of processor-executable instructions for sequential recommendation based on user intent modeling, the instructions being executed by a processor to perform operations comprising:

receiving a plurality of user behavior sequences;

encoding, via an encoder, the plurality of user behavior sequences into a plurality of user interest representations;

clustering the plurality of user interest representations into a plurality of clusters based on mutual distances among the user interest representations in a representation space;

determining a plurality of intention prototypes based on centroids of the plurality of clusters;

constructing a set of augmented views for a first user behavior sequence from the plurality of user behavior sequences;

encoding, via the encoder, the set of augmented views into a set of view representations;

computing a contrastive loss based on a summation of user contrastive losses corresponding to a number of users, wherein each user contrastive loss is computed based on a first similarity between a first positive view representation and an intention prototype of the plurality of intention prototypes corresponding to a respective user, and a plurality of similarities between the first positive view representation and a set of intention prototypes of the plurality of intention prototypes that do not correspond to the respective user; and

iteratively updating the encoder to minimize the contrastive loss alone or in weighted combination with an additional loss component.

Assignments (2)
CHANGE OF NAME Recorded Aug 4, 2026
From: SALESFORCE.COM, INC.
To: SALESFORCE, INC.
Reel/Frame 076118/0548 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2022
From: CHEN, YONGJUN; LIU, ZHIWEI; LI, JIA; XIONG, CAIMING
To: SALESFORCE.COM, INC.
Reel/Frame 059405/0188 →
Continuity (2)
Provisional Application 63233164 · Aug 13, 2021
Related Publication 20230073754A1 · Mar 9, 2023
References Cited (16)
US 20100312726A1 · Thompson · 2010 [cited by examiner]
US 20180032606A1 · Tolman · 2018 [cited by examiner]
US 20220019888A1 · Aggarwal · 2022 [cited by examiner]
US 20220237682A1 · Zhao · 2022 [cited by examiner]
Xie et al. (Xie) “Contrastive Learning for Sequential Recommendation” arXiv preprint arXiv:2010.14395 (2020). 11 pages (this is the same as the applicant provided NPL) (Year: 2020). [cited by examiner]
Renqin Cai et al., Category-aware Collaborative Sequential Recommendation. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. 388-397 .(2021). [cited by applicant]
Md Mehrab Tanjim et al., sequential models of latent intent for next item recommendation. In Proceedings of The Web Conference 2020. 2528-2534. (2020). [cited by applicant]
Jianxin Ma et al., Disentangled self-supervision in sequential recommenders. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 483-491. (2020). [cited by applicant]
Shoujin Wang et al., Modeling multi-purpose sessions for next-item recommendations via mixture-channel purpose routing networks. In International Joint Conference on Artificial Intelligence. International Joint Conferen… [cited by applicant]
Zhiwei Liu et al., . Basket recommendation with multi-intent translation graph neural network. In 2020 IEEE International Conference on Big Data (Big Data). IEEE, 728-737. (2020). [cited by applicant]
Zhiqiang Pan et al., An intentguided collaborative machine for session-based recommendation. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval. 1833-1836.… [cited by applicant]
Jiancan Wun et al., Self-supervised graph learning for recommendation. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. 726-735.(2021). [cited by applicant]
Xu Xie et al., Contrastive Learning for Sequential Recommendation. arXiv preprint arXiv:2010.14395 (2020). [cited by applicant]
Tiansheng Yao et al., Self-supervised Learning for Large-scale Item Recommendations. arXiv preprint arXiv:2007.12865 (2020). [cited by applicant]
Kun Zhou et al., S3-rec: Self-supervised learning for sequential recommendation with mutual information maximization. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management. 1893-1… [cited by applicant]
Xiao Liu et al., Self-supervised learning: Generative or contrastive. IEEE Transactions on Knowledge and Data Engineering (2021). [cited by applicant]