IP Library Granted Patent US 12,423,381
Granted Patent B2
US 12,423,381 · App. 17/543,065 · Granted Sep 23, 2025

Model with usage data compensation

Inventors: Oren Barkan (Tel-Aviv, IL); Roy Hirsch (Tel Aviv, IL); Ori Katz (Tel-Aviv, IL); Avi Caciularu (Tel-Aviv, IL); Yonathan Weill (Tel-Aviv, IL); Noam Koenigstein (Tel Aviv, IL); Nir Nice (Salit, IL)
Assignee: Microsoft Technology Licensing, LLC
G06F18/2115G06F18/2148G06F18/217G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,423,381
App. No.
17/543,065
Granted
Sep 23, 2025
Kind
B2
Abstract

A method of training a machine learning model is provided. The method includes receiving labeled training data in the machine learning model, the received labeled training data including content data for items accessible to a user and input usage data representing recorded interaction between the user and the items, wherein the received content data for each item includes data representing intrinsic attributes of the item. The method further includes selecting a set of the input usage data that excludes input usage data for a proper subset of the items and training the machine learning model based on both the content data and the selected set of input usage data of the received labeled training data for the items.

Claims (46)

1. A method of training a machine learning model, the method comprising:

receiving training data in the machine learning model, the received training data including content data for one or more items accessible to one or more users and input usage data labels representing recorded interaction between each user and each item, wherein the content data for each item includes data representing intrinsic attributes of the item;

generating modified training data by including the content data and the input usage data labels for a second proper subset of items and the input usage data labels for the first proper subset of the items and by excluding the input usage data labels for the first proper subset of the items, the second proper subset of the items corresponding to the items not included in the first proper subset;

simulating, by a usage data simulator of the machine learning model and for the first proper subset of items, simulated usage data labels, based on the content data for the first proper subset of the items;

adding the simulated usage data labels to the modified training data; and

training the machine learning model using the modified training data to predict input usage data of an input item based on input content data of the input item.

2. The method of claim 1 , wherein the operation of training further trains the machine learning model based on the simulated usage data labels for the first proper subset of the items.

3. The method of claim 1 , wherein the operation of generating excludes the input usage data labels for the first proper subset of the items based on a random variable.

4. The method of claim 3 , wherein the random variable is based on a modifiable popularity bias compensation parameter.

5. The method of claim 1 , further comprising:

generating, by an aggregate content analyzer of the machine learning model, an aggregated content data representation based on a plurality of content elements of the content data, wherein the operation of training is based on the aggregated content data representation.

6. The method of claim 1 , wherein the operation of training further comprises:

determining a loss between an excluded usage data label of the training data for a given item of the first proper subset of the items and a predicted usage data label output by the machine learning model for the given item of the first proper subset of the items; and

modifying the modified training data based on the determined loss.

7. The method of claim 1 , wherein the training data further includes user data that identifies the user.

8. A computing device having a processor and memory, the processor configured to execute instructions stored in the memory, the computing device comprising:

a communication interface operable to receive training data in a machine learning model, the received training data including content data for items accessible to a user and input usage data labels representing recorded interaction between the user and the items, wherein the content data for each item includes data representing intrinsic attributes of the item;

a selector executable by the processor and operable to generate modified training data by excluding, from the training data, the input usage data labels for a first proper subset of the items of the received training data, the modified training data including the input usage data labels for a second proper subset of the items and the content data for the first proper subset of the items and the second proper subset of the items, the second proper subset of the items corresponding to the items not included in the first proper subset;

a usage data simulator of the machine learning model executable by the processor and operable to:

simulate, for the first proper subset of the items, simulated usage data labels, based on the content data for the first proper subset of the items; and

add the simulated usage data labels to the modified training data; and

a model tuner executable by the processor and operable to train, using the modified training data, the machine learning model to predict input usage data of input items based on content data of the input items.

9. The computing device of claim 8 , wherein the model tuner trains the machine learning model further based on the simulated usage data labels for the first proper subset of the items.

10. The computing device of claim 9 , wherein the selector excludes the input usage data labels for the first proper subset of the items based on a random variable.

11. The computing device of claim 10 , wherein the random variable is based on a modifiable popularity bias compensation parameter.

12. The computing device of claim 8 , further comprising:

an aggregate content analyzer of the machine learning model executable by the processor and operable to generate an aggregated content data representation based on a plurality of content elements of the content data, wherein the model tuner trains further based on the aggregated content data representation.

13. The computing device of claim 8 , wherein the model tuner is operable to:

determine a loss between an excluded usage data label of the training data for a given item of the first proper subset of the items and a predicted usage data label output by the machine learning model for the given item of the first proper subset of the items; and

modify the modified training data based on the determined loss.

14. The computing device of claim 8 , wherein the training data further includes user data that identifies the user.

15. One or more tangible processor-readable storage media embodied with instructions for executing on one or more processors of a computing device a process for training a machine learning model, the process comprising:

receiving training data in the machine learning model, the received training data including content data for items accessible to a user and input usage data labels representing recorded interaction between the user and the items, wherein the content data for each item includes data representing intrinsic attributes of the item;

generating modified training data by excluding, from the training data, the input usage data labels for a first proper subset of the items of the training data, the modified training data including the input usage data labels for a second proper subset of the items and the content data for the first proper subset of the items and the second proper subset of the items, the second proper subset of the items corresponding to the items not included in the first proper subset;

simulating, by a usage data simulator of the machine learning model and for the first proper subset of the items, simulated usage data labels, based on the content data for the first proper subset of the items;

adding the simulated usage data labels to the modified training data; and

training the machine learning model to predict input usage data of input items based on content data of the input items, the training operation using the modified training data.

16. The one or more tangible processor-readable storage media of claim 15 ,

wherein the operation of training further trains the machine learning model based on the simulated usage data labels for the first proper subset of the items.

17. The one or more tangible processor-readable storage media of claim 16 , wherein the operation of generating excludes the input usage data labels for the first proper subset of the items based on a random variable.

18. The one or more tangible processor-readable storage media of claim 17 , wherein the random variable is based on a modifiable popularity bias compensation parameter.

19. The one or more tangible processor-readable storage media of claim 15 , the process further comprising:

generating, by an aggregate content analyzer of the machine learning model, an aggregated content data representation based on a plurality of content elements of the content data, wherein the operation of training is based on the aggregated content data representation.

20. The one or more tangible processor-readable storage media of claim 15 , wherein the training further comprises:

determining a loss between an excluded usage data label of the training data for a given item of the first proper subset of the items and a predicted usage data label output by the machine learning model for the given item of the first proper subset of the items; and

modifying the modified training data based on the determined loss.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 13, 2021
From: BARKAN, OREN; HIRSCH, ROY; KATZ, ORI; CACIULARU, AVI; WEILL, YONATHAN; KOENIGSTEIN, NOAM; NICE, NIR
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 058371/0263 →
Continuity (1)
Related Publication 20230177111A1 · Jun 8, 2023
References Cited (57)
US 10945012B2 · Schneck et al. · 2021 [cited by applicant]
US 20140181121A1 · Nice et al. · 2014 [cited by applicant]
US 20180218428A1 · Xie et al. · 2018 [cited by applicant]
US 20190043493A1 · Mohajer et al. · 2019 [cited by applicant]
US 20210004021A1 · Zhang · 2021 [cited by examiner]
US 20210165848A1 · Sror et al. · 2021 [cited by applicant]
US 20210201208A1 · Bhole · 2021 [cited by examiner]
US 20220172426A1 · Lissi · 2022 [cited by examiner]
US 20220180186A1 · Basilico · 2022 [cited by examiner]
CN 111310028A · 2020 [cited by applicant]
CN 112632397A · 2021 [cited by applicant]
WO WO2019022840A1 · 2019 [cited by examiner]
“Cold Start Revisited: A Deep Hybrid Recommender with Cold-Warm Item Harmonization”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing {ICASSP), Jun. 6, 2021, pp. 3260-3264.) (Ye… [cited by examiner]
Barkan, et al., “Cold Start Revisited: A Deep Hybrid Recommender with Cold-Warm Item Harmonization”, In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Jun. 6, 2021, pp.… [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US22/044307”, Mailed Date: Jan. 5, 2023, 16 Pages. [cited by applicant]
Duricic, et al., “Trust-Based Collaborative Filtering: Tackling the Cold Start Problem Using Regular Equivalence”, In Proceedings of the 12th ACM Conference on Recommender Systems, Oct. 2, 2018, pp. 446-450. [cited by applicant]
Barkan, et al., “CB2CF: A Neural Multiview Content-to-Collaborative Filtering Model for Completely Cold Item Recommendations”, In Proceedings of the 13th ACM Conference on Recommender System, Sep. 16, 2019, pp. 228-236. [cited by applicant]
Bennett, et al., “The Netflix Prize”, In Proceedings of KDD Cup and Workshop , vol. 2007, Aug. 12, 2007, 4 Pages. [cited by applicant]
Blei, et al., “Latent Dirichlet Allocation”, In Journal of Machine Learning Research, vol. 3, Mar. 1, 2003, pp. 993-1022. [cited by applicant]
Braunhofer, Matthias, “Hybridisation Techniques for Cold-Starting Context-Aware Recommender Systems”, In Proceedings of the 8th ACM Conference on Recommender Systems, Oct. 6, 2014, pp. 405-408. [cited by applicant]
Burke, Robin, “Hybrid Recommender Systems: Survey and Experiments”, In Journal of User Modeling and User-Adapted Interaction, vol. 12, Issue 4, Nov. 2002, pp. 331-370. [cited by applicant]
Barkan, et al., “Item2vec: Neural Item Embedding for Collaborative Filtering”, In Proceedings of the IEEE 26th International Workshop on Machine Learning for Signal Processing, Sep. 13, 2016, 6 Pages. [cited by applicant]
Dacrema, et al., “Are We Really Making Much Progress? A Worrying Analysis of Recent Neural Recommendation Approaches”, In Proceedings of the 13th ACM Conference on Recommender Systems, Sep. 16, 2019, pp. 101-109. [cited by applicant]
Day, George S., “The Product Life Cycle: Analysis and Applications Issues”, In Journal of Marketing, vol. 45, Issue 4, Sep. 1, 1981, pp. 60-67. [cited by applicant]
Deldjoo, et al., “Using Visual Features Based on MPEG-7 and Deep Learning for Movie Recommendation”, In International Journal of Multimedia Information Retrieval, vol. 7, Issue 4, Jun. 14, 2018, 13 Pages. [cited by applicant]
Devlin, et al., “BERT: Pretraining of Deep Bidirectional Transformers for Language Understanding”, In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human L… [cited by applicant]
Devlin, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, In Repository of arXiv:1810.04805v1, Oct. 11, 2018, 14 Pages. [cited by applicant]
Dror, et al., “The Yahoo! Music Dataset and KDD-Cup'11”, In Proceedings of KDD Cup of Machine Learning Research, vol. 18, Jun. 2012, pp. 3-18. [cited by applicant]
“IMDb”, Retrieved from: https://web.archive.org/web/20211112182844/https:/www.imdb.com/, Nov. 12, 2021, 5 Pages. [cited by applicant]
Frolov, et al., “HybridSVD: When Collaborative Information is Not Enough”, In Proceedings of the 13th ACM Conference on Recommender Systems, Sep. 16, 2019, pp. 331-339. [cited by applicant]
Harper, et al., “The MovieLens Datasets: History and Context”, In Journal of ACM Transactions on Interactive Intelligent Systems, vol. 5, Issue 4, Article 19, Dec. 22, 2015, 19 Pages. [cited by applicant]
He, et al., “Deep Residual Learning for Image Recognition”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 27, 2016, pp. 770-778. [cited by applicant]
He, et al., “Neural Collaborative Filtering”, In Proceedings of the 26th International Conference on World Wide Web, Apr. 3, 2017, pp. 173-182. [cited by applicant]
Kingma, et al., “Adam: A Method for Stochastic Optimization”, In Repository of arXiv:1412.6980v1, Dec. 22, 2014, 9 Pages. [cited by applicant]
Koren, et al., “Matrix Factorization Techniques for Recommender Systems”, In Journal of Computer, vol. 42, Issue 8, Aug. 7, 2009, pp. 30-37. [cited by applicant]
Lam, et al., “Addressing Cold-Start Problem in Recommendation Systems”, In Proceedings of the 2nd International Conference on Ubiquitous Information Management and Communication, Jan. 31, 2008, pp. 208-211. [cited by applicant]
Lee, et al., “MeLU: Meta-Learned User Preference Estimator for Cold-Start Recommendation”, In Repository of arXiv:1908.00413v1, Aug. 4, 2019, 10 Pages. [cited by applicant]
Li, et al., “Collaborative Variational Autoencoder for Recommender Systems”, In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Aug. 13, 2017, pp. 305-314. [cited by applicant]
Çano, et al., “Hybrid Recommender Systems: A Systematic Literature Review”, In Journal of Intelligent Data Analysis, vol. 21, Issue 6, Nov. 15, 2017, 38 Pages. [cited by applicant]
Malkiel, et al., “RecoBERT: A Catalog Language Model for Text-Based Recommendations”, In Repository of arXiv:2009.13292v1, Sep. 25, 2020, 12 Pages. [cited by applicant]
Oord, et al., “Deep Content-Based Music Recommendation”, In Proceedings of the 26th International Conference on Neural Information Processing Systems, vol. 2, Dec. 5, 2013, 9 Pages. [cited by applicant]
Park, et al., “The Long Tail of Recommender Systems and How to Leverage It”, In Proceedings of the ACM Conference on Recommender Systems, Oct. 23, 2008, pp. 11-18. [cited by applicant]
Resnick, et al., “Recommender Systems”, In Journal of Communications of the ACM, vol. 40, Issue 3, Mar. 1, 1997, pp. 56-58. [cited by applicant]
Ricci, et al., “Introduction to Recommender Systems Handbook”, In Publication of Springer, Oct. 5, 2010, 35 Pages. [cited by applicant]
Rink, et al., “Product Life Cycle Research: A Literature Review”, In Journal of Business Research, vol. 7, Issue 3, Sep. 1, 1979, pp. 219-242. [cited by applicant]
Schein, et al., “Methods and Metrics for Cold-Start Recommendations”, In Proceedings of the 25th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, Aug. 11, 2002, pp. 253-260. [cited by applicant]
Tsukuda, et al., “DualDiv: Diversifying Items and Explanation Styles in Explainable Hybrid Recommendation”, In Proceedings of the ACM Conference on Recommender Systems, Sep. 16, 2019, pp. 398-402. [cited by applicant]
Vincent, et al., “Stacked Denoising Autoencoders: Learning Useful Representations in a Deep Network with a Local Denoising Criterion”, In Journal of Machine Learning Research, vol. 11, Dec. 2010, pp. 3371-3408. [cited by applicant]
Wang, et al., “Collaborative Deep Learning for Recommender Systems”, In Proceedings of the 21st ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Aug. 10, 2015, pp. 1235-1244. [cited by applicant]
Wang, et al., “Collaborative Topic Modeling for Recommending Scientific Articles”, In Proceedings of the 17th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Aug. 21, 2011, pp. 448-456. [cited by applicant]
Wang, et al., “Collaborative Topic Regression with Social Regularization for Tag Recommendation”, In Proceedings of the 23rd International Joint Conference on Artificial Intelligence, Aug. 3, 2013, 7 Pages. [cited by applicant]
Wang, et al., “Neural Graph Collaborative Filtering”, In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval, Jul. 21, 2019, pp. 165-174. [cited by applicant]
Wei, et al., “Collaborative Filtering and Deep Learning Based Recommendation System for Cold Start Items”, In Journal of Expert Systems with Applications, vol. 69, Issue 1, Mar. 1, 2017, pp. 29-39. [cited by applicant]
Wolf, et al., “Huggingface's Tansformers: State-of-the-Art Natural Language Processing”, In Repository of arXiv:1910.03771v1, Oct. 9, 2019, 11 Pages. [cited by applicant]
Wu, et al., “Collaborative Denoising Auto-Encoders for Top-N Recommender Systems”, In Proceedings of the 9th ACM International Conference on Web Search and Data Mining, Feb. 22, 2016, pp. 153-162. [cited by applicant]
Zhang, et al., “Content-Collaborative Disentanglement Representation Learning for Enhanced Recommendation”, In Proceedings of 14th ACM Conference on Recommender Systems, Sep. 22, 2020, pp. 43-52. [cited by applicant]
Lika, et al., “Facing the Cold Start Problem in Recommender Systems”, In Journal of Expert Systems with Applications, vol. 41, Issue 4, Part 2, Mar. 2014, pp. 2065-2073. [cited by applicant]