IP Library Granted Patent US 12,389,079
Granted Patent B2
US 12,389,079 · App. 18/402,643 · Granted Aug 12, 2025

Masked model training of a prediction network

Inventors: Pengyu Zhao (Beijing, CN); Chunxu Xu (Beijing, CN); Xianghui Mao (Beijing, CN); Xiaohui Xie (Beijing, CN)
Assignee: HULU, LLC
H04N21/4826G06N3/045G06N3/084H04L65/612
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,389,079
App. No.
18/402,643
Granted
Aug 12, 2025
Kind
B2
Abstract

In some embodiments, a method receives a first sequence of inputs for processing via a sub-model of a plurality of sub-model. The plurality of sub-models are part of a main model. An input in the sequence of inputs is masked with a masked value to generate a second sequence of inputs. The method processes the second sequence of inputs using the sub-model to generate a sequence of features that correspond to the second sequence of inputs and processes the sequence of features to generate a first output. The first output is processed to generate a second output of the main model. The sub-model is trained based on a feature in the sequence of features that corresponds to the masked input and the second output.

Claims (66)

1. A method comprising:

receiving, by a computing device, a request for a recommendation of an item;

applying, by the computing device, the request to a trained main model to obtain a prediction relating to data in the request;

generating, by the computing device, a recommendation based on the prediction; and

sending, by the computing device, the recommendation as a response to the request, wherein the trained model comprises a trained sub-model of a plurality of sub-models, and wherein the sub-model is trained by:

masking an input in a sequence of inputs with a masked value in the sub-model, wherein the masked value is different from a value of the input;

generating, using the sub-model, a sequence of features using the sequence of inputs with the masked value;

generating a first value based on a comparison of an output of the main model to an expected output, wherein the output is determined using the sequence of features that is generated by the sub-model;

determining a second value that is based on a comparison of a feature in the sequence of features that corresponds to the masked input to an expected feature for the input; and

training a parameter of the sub-model using the second value and the first value, wherein the parameter is used by the sub-model to generate the sequence of features from the sequence of inputs.

2. The method of claim 1 , wherein:

the sequence of inputs comprises a first sequence of inputs,

masking the input in the first sequence of inputs with the masked value generates a second sequence of inputs, and

the sequence of features is generated using the second sequence of inputs.

3. The method of claim 2 , wherein training the sub-model comprises:

comparing a first output that corresponds to the masked input to the second sequence of inputs,

calculating a similarity between the first output and the second sequence of inputs based on the comparing; and

using the similarity to train the sub-model.

4. The method of claim 3 , wherein using the similarity to train the sub-model comprises:

adjusting the parameter of the sub-model to increase a similarity of the feature that corresponds to the masked input to the corresponding input in the sequence of inputs compared to a similarity for other inputs in the sequence of inputs.

5. The method of claim 1 , wherein the sequence of inputs is based on behavior of a user account that is using a video delivery system.

6. The method of claim 5 , wherein:

the masked value is based on information separate from the behavior of the user account.

7. The method of claim 1 , wherein the output represents a relevance of the item to the sequence of inputs and other input processed by the main model.

8. The method of claim 7 , wherein the relevance comprises a probability of selection of the item.

9. The method of claim 1 , wherein the sequence of features is a representation of the sequence of inputs.

10. The method of claim 1 , wherein the output comprises a first output, the method further comprising:

generating a second output from the sequence of features, wherein the second output is used to generate the first output of the main model.

11. The method of claim 1 , wherein masking the input comprises:

randomly masking one of the sequence of inputs.

12. The method of claim 1 , further comprising:

applying attention to the sequence of inputs based on a relationship between inputs, wherein the attention applies weights to the inputs differently.

13. The method of claim 1 , further comprising:

processing the sequence of inputs using multiple layers in the sub-model to generate the sequence of features.

14. The method of claim 1 , wherein the sequence of features is processed by one or more other sub-models in the plurality of sub-models to generate the output of the main model.

15. The method of claim 14 , wherein the one or more other sub-models receive input other than the sequence of inputs.

16. The method of claim 1 , wherein

the second value is a first loss based on the comparison of the feature in the sequence of features that corresponds to the masked input to the expected feature for the input;

the first value is a second loss based on the comparison of the output of the main model to the expected out.

17. The method of claim 1 , wherein the sub-model comprises a self-attention sub-model that applies attention to inputs in the sequence of inputs based on relationships between the inputs.

18. A non-transitory computer-readable storage medium having stored thereon computer executable instructions, which when executed by a computing device, cause the computing device to be operable for:

receiving a request for a recommendation of an item;

applying the request to a trained main model to obtain a prediction relating to data in the request;

generating a recommendation based on the prediction; and

sending the recommendation as a response to the request, wherein the trained model comprises a trained sub-model of a plurality of sub-models, and wherein the sub-model is trained by:

masking an input in a sequence of inputs with a masked value in the sub-model, wherein the masked value is different from a value of the input;

generating, using the sub-model, a sequence of features using the sequence of inputs with the masked value;

generating a first value based on a comparison of an output of the main model to an expected output, wherein the output is determined using the sequence of features that is generated by the sub-model;

determining a second value that is based on a comparison of a feature in the sequence of features that corresponds to the masked input to an expected feature for the input; and

training a parameter of the sub-model using the second value and the first value, wherein the meter is used by the sub-model to generate the sequence of features from the sequence of inputs.

19. The non-transitory computer-readable storage medium of claim 18 , wherein:

the sequence of inputs comprises a first sequence of inputs,

masking the input in the first sequence of inputs with the masked value generates a second sequence of inputs, and

the sequence of features is generated using the second sequence of inputs.

20. An apparatus comprising:

one or more computer processors; and

a computer-readable storage medium comprising instructions for controlling the one or more computer processors to be operable for:

receiving a request for a recommendation of an item;

applying the request to a trained main model to obtain a prediction relating to data in the request;

generating a recommendation based on the prediction; and

sending the recommendation as a response to the request, wherein the trained model comprises a trained sub-model of a plurality of sub-models, and wherein the sub-model is trained by:

masking an input in a sequence of inputs with a masked value in the sub-model, wherein the masked value is different from a value of the input;

generating, using the sub-model, a sequence of features using the sequence of inputs with the masked value;

generating a first value based on a comparison of an output of the main model to an expected output wherein the output is determined using the sequence of features that is generated by the sub-model;

determining a second value that is based on a comparison of a feature in the sequence of features that corresponds to the masked input to an expected feature for the input; and

training a parameter of the sub-model using the second value and the first value, wherein the parameter is used by the sub-model to generate the sequence of features from the sequence of inputs.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 2, 2024
From: ZHAO, PENGYU; XU, CHUNXU; MAO, XIANGHUI; XIE, XIAOHUI
To: HULU, LLC
Reel/Frame 065999/0186 →
Continuity (2)
Division 17374606 · Jul 13, 2021
Related Publication 20240137620A1 · Apr 25, 2024
References Cited (36)
US 10290040B1 · Misra · 2019 [cited by examiner]
US 10762422B2 · Shaked · 2020 [cited by examiner]
US 10860860B1 · Huynh · 2020 [cited by examiner]
US 11507876B1 · Kuo · 2022 [cited by examiner]
US 11580585B1 · Rajana · 2023 [cited by examiner]
US 11901080B1 · Matt · 2024 [cited by examiner]
US 20170054779A1 · Ehmann · 2017 [cited by examiner]
US 20190026274A1 · Deng · 2019 [cited by examiner]
US 20190028766A1 · Wold · 2019 [cited by examiner]
US 20200193288A1 · Li · 2020 [cited by examiner]
US 20200288204A1 · Duersch · 2020 [cited by examiner]
US 20200342955A1 · Guo · 2020 [cited by examiner]
US 20210029391A1 · Choudhari · 2021 [cited by examiner]
US 20210089963A1 · Baek · 2021 [cited by examiner]
US 20210160572A1 · Menendez · 2021 [cited by examiner]
US 20210233144A1 · Gaur · 2021 [cited by examiner]
US 20210303777A1 · Tong et al. · 2021 [cited by applicant]
US 20210383223A1 · Tan · 2021 [cited by examiner]
US 20210398016A1 · Tsimerman · 2021 [cited by examiner]
US 20220027562A1 · Zhang · 2022 [cited by examiner]
US 20220043823A1 · Belli · 2022 [cited by examiner]
US 20220100867A1 · Sinn · 2022 [cited by examiner]
US 20220207557A1 · Siddhartha · 2022 [cited by examiner]
US 20220222706A1 · Arora · 2022 [cited by examiner]
US 20220237892A1 · Anderton-Yang · 2022 [cited by examiner]
US 20220245160A1 · Mohamed · 2022 [cited by examiner]
US 20220254083A1 · Akhoundi · 2022 [cited by applicant]
US 20220394340A1 · Jiang · 2022 [cited by examiner]
US 20220408131A1 · Gardner · 2022 [cited by examiner]
US 20230019564A1 · Zhao et al. · 2023 [cited by applicant]
US 20230230378A1 · Rüfenacht et al. · 2023 [cited by applicant]
Ashish Vaswani, et al., “Attention Is All You Need”, https://arxiv.org/pdf/1706.03762.pdf, Dec. 6, 2017. [cited by applicant]
Jacob Devlin, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, https://arxiv.org/pdf/1810.04805.pdf, May 24, 2019. [cited by applicant]
Jason Brownlee, “A Gentle Introduction to Cross-Entropy for Machine Learning”, https://machinelearningmastery.com/cross-entropy-for-machine-learning/, Dec. 22, 2020. [cited by applicant]
Qiwei Chen, et al., “Behavior Sequence Transformer for E-commerce Recommendation in Alibaba”, https://arxiv.org/pdf/1905.06874, May 15, 2019. [cited by applicant]
Wikipedia contributors, “Receiver operating characteristic”, Wikipedia, The Free Encyclopedia, https://en.wikipedia.org/wiki/Receiver_operating_characteristic, Jun. 22, 2021. [cited by applicant]