IP Library Granted Patent US 11,902,628
Granted Patent B2
US 11,902,628 · App. 17/374,606 · Granted Feb 13, 2024

Masked model training of a prediction network

Inventors: Pengyu Zhao (Beijing, CN); Chunxu Xu (Beijing, CN); Xianghui Mao (Beijing, CN); Xiaohui Xie (Beijing, CN)
Assignee: HULU, LLC
H04N21/4826G06N3/045G06N3/084H04L65/612
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,902,628
App. No.
17/374,606
Granted
Feb 13, 2024
Kind
B2
Abstract

In some embodiments, a method receives a first sequence of inputs for processing via a sub-model of a plurality of sub-model. The plurality of sub-models are part of a main model. An input in the sequence of inputs is masked with a masked value to generate a second sequence of inputs. The method processes the second sequence of inputs using the sub-model to generate a sequence of features that correspond to the second sequence of inputs and processes the sequence of features to generate a first output. The first output is processed to generate a second output of the main model. The sub-model is trained based on a feature in the sequence of features that corresponds to the masked input and the second output.

Claims (56)

1. A method comprising:

receiving, by a computing device, a first sequence of inputs for processing via a sub-model of a plurality of sub-models, wherein the plurality of sub-models are part of a main model;

masking, by the computing device, an input in the first sequence of inputs with a masked value to generate a second sequence of inputs;

processing, by the computing device, the second sequence of inputs using the sub-model to generate a sequence of features that correspond to the second sequence of inputs, wherein the sub-model comprises a self-attention sub-model that applies attention to inputs in the second sequence of inputs based on relationships between the inputs;

generating, by the computing device, a first output of the main model based on the sequence of features; and

training, by the computing device, the sub-model based on a feature in the sequence of features that corresponds to the masked input and the first output, wherein training the sub-model comprises:

generating a first value based on a feature in the sequence of features that corresponds to the masked input;

generating a second value based on the first output; and

adjusting a parameter of the sub-model using the first value and the second value.

2. The method of claim 1 , wherein the first sequence of inputs is based on behavior of a user account that is using a video delivery system.

3. The method of claim 2 , wherein:

the masked value is based on information separate from the behavior of the user account.

4. The method of claim 1 , wherein the first output represents a relevance of an item to the first sequence of inputs and other input processed by the main model.

5. The method of claim 4 , wherein the relevance comprises a probability of selection of the item.

6. The method of claim 1 , wherein the sequence of features is a representation of the second sequence of inputs.

7. The method of claim 1 , further comprising:

generating a second output from the sequence of features that represents the second sequence of inputs, wherein the second output is used to generate the first output of the main model.

8. The method of claim 1 , wherein masking the input comprises:

randomly masking one of the first sequence of inputs.

9. The method of claim 1 , wherein the attention applies weights to the inputs differently.

10. The method of claim 1 , wherein processing the second sequence of inputs comprises:

processing the second sequence of inputs using multiple layers in the sub-model to generate the sequence of features.

11. The method of claim 7 , wherein the first output is processed by one or more other sub-models in the plurality of sub-models to generate the first output.

12. The method of claim 11 , wherein the one or more other sub-models receive input other than the first sequence of inputs.

13. A method comprising:

receiving, by a computing device, a first sequence of inputs for processing via a sub-model of a plurality of sub-models, wherein the plurality of sub-models are part of a main model;

masking, by the computing device, an input in the first sequence of inputs with a masked value to generate a second sequence of inputs;

processing, by the computing device, the second sequence of inputs using the sub-model to generate a sequence of features that correspond to the second sequence of inputs;

generating, by the computing device, a first output of the main model based on the sequence of features; and

training, by the computing device, the sub-model based on a feature in the sequence of features that corresponds to the masked input and the first output,

wherein training the sub-model comprises:

comparing the feature that corresponds to the masked input to the first sequence of inputs,

calculating a similarity between the feature that corresponds to the masked input and the first sequence of inputs; and

using the similarity to train the sub-model, wherein using the similarity to train the sub-model comprises:

adjusting a parameter of the sub-model to increase a similarity of the feature that corresponds to the masked input to the corresponding input in the first sequence of inputs compared to a similarity for other inputs in the first sequence of inputs.

14. A non-transitory computer-readable storage medium having stored thereon computer executable instructions, which when executed by a computing device, cause the computing device to be operable for:

receiving a first sequence of inputs for processing via a sub-model of a plurality of sub-models, wherein the plurality of sub-models are part of a main model;

masking an input in the first sequence of inputs with a masked value to generate a second sequence of inputs;

processing the second sequence of inputs using the sub-model to generate a sequence of features that correspond to the second sequence of inputs, wherein the sub-model comprises a self-attention sub-model that applies attention to inputs in the second sequence of inputs based on relationships between the inputs;

generating a first output of the main model based on the sequence of features; and

training the sub-model based on a feature in the sequence of features that corresponds to the masked input and the first output, wherein training the sub-model comprises:

generating a first value based on a feature in the sequence of features that corresponds to the masked input;

generating a second value based on the first output; and

adjusting a parameter of the sub-model using the first value and the second value.

15. The non-transitory computer-readable storage medium of claim 14 , wherein the first output represents a relevance of an item to the first sequence of inputs and other input processed by the main model.

16. The non-transitory computer-readable storage medium of claim 14 , wherein training the sub-model comprises:

comparing the feature that corresponds to the masked input to the first sequence of inputs,

calculating a similarity between the feature that corresponds to the masked input and the first sequence of inputs; and

using the similarity to train the sub-model.

17. The non-transitory computer-readable storage medium of claim 16 , wherein using the similarity to train the sub-model comprises:

adjusting a parameter of the sub-model to increase a similarity of the feature that corresponds to the masked input to the corresponding input in the first sequence of inputs compared to a similarity for other inputs in the first sequence of inputs.

18. The non-transitory computer-readable storage medium of claim 14 , further operable for:

generating a second output from the sequence of features that represents the second sequence of inputs, wherein the second output is used to generate the first output of the main model.

19. The non-transitory computer-readable storage medium of claim 14 , wherein masking the input comprises: randomly masking one of the first sequence of inputs.

20. The non-transitory computer-readable storage medium of claim 14 , wherein processing the second sequence of inputs comprises:

processing the second sequence of inputs using multiple layers in the sub-model to generate the sequence of features.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 13, 2021
From: ZHAO, PENGYU; XU, CHUNXU; MAO, XIANGHUI; XIE, XIAOHUI
To: HULU, LLC
Reel/Frame 056842/0278 →
Continuity (1)
Related Publication 20230019564A1 · Jan 19, 2023