IP Library Granted Patent US 11,546,665
Granted Patent B2
US 11,546,665 · App. 17/315,107 · Granted Jan 3, 2023

Reinforcement learning for guaranteed delivery of supplemental content

Inventors: Pengfei Gao (Beijing, CN); Dingming Wu (Beijing, CN); Chunyang Wei (Beijing, CN); Xiaohui Xie (Beijing, CN); Shulei Ma (Beijing, CN)
Assignee: HULU, LLC
H04N21/472G06N5/04G06N20/00G06Q30/0212H04N21/44222H04N21/4532H04N21/4662H04N21/4667
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,546,665
App. No.
17/315,107
Granted
Jan 3, 2023
Kind
B2
Abstract

In some embodiments, a method receives a request for supplemental content to be provided in association with main content. The method selects an instance of supplemental content based on a long-term reward metric and a short-term reward metric. The long-term reward metric is based on feedback from delivery of a plurality of instances of supplemental content and a delivery status for a delivery constraint of one instance of supplemental content. The short-term reward metric is based on feedback from delivery of the one instance of supplemental content. The long-term reward metric is based on feedback from delivery of a plurality of instances of supplemental content and the short-term reward metric is based on feedback from delivery of one instance of supplemental content. The instance of supplemental content is provided to a client device.

Claims (54)

1. A method comprising:

calculating, by a computing device, a short-term reward metric for an instance of supplemental content, wherein the short-term reward metric is calculated using first feedback that is based on a display of the instance of supplemental content; and

calculating, by a computing device, a long-term reward metric for an instance of supplemental content, wherein the long-term reward metric is calculated using second feedback that is received after the display of the instance of the supplemental content and a delivery status for a delivery constraint of the instance of supplemental content;

receiving, by a computing device, a request for supplemental content to be provided in association with main content;

selecting, by the computing device, from a plurality of instances of supplemental content to select a selected instance of supplemental content, wherein the long-term reward metric and the short-term reward metric for the instance of supplemental content are used in the selecting; and

causing, by the computing device, delivery of the selected instance of supplemental content to a client device for the request.

2. The method of claim 1 , further comprising:

receiving a state that is associated with characteristics of an account that is using the client device, wherein the state is used to select the instance of supplemental content.

3. The method of claim 1 , further comprising:

receiving a state that is associated the delivery status of instance of supplemental content, wherein the state is used to select the selected instance of supplemental content.

4. The method of claim 1 , further comprising:

receiving a first state that is associated characteristics of an account that is using the client device;

receiving a second state that is associated the delivery status of instance of supplemental content; and

combining the first state and the second state into a third state that is used to select the selected instance of supplemental content.

5. The method of claim 1 , wherein selecting from the plurality of instances instance of supplemental content comprises:

outputting information that rates instances in the plurality of instances of supplemental content that are eligible to be delivered for the request, wherein the information is used to select the selected instance of supplemental content.

6. The method of claim 5 , wherein the information comprises an allocation plan that rates the instances in the plurality of instances of supplemental content for selection.

7. The method of claim 6 , wherein the information comprises relevance weights that are based on a relevance to an account associated with the client device.

8. The method of claim 7 , wherein selecting from the plurality of instances of supplemental content comprises:

generating selection weights from the allocation plan and the relevance weights; and

selecting the selected instance of supplemental content based on the selection weights for instances of the plurality of instances of supplemental content.

9. The method of claim 1 , wherein the short-term reward metric and the long-term reward metric are based on feedback from an account that is using the client device during a session associated with viewing the main content.

10. The method of claim 1 , further comprising:

training a model that is used to select the selected instance of supplemental content based on feedback from an account that is using the client device.

11. The method of claim 10 , wherein training the model comprises:

receiving the long-term reward metric and the delivery constraint for the instance of supplemental content, wherein the delivery constraint specifies a delivery goal of the instance of supplement content; and

calculating a constrained reward metric based on the long-term reward metric and an unsatisfied delivery constraint based on the delivery status, wherein the constrained reward is used to train the model.

12. The method of claim 11 , wherein training the model comprises:

calculating a general reward metric to use to train the model based on the long-term reward metric and the constrained reward metric.

13. The method of claim 1 , wherein the long-term reward metric assigns a value to the instance of supplemental content that is measured over delivery of multiple instances of supplemental content.

14. The method of claim 1 , wherein the long-term reward metric assigns a value to the instance of supplemental content based on a time that a user account continued to view the main content after the display of the instance of supplemental content.

15. The method of claim 14 , wherein the short-term reward metric assigns a value to the instance of supplemental content based on interaction with the instance of supplemental content while being displayed at the client device.

16. A non-transitory computer-readable storage medium containing instructions, that when executed, control a computer system to be operable for:

calculating a short-term reward metric for an instance of supplemental content, wherein the short-term reward metric is calculated using first feedback that is based on a display of the instance of supplemental content; and

calculating a long-term reward metric for the instance of supplemental content, wherein the long-term reward metric is calculated using second feedback that is received after the display of the instance of the supplemental content and a delivery status for a delivery constraint of the instance of supplemental content;

receiving a request for supplemental content to be provided in association with main content;

selecting from a plurality of instances of supplemental content to select a selected instance of supplemental content, wherein the long-term reward metric and the short-term reward metric for the instance of supplemental content are used in the selecting; and

causing delivery of the selected instance of supplemental content to a client device for the request.

17. The non-transitory computer-readable storage medium of claim 16 , further operable for:

receiving a first state that is associated characteristics of an account that is using the client device;

receiving a second state that is associated the delivery status of instance of supplemental content; and

combining the first state and the second state into a third state that is used to select the selected instance of supplemental content.

18. The non-transitory computer-readable storage medium of claim 16 , wherein selecting from the plurality of instances of supplemental content comprises:

outputting information that rates instances in the plurality of instances of supplemental content that are eligible to be delivered for the request, wherein the information is used to select the selected instance of supplemental content.

19. The non-transitory computer-readable storage medium of claim 16 , further operable for:

training a model that is used to select the selected instance of supplemental content based on feedback from an account that is using the client device.

20. An apparatus comprising:

one or more computer processors; and

a non-transitory computer-readable storage medium comprising instructions, that when executed, control the one or more computer processors to be operable for:

calculating a short-term reward metric for an instance of supplemental content, wherein the short-term reward metric is calculated using first feedback that is based on a display of the instance of supplemental content; and

calculating a long-term reward metric for an instance of supplemental content, wherein the long-term reward metric is calculated using second feedback that is received after the display of the instance of the supplemental content and a delivery status for a delivery constraint of the instance of supplemental content;

receiving a request for supplemental content to be provided in association with main content;

selecting from a plurality of instances of supplemental content to select a selected instance of supplemental content, wherein the long-term reward metric and the short-term reward metric for the instance of supplemental content are used in the selecting; and

causing delivery of the selected instance of supplemental content to a client device for the request.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 7, 2026
From: HULU, LLC
To: ADEIA MEDIA HOLDINGS INC.
Reel/Frame 075567/0854 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 7, 2021
From: GAO, PENGFEI; WU, DINGMING; WEI, CHUNYANG; XIE, XIAOHUI; MA, SHULEI
To: HULU, LLC
Reel/Frame 056177/0127 →
Continuity (1)
Related Publication 20220360854A1 · Nov 10, 2022