IP Library Granted Patent US 10,616,625
Granted Patent B2
US 10,616,625 · App. 15/985,473 · Granted Apr 7, 2020

Reinforcement learning network for recommendation system in video delivery system

Inventors: Siguang Huang (Beijing, CN); Guoxin Zhang (Beijing, CN); Hanning Zhou (Beijing, CN)
Assignee: HULU, LLC
H04N21/251H04N21/4532H04N21/472
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,616,625
App. No.
15/985,473
Granted
Apr 7, 2020
Kind
B2
Abstract

A method receives user behavior information at a first system. The user behavior information is determined by user interaction with a first list sent to the user by a first network on a video delivery service. A first state is generated using the received user behavior information and prior user behavior information by the user from cells that store the prior user behavior. The method inputs the first state into a second network with the first recommendation list to generate a value that evaluates a performance of recommending the first recommendation list. An update to parameters is generated for the first network and provided to the first network. The first network generates a second state from the received user behavior information and prior user behavior information derived from cells that store the prior user behavior and outputs a second recommendation list using the second state and updated parameters.

Claims (55)

1. A method comprising:

receiving, by a computing device, user behavior information for a user at a first system, the user behavior information determined by user interaction with a first recommendation list sent to an interface being used by the user, wherein the first recommendation list is generated by a first network on a video delivery service;

generating, by the computing device, a first state using the received user behavior information and prior user behavior information by the user from a first set of cells that store the prior user behavior;

inputting, by the computing device, the first state into a second network with the first recommendation list from the first network to generate a value that evaluates a performance of recommending the first recommendation list;

generating, by the computing device, a first evaluation score for the first network, wherein first network parameters are updated based on the first evaluation score to generate a first update to the first network parameters;

generating, by the computing device, a second state using the received user behavior information and the prior user behavior information by the user from a second set of cells that store the prior user behavior; and

outputting, by the computing device, a second recommendation list using the first network based on the first update to the first network parameters and the second state without being prompted by the user, wherein the second network uses the second recommendation list to generate a second evaluation score for the first network, wherein the first update of the first network parameters is updated based on the second evaluation score to generate a second update to the first network parameters during a live environment in which the first network is being used to provide recommendation lists to the interface.

2. The method of claim 1 , further comprising:

providing the first recommendation list to the user on the interface for the video delivery service; and

receiving the user behavior information determined by user interaction with the first recommendation list on the interface.

3. The method of claim 2 , wherein the user behavior information is determined by the user selecting a recommendation in the first recommendation list on the interface.

4. The method of claim 1 , wherein outputting the second recommendation list using the first network comprises:

generating a first vector in a continuous embedding space; and

selecting N videos that have respective second vectors closest to the first vector.

5. The method of claim 4 , wherein the first vector and the respective second vectors define characteristics for videos.

6. The method of claim 1 , wherein the first set of cells and the second set of cells store sequential user behavior that is used to generate the first state and the second state, respectively.

7. The method of claim 1 , further comprising:

sending the second recommendation list to the second network; and

inputting the second state into the second network with the second recommendation list from the first network to generate a second evaluation value that evaluates the performance of recommending the second recommendation list.

8. The method of claim 7 , wherein the second recommendation list is not sent to the user.

9. The method of claim 7 , wherein the second evaluation score is used to further update the first update to the first network parameters to generate a second update to the first network parameters.

10. The method of claim 9 , further comprising:

inputting the second state into the first network to output a third recommendation list using the first network based on the second update to the first network parameters without being prompted by the user.

11. The method of claim 10 , wherein the third recommendation list is not sent to the user.

12. The method of claim 10 , further comprising:

continuing to evaluate recommendations output by the first network using the second network until a change to a gradient used to update the first network parameters is minimized.

13. A non-transitory computer-readable storage medium containing instructions, that when executed, control a computer system to be configured for:

receiving user behavior information for a user at a first system, the user behavior information determined by user interaction with a first recommendation list sent to an interface being used by the user, wherein the first recommendation list is generated by a first network on a video delivery service;

generating a first state using the received user behavior information and prior user behavior information by the user from a first set of cells that store the prior user behavior;

inputting the first state into a second network with the first recommendation list from the first network to generate a value that evaluates a performance of recommending the first recommendation list;

generating a first evaluation score for the first network, wherein first network parameters are updated based on the first evaluation score to generate a first update to the first network parameters;

generating a second state using the received user behavior information and the prior user behavior information by the user from a second set of cells that store the prior user behavior; and

outputting a second recommendation list using the first network based on the first update to the first network parameters and the second state without being prompted by the user, wherein the second network uses the second recommendation list to generate a second evaluation score for the first network, wherein the first update of the first network parameters is updated based on the second evaluation score to generate a second update to the first network parameters during a live environment in which the first network is being used to provide recommendation lists to the interface.

14. The non-transitory computer-readable storage medium of claim 13 , further configured for:

providing the first recommendation list to the user on the interface for the video delivery service; and

receiving the user behavior information determined by user interaction with the first recommendation list on the interface.

15. The non-transitory computer-readable storage medium of claim 13 , wherein outputting the second recommendation list using the first network comprises:

generating a first vector in a continuous embedding space; and

selecting N videos that have respective second vectors closest to the vector.

16. The non-transitory computer-readable storage medium of claim 13 , further configured for:

sending the second recommendation list to the second network; and

inputting the second state into the second network with the second recommendation list from the first network to generate a second evaluation value that evaluates the performance of recommending the second recommendation list.

17. The non-transitory computer-readable storage medium of claim 16 , wherein the second recommendation list is not sent to the user.

18. The non-transitory computer-readable storage medium of claim 16 , wherein the second evaluation score is used to further update the first update to the first network parameters to generate a second update to the first network parameters.

19. The non-transitory computer-readable storage medium of claim 16 , further configured for:

inputting the second state into the first network to output a third recommendation list using the first network based on the second update to the first network parameters without being prompted by the user.

20. An apparatus comprising:

one or more computer processors; and

a non-transitory computer-readable storage medium comprising instructions, that when executed, control the one or more computer processors to be configured for:

receiving user behavior information for a user at a first system, the user behavior information determined by user interaction with a first recommendation list sent to an interface being used by the user, wherein the first recommendation list is generated by a first network on a video delivery service;

generating a first state using the received user behavior information and prior user behavior information by the user from a first set of cells that store the prior user behavior;

inputting the first state into a second network with the first recommendation list from the first network to generate a value that evaluates a performance of recommending the first recommendation list;

generating a first evaluation score for the first network, wherein first network parameters are updated based on the first evaluation score to generate a first update to the first network parameters;

generating a second state using the received user behavior information and the prior user behavior information by the user from a second set of cells that store the prior user behavior; and

outputting a second recommendation list using the first network based on the first update to the first network parameters and the second state without being prompted by the user, wherein the second network uses the second recommendation list to generate a second evaluation score for the first network, wherein the first update of the first network parameters is updated based on the second evaluation score to generate a second update to the first network parameters during a live environment in which the first network is being used to provide recommendation lists to the interface.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 21, 2018
From: HUANG, SIGUANG; ZHANG, GUOXIN; ZHOU, HANNING
To: HULU, LLC
Reel/Frame 045865/0329 →
Continuity (1)
Related Publication 20190356937A1 · Nov 21, 2019
Cited By (2)
US 12,412,105 US 12,413,799