IP Library Granted Patent US 10,608,976
Granted Patent B2
US 10,608,976 · App. 15/793,787 · Granted Mar 31, 2020

Delayed processing for arm policy determination for content management system messaging

Inventors: Aditi Jain (San Francisco, CA); Manveer Singh Chawla (San Francisco, CA); Thomas Berg (San Francisco, CA); Swapnil Zarekar (San Francisco, CA); Robert Kajic (San Francisco, CA); Karandeep Johar (San Francisco, CA); Aaron Feldstein (San Francisco, CA); Walter Kim (San Francisco, CA); Joe Nudell (San Francisco, CA); Jenny Dong (San Francisco, CA); Jared Wilson (San Francisco, CA); Luke Thompson (San Francisco, CA); David Kriegman (San Francisco, CA)
Assignee: DROPBOX, INC.
H04L51/26G06F9/5038H04L51/36H04L63/102H04L67/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,608,976
App. No.
15/793,787
Granted
Mar 31, 2020
Kind
B2
Abstract

Computer-implemented techniques include, during a delayed processing window, receiving reward data for arm actions taken, where the arm actions were chosen based on a previous version of an arm choice policy, and the previous version of the arm choice policy was determined based on a previous set of reward data for a previous set of arm actions taken. When the delayed processing window has closed, a new arm choice policy is determined based at least in part on the action-reward data, and the previous set of reward data and/or the previous arm choice policy. After a request to choose an arm choice is received, a particular arm action to take is determined based on the new arm choice policy. This chosen arm is provided in response to the request.

Claims (62)

1. A method comprising:

while a delayed processing timing has not been met, receiving reward data for arm actions taken as a first set of action-reward data, wherein the arm actions were chosen based on a previous version of an arm choice policy, and wherein the previous version of the arm choice policy was determined at least in part based on a previous set of reward data for a previous set of arm actions taken;

when the delayed processing timing has been met, determining a new arm choice policy based at least in part on (1) the first set of action-reward data, and (2) one or more of the previous set of reward data for the previous set of arm actions taken or the previous version of the arm choice policy;

receiving a request to choose an arm from among arms in the new arm choice policy;

determining a particular arm action to take based at least in part on the new arm choice policy; and

providing the particular arm action to be taken in response to the request;

wherein the method is performed by one or more computing devices.

2. The method of claim 1 , further comprising performing an action of the determined particular arm action.

3. The method of claim 1 , further comprising:

while the delayed processing timing has not been met, receiving multiple requests wherein each request of the multiple requests is for choice of arm to be taken; and

when the delayed processing timing has been met, determining one or more corresponding arm actions to be taken in response to each request, of the multiple request, based at least in part on the new arm choice policy.

4. The method of claim 1 , wherein determining the particular arm action to take based at least in part on the new arm choice policy comprises using statistical variance in the new arm choice policy during the determining.

5. The method of claim 1 , wherein receiving reward data for arm actions taken comprises receiving first reward data for a first arm action, wherein the first reward data was determined based on passage of a particular timeout period after the first arm action was taken.

6. The method of claim 1 , further comprising determining multiple, ranked arms to take in response to the received request.

7. The method of claim 6 , further comprising choosing which action to take based at least in part on ranking of the multiple, ranked arms.

8. The method of claim 1 , wherein:

the first set of action-reward data comprises first context data for the arm actions taken; and

the new arm choice policy is determined based at least in part on the first set of action-reward data that includes the first context data for the arm actions taken; and

the method further comprises:

receiving second context data as part of the request, and

wherein determining the particular arm action to take comprises determining the particular arm action to take based at least in part on the new arm choice policy and the second context data.

9. A system comprising:

one or more computing devices;

memory; and

instructions, stored in the memory, and which, when executed by the system, cause the system to perform:

receiving, during a batch window, reward data for previous arm actions, wherein the previous arm actions were chosen based on a previous arm policy, and wherein the previous arm policy was chosen based on a previous set of arm choice-reward data;

determining a first set of arm choice-reward data based at least in part on the reward data for the previous arm actions received during the batch window;

after the batch window, determining a new arm choice policy based at least in part on the first set of arm choice-reward data, and at least one of: the previous set of arm choice-reward data for the previous set of arm choice-reward data or the previous arm policy;

receiving a request for an arm action to take;

determining a particular arm action to take based at least in part on the new arm choice policy; and

providing the particular arm action to be taken in response to the request.

10. The system of claim 9 , further comprising instructions which, when executed by the system, cause the system to perform performing an action of the determined particular arm action.

11. The system of claim 9 , further comprising instructions which, when executed by the system, cause the system to perform:

during the batch window, receiving multiple requests wherein each request of the multiple requests is for choice of arm to be taken; and

after the batch window, determining one or more corresponding arm actions to be taken in response to each request, of the multiple requests, based at least in part on the new arm choice policy.

12. The system of claim 9 , wherein receiving reward data for previous arm actions comprises receiving first reward data for a first arm action, wherein the first reward data was determined based on passage of a particular timeout period after the first arm action was taken.

13. The system of claim 9 , further comprising instructions which, when executed by the system, cause the system to perform determining multiple, ranked arms to take in response to the received request.

14. The system of claim 9 , wherein:

the first set of arm choice-reward data comprises first context data for the previous arm actions; and

the new arm choice policy is determined based at least in part on the first set of arm choice-reward data that includes the first context data for the previous arm actions; and

the system further comprises instructions which, when executed by the system, cause the system to perform:

receiving second context data as part of the request, and

wherein determining the particular arm action to take comprises determining the particular arm action to take based at least in part on the new arm choice policy and the second context data.

15. One or more non-transitory media comprising instructions which, when executed by a system having one or more computing devices, cause the system to perform:

while a batch window timing has not been met, receiving a first set of action-reward data, wherein the first set of action-reward data is associated with arm actions chosen based on a previous version of an arm choice policy, and wherein the previous version of the arm choice policy was determined at least in part on a previous set of action-reward data for a previous set of arm actions taken;

when the batch window timing has been met, determining a new arm choice policy based at least in part on (1) the first set of action-reward data, and (2) the previous set of action-reward data for the previous set of arm actions taken or the previous version of the arm choice policy;

receiving a request to choose an arm from among arms in the new arm choice policy;

determining a particular arm action to take based at least in part on the new arm choice policy; and

providing the determined particular arm action to be taken in response to the request.

16. The one or more non-transitory media of claim 15 , further comprising instructions which, when executed by the system, cause the system to perform performing an action of the determined particular arm action.

17. The one or more non-transitory media of claim 15 , further comprising instructions which, when executed by the system, cause the system to perform:

while the batch window timing has not been met, receiving multiple requests wherein each request of the multiple requests is for choice of arm to be taken; and

when the batch window timing has been met, determining one or more corresponding arm actions to be taken in response to each request, of the multiple requests, based at least in part on the new arm choice policy.

18. The one or more non-transitory media of claim 15 , wherein determining the particular arm action to take based at least in part on the new arm choice policy comprises using statistical variance in the new arm choice policy during the determining.

19. The one or more non-transitory media of claim 15 , wherein receiving the first set of action-reward data comprises receiving first reward data for a first arm action, wherein the first reward data was determined based on passage of a particular timeout period after the first arm action was taken.

20. The one or more non-transitory media of claim 15 , further comprising instructions which, when executed by the system, cause the system to perform determining multiple, ranked arms to take in response to the received request and choosing which action to take based at least in part on ranking of the multiple, ranked arms.

21. The one or more non-transitory media of claim 15 , wherein:

the first set of action-reward data comprises first context data for the arm actions chosen; and

the new arm choice policy is determined based at least in part on the first set of action-reward data that includes the first context data for the arm actions chosen; and

the one or more non-transitory media further comprises instructions which, when executed by the system, cause the system to perform:

receiving second context data as part of the request, and

wherein determining the particular arm action to take comprises determining the particular arm action to take based at least in part on the new arm choice policy and the second context data.

Assignments (4)
RELEASE OF SECURITY INTEREST Recorded Dec 13, 2024
From: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
To: DROPBOX, INC.
Reel/Frame 069635/0332 →
SECURITY INTEREST Recorded Dec 12, 2024
From: DROPBOX, INC.
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 069604/0611 →
PATENT SECURITY AGREEMENT Recorded Mar 10, 2021
From: DROPBOX, INC.
To: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 055670/0219 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 7, 2019
From: JAIN, ADITI; CHAWLA, MANVEER SINGH; BERG, THOMAS; ZAREKAR, SWAPNIL; KAJIC, ROBERT; JOHAR, KARANDEEP; FELDSTEIN, AARON; KIM, WALTER; NUDELL, JOE; DONG, JENNY; WILSON, JARED; THOMPSON, LUKE; KRIEGMAN, DAVID
To: DROPBOX, INC.
Reel/Frame 048271/0456 →