IP Library Granted Patent US 11,171,909
Granted Patent B2
US 11,171,909 · App. 16/789,200 · Granted Nov 9, 2021

Delayed processing for arm policy determination for content management system messaging

Inventors: Aditi Jain (San Francisco, CA); Manveer Singh Chawla (San Francisco, CA); Thomas Berg (San Francisco, CA); Swapnil Zarekar (San Francisco, CA); Robert Kajic (San Francisco, CA); Karandeep Johar (San Francisco, CA); Aaron Feldstein (San Francisco, CA); Walter Kim (San Francisco, CA); Joe Nudell (San Francisco, CA); Jenny Dong (San Francisco, CA); Jared Wilson (San Francisco, CA); Luke Thompson (San Francisco, CA); David Kriegman (San Francisco, CA)
Assignee: Dropbox, Inc.
H04L51/26G06F9/5038H04L51/36H04L63/102H04L67/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,171,909
App. No.
16/789,200
Granted
Nov 9, 2021
Kind
B2
Abstract

Techniques are provided for delayed processing for arm policy determination for content management system messaging, including, during a delayed processing window, receiving reward data for arm actions taken, where the arm actions were chosen based on a previous version of an arm choice policy, and the previous version of the arm choice policy was determined based on a previous set of reward data for a previous set of arm actions taken. When the delayed processing window has closed, a new arm choice policy is determined based at least in part on the action-reward data, and the previous set of reward data and/or the previous arm choice policy. After a request to choose an arm choice is received, a particular arm action to take is determined based on the new arm choice policy. This chosen arm is provided in response to the request.

Claims (80)

1. A method comprising:

during a first time period:

making a first set of one or more communication decisions based on a first communication decision arm choice policy, and

receiving a first set of action-reward data that indicates a first set of one or more outcomes of the first set of one or more communication decisions,

wherein the first communication decision arm choice policy was determined based, at least in part, on a second set of action-reward data that indicates a second set of one or more outcomes of a second set of one or more communication decisions made previous to said making the first set of one or more communication decisions;

after the first time period, determining a second communication decision arm choice policy based, at least in part, on the first set of action-reward data;

receiving a request to make a communication decision;

determining a particular communication decision based, at least in part, on the second communication decision arm choice policy; and

providing the particular communication decision as a response to the request;

wherein the method is performed by one or more computing devices.

2. The method of claim 1 , wherein:

the request is associated with a particular user;

the particular communication decision includes one or more communication features; and

after providing the particular communication decision as a response to the request, a communication, having the one or more communication features, is sent to the particular user.

3. The method of claim 1 , further comprising:

during the first time period, receiving multiple requests wherein each request of the multiple requests is for a respective communication decision; and

after the first time period, determining one or more communication decisions for each request, of the multiple requests, based at least in part on the second communication decision arm choice policy.

4. The method of claim 1 , wherein determining the particular communication decision based, at least in part, on the second communication decision arm choice policy comprises determining the particular communication decision based, at least in part, on statistical variance in the second communication decision arm choice policy.

5. The method of claim 1 , wherein:

the second set of action-reward data comprises first reward data for a first communication decision of the second set of one or more communication decisions; and

the first reward data was determined based, at least in part, on passage of a particular timeout period after a communication, based on the first communication decision, was delivered.

6. The method of claim 1 , further comprising:

determining multiple, ranked communication decisions;

wherein determining the particular communication decision is further based, at least in part, on a ranking of the multiple, ranked communication decisions.

7. The method of claim 1 , wherein:

the first set of action-reward data comprises first context data for one or more communication decisions of the second set of one or more communication decisions; and

the second communication decision arm choice policy is determined based, at least in part, on the first set of action-reward data that includes the first context data; and

the method further comprises:

receiving second context data as part of the request, and

wherein determining the particular communication decision comprises determining the particular communication decision based, at least in part, on the second communication decision arm choice policy and the second context data.

8. The method of claim 1 , wherein a communication decision comprises determining at least one of:

a time for sending an electronic communication to one or more users; or

a type of electronic communication to send to one or more users.

9. A system comprising:

one or more computing devices;

memory; and

instructions, stored in the memory, and which, when executed by the system, cause the system to perform:

receiving, during a batch window, a first set of reward data for a first set of communication decisions, wherein the first set of communication decisions were chosen based on a first communication decision arm policy, and wherein the first communication decision arm policy was chosen based on a second set of reward data received prior to the batch window;

after the batch window, determining a new communication decision arm policy based, at least in part, on the first set of reward data;

receiving a request for a communication decision;

determining a particular communication decision based, at least in part, on the new communication decision arm policy; and

providing the particular communication decision as a response to the request.

10. The system of claim 9 , wherein:

the request is associated with a particular user;

the particular communication decision includes one or more communication features; and

after providing the particular communication decision as a response to the request, a communication, having the one or more communication features, is sent to the particular user.

11. The system of claim 9 , further comprising instructions which, when executed by the system, cause the system to perform:

during the batch window, receiving multiple requests, wherein each request of the multiple requests is for a respective communication decision; and

after the batch window, determining one or more corresponding communication decisions for each request, of the multiple requests, based at least in part on the new communication decision arm policy.

12. The system of claim 9 , wherein:

the first set of reward data comprises first reward data for a first communication decision of the first set of communication decisions; and

the first reward data was determined based, at least in part, on passage of a particular timeout period after a communication, based on the first communication decision, was delivered.

13. The system of claim 9 , further comprising instructions which, when executed by the system, cause the system to perform:

determining multiple, ranked communication decisions;

wherein determining the particular communication decision is further based, at least in part, on a ranking of the multiple, ranked communication decisions.

14. The system of claim 9 , wherein:

the first set of reward data comprises first context data for one or more communication decisions of the first set of communication decisions; and

the new communication decision arm policy is determined based at least in part on the first set of reward data that includes the first context data; and

the system further comprises instructions which, when executed by the system, cause the system to perform:

receiving second context data as part of the request, and

wherein determining the particular communication decision comprises determining the particular communication decision based at least in part on the new communication decision arm policy and the second context data.

15. One or more non-transitory media comprising instructions which, when executed by a system having one or more computing devices, cause the system to perform:

during a batch window time period, receiving a first set of action-reward data, wherein the first set of action-reward data is associated with a first set of communication decisions chosen based on a first communication decision arm policy, and wherein the first communication decision arm policy was determined based, at least in part on, a second set of action-reward data for a second set of communication decisions taken previous to the first set of communication decisions;

after the batch window time period, determining a new communication decision arm policy based at least in part on the first set of action-reward data;

determining to choose a communication decision arm from among communication decision arms in the new communication decision arm policy;

in response to determining to choose a communication decision arm, determining a particular communication decision based, at least in part, on the new communication decision arm policy.

16. The one or more non-transitory media of claim 15 , wherein:

determining to choose a communication decision arm from among communication decision arms in the new communication decision arm policy is based on a request that is associated with a particular user;

the particular communication decision includes one or more communication features; and

after providing the particular communication decision as a response to the request, a communication, having the one or more communication features, is sent to the particular user.

17. The one or more non-transitory media of claim 15 , further comprising instructions which, when executed by the system, cause the system to perform:

during the batch window time period, receiving multiple requests wherein each request, of the multiple requests, is for a respective communication decision; and

after the batch window time period, determining one or more corresponding communication decisions for each request, of the multiple requests, based, at least in part, on the new communication decision arm policy.

18. The one or more non-transitory media of claim 15 , wherein determining the particular communication decision based at least in part on the new communication decision arm policy comprises using statistical variance in the new communication decision arm policy during the determining.

19. The one or more non-transitory media of claim 15 , wherein:

the first set of action-reward data comprises first action-reward data for a first communication decision; and

the first action-reward data was determined based, at least in part, on passage of a particular timeout period after a communication, based on the first communication decision, was delivered.

20. The one or more non-transitory media of claim 15 , further comprising instructions which, when executed by the system, cause the system to perform:

determining multiple, ranked communication decisions; and

wherein determining the particular communication decision is further based, at least in part, on a ranking of the multiple, ranked arms.

Assignments (4)
RELEASE OF SECURITY INTEREST Recorded Dec 13, 2024
From: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
To: DROPBOX, INC.
Reel/Frame 069635/0332 →
SECURITY INTEREST Recorded Dec 12, 2024
From: DROPBOX, INC.
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 069604/0611 →
PATENT SECURITY AGREEMENT Recorded Mar 10, 2021
From: DROPBOX, INC.
To: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 055670/0219 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 12, 2020
From: JAIN, ADITI; CHAWLA, MANVEER SINGH; BERG, THOMAS; ZAREKAR, SWAPNIL; KAJIC, ROBERT; JOHAR, KARANDEEP; FELDSTEIN, AARON; KIM, WALTER; NUDELL, JOE; DONG, JENNY; WILSON, JARED; THOMPSON, LUKE; KRIEGMAN, DAVID
To: DROPBOX, INC.
Reel/Frame 051804/0964 →
Continuity (2)
Continuation 15793787 · Oct 25, 2017
Related Publication 20200259777A1 · Aug 13, 2020