IP Library › Granted Patent US 12,511,677
Granted Patent B2
US 12,511,677 · App. 18/108,916 · Granted Dec 30, 2025

Automated policy function adjustment using reinforcement learning algorithm

Inventors: Tilman Drerup (Palo Alto, CA); Nour Alkhatib (Mississauga, CA); Jonathan Gu (San Francisco, CA); Amin Akbari (San Francisco, CA); Changyao Chen (New York, NY)
Assignee: Maplebear Inc.
G06Q30/0617G06N3/092
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,511,677
App. No.
18/108,916
Granted
Dec 30, 2025
Kind
B2
Abstract

An online system may receive, from a content provider, a content presentation campaign that includes one or more objectives. The online system may define a set of one or more policy functions that automatically controls the content presentation campaign. A policy function may control one or more criteria in bidding content slots. The online system may monitor a realized outcome of the content presentation campaign. The online system may apply a reinforcement learning algorithm in adjusting the set of policy functions. The reinforcement learning algorithm adjusts one or more parameters in the set of policy functions to reduce a difference between the realized outcome and the desired outcome set by the content provider. The online system generates an adjusted set of policy functions and uses the adjusted set of policy functions in bidding content slots to present one or more content items provided by the content provider.

Claims (56)

1 . A method for creating an adaptive autonomous system, the method comprising:

at an online concierge system comprising a processor and a computer-readable medium:

receiving, by the online concierge system from a computer system associated with a content provider, a content presentation campaign that includes one or more objectives set by the content provider, at least one of the objectives defining a desired outcome;

defining a set of one or more policy functions that automatically control the content presentation campaign, each policy function comprising one or more parameters and a plurality of states, each state corresponding to a previously used parameter configuration and an associated realized outcome of the content presentation campaign;

monitoring a realized outcome of the content presentation campaign that is controlled by the set of policy functions;

storing, by the online concierge system, data indicating the realized outcomes and the corresponding states of the policy functions associated with the campaign;

applying a reinforcement learning algorithm that autonomously adjusts one or more parameters of the policy functions using a state-aware counterfactual estimation process that references previously stored states and corresponding outcomes to reduce a difference between the realized outcomes and the desired outcome set by the content provider;

generating an adjusted set of policy functions by the reinforcement learning algorithm; and

using the adjusted set of policy functions to control subsequent content selection operations of the online concierge system, thereby enabling autonomous system adaptation to changing campaign performance.

2 . The method of claim 1 , wherein using the adjusted set of policy functions comprises inputting estimates of a content slot to a first policy function of the policy functions to generate a bid value.

3 . The method of claim 2 , wherein inputting estimates of a content slot to the first policy function to generate the bid value comprises:

generating a plurality of features related to the content slot;

inputting the plurality of features to a machine learning model to predict the estimate; and

using the estimates generated by the machine learning model as inputs to the first policy function.

4 . The method of claim 1 , wherein the set of policy functions comprises a plurality of policy functions, each policy function being defined based on an objective provided by the content provider.

5 . The method of claim 1 , wherein at least one of the policy functions comprises a plurality of states that record past actions and outcomes associated with the policy function.

6 . The method of claim 1 , wherein the reinforcement learning algorithm updates the set of policy functions through a counterfactual policy estimation.

7 . The method of claim 1 , wherein the reinforcement learning algorithm is heuristic based, and wherein applying the heuristic based reinforcement learning algorithm comprises:

defining a rule in adjusting a parameter in a policy function;

examining the policy function at a previous state that has a known realized outcome; and

adjusting the parameter based on the rule and the known realized outcome of the previous state.

8 . The method of claim 1 , wherein the reinforcement learning algorithm is a machine learning based.

9 . The method of claim 1 , wherein one of the content items in the content presentation campaign is a sponsored item offered on one or more interfaces hosted by the online concierge system.

10 . A non-transitory computer readable medium configured to store code comprising instructions, the instructions, when executed by one or more processors, cause the one or more processors to:

receive, by an online concierge system from a computer system associated with a content provider, a content presentation campaign that includes one or more objectives set by the content provider, at least one of the objectives defining a desired outcome;

define a set of one or more policy functions that automatically control the content presentation campaign, each policy function comprising one or more parameters and a plurality of states, each state corresponding to a previously used parameter configuration and an associated realized outcome of the content presentation campaign;

monitor a realized outcome of the content presentation campaign that is controlled by the set of policy functions;

store, by the online concierge system, data indicating the realized outcomes and the corresponding states of the policy functions associated with the campaign;

apply a reinforcement learning algorithm that autonomously adjusts one or more parameters of the policy functions using a state-aware counterfactual estimation process that references previously stored states and corresponding outcomes to reduce a difference between the realized outcomes and the desired outcome set by the content provider;

generate an adjusted set of policy functions by the reinforcement learning algorithm; and

use the adjusted set of policy functions to control subsequent content selection operations of the online concierge system, thereby enabling autonomous system adaptation to changing campaign performance.

11 . The non-transitory computer readable medium of claim 10 , wherein using the adjusted set of policy functions comprises inputting estimates of a content slot to a first policy function of the policy functions to generate a bid value.

12 . The non-transitory computer readable medium of claim 11 , inputting estimates of a content slot to the first policy function to generate the bid value comprises:

generating a plurality of features related to the content slot;

inputting the plurality of features to a machine learning model to predict the estimate; and

using the estimates generated by the machine learning model as inputs to the first policy function.

13 . The non-transitory computer readable medium of claim 10 , wherein the set of policy functions comprises a plurality of policy functions, each policy function being defined based on an objective provided by the content provider.

14 . The non-transitory computer readable medium of claim 10 , wherein at least one of the policy functions comprises a plurality of states that record past actions and outcomes associated with the policy function.

15 . The non-transitory computer readable medium of claim 10 , wherein the reinforcement learning algorithm updates the set of policy functions through a counterfactual policy estimation.

16 . The non-transitory computer readable medium of claim 10 , wherein the reinforcement learning algorithm is heuristic based, and wherein applying the heuristic based reinforcement learning algorithm comprises:

defining a rule in adjusting a parameter in a policy function;

examining the policy function at a previous state that has a known realized outcome; and

adjusting the parameter based on the rule and the known realized outcome of the previous state.

17 . The non-transitory computer readable medium of claim 10 , wherein the reinforcement learning algorithm is a machine learning based.

18 . The non-transitory computer readable medium of claim 10 , wherein one of the content items in the content presentation campaign is a sponsored item offered on one or more interfaces hosted by the online concierge system.

19 . An online concierge system comprising:

one or more processors; and

memory configured to store code comprising instructions, the instructions, when executed by the one or more processors, cause the one or more processors to:

receive, by an online concierge system from a computer system associated with a content provider, a content presentation campaign that includes one or more objectives set by the content provider, at least one of the objectives defining a desired outcome;

define a set of one or more policy functions that automatically control the content presentation campaign, each policy function comprising one or more parameters and a plurality of states, each state corresponding to a previously used parameter configuration and an associated realized outcome of the content presentation campaign;

monitor a realized outcome of the content presentation campaign that is controlled by the set of policy functions;

store, by the online concierge system, data indicating the realized outcomes and the corresponding states of the policy functions associated with the campaign;

apply a reinforcement learning algorithm that autonomously adjusts one or more parameters of the policy functions using a state-aware counterfactual estimation process that references previously stored states and corresponding outcomes to reduce a difference between the realized outcomes and the desired outcome set by the content provider;

generate an adjusted set of policy functions by the reinforcement learning algorithm; and

use the adjusted set of policy functions to control subsequent content selection operations of the online concierge system, thereby enabling autonomous system adaptation to changing campaign performance.

20 . The online concierge system of claim 19 , wherein using the adjusted set of policy functions comprises inputting estimates of a content slot to a first policy function of the policy functions to generate a bid value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 22, 2023
From: DRERUP, TILMAN; ALKHATIB, NOUR; GU, JONATHAN; AKBARI, AMIN; CHEN, CHANGYAO
To: MAPLEBEAR INC. (DBA INSTACART)
Reel/Frame 062764/0689 →
Continuity (2)
Provisional Application 63310022 · Feb 14, 2022
Related Publication 20230298080A1 · Sep 21, 2023
References Cited (7)
US 11715042B1 · Liu · 2023 [cited by examiner]
US 20160210655A1 · Cottle · 2016 [cited by examiner]
US 20210073647A1 · Hunter · 2021 [cited by examiner]
US 20210182976A1 · Poslavsky · 2021 [cited by examiner]
US 20220043742A1 · van Adelsberg · 2022 [cited by examiner]
US 20220044299A1 · Tate · 2022 [cited by examiner]
Vargese, N., & Mahmoud, Q., “A Survey of Multi-Task Deep Reinforcement Learning”, Electronics 9.9: 1363 MDPI AG, 2020 (Year: 2020). [cited by examiner]