IP Library Granted Patent US 11,645,679
Granted Patent B2
US 11,645,679 · App. 17/379,959 · Granted May 9, 2023

Ad exchange bid optimization with reinforcement learning

Inventors: Danny Portman (Atlanta, GA); Zachary D Jones (Atlanta, GA); David Rose (Atlanta, GA)
Assignee: Zeta Global Corp.
G06Q30/0275G06N3/0454G06N3/084G06Q30/0246G06Q30/0205G06Q30/0249G06Q30/0256G06Q30/0276
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,645,679
App. No.
17/379,959
Granted
May 9, 2023
Kind
B2
Abstract

A system for training a bidding model comprising: a plurality of tactics stored on at least one database; a plurality of hyperparameters; in response to an available inventory from a publisher relayed through a real time bid server, computing a bid on the available inventory; sending the bid to the real time bid server; receiving an auction result in response to the bid; calculating a plurality of rewards based on the auction result and the tactics; calculate a plurality of q values based on the rewards; calculate a plurality of losses; backpropogating the losses through the bidding model.

Claims (11)

1. A system for training a bidding model by reinforcement learning using first and second neural networks, the system configured to:

perform operations, in response to an available inventory from a publisher relayed through a real time bid server, including processing a bid on the available inventory, the processing including:

sending the bid to the real time bid server;

receiving a bid result in response to the bid;

storing state data including a sequence of bids sent to the real time bid server, the bid result, and a response rate for the available inventory;

using the first neural network, determining a plurality of target action Q-values based on the state data, the plurality of target action Q-values including at least one Q-value for each possible action at a current state of the bid server;

selecting an action based on a maximum target action Q-value;

using a second machine learning model, determining a target Q-value for the selected action based on the state data and experience data, the experience data including the selected action and a reward earned for the selected action;

training the first neural network to update the plurality of target action Q-values based on the target Q-value, the training using a stochastic gradient descent;

calculating a plurality of losses; and

back propagating the losses through the bidding model.

Assignments (2)
NOTICE OF GRANT OF SECURITY INTEREST IN PATENTS Recorded Aug 30, 2024
From: ZETA GLOBAL CORP.; ZSTREAM ACQUISITION LLC
To: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 068822/0154 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 20, 2022
From: PORTMAN, DANNY; JONES, ZACHARY D; ROSE, DAVID
To: ZETA GLOBAL CORP.
Reel/Frame 058704/0604 →