IP Library › Granted Patent US 11,562,386
Granted Patent B2
US 11,562,386 · App. 15/795,821 · Granted Jan 24, 2023

Intelligent agent system and method

Inventor: Kari Saarenvirta (Ajax, CA)
Assignee: Daisy Intelligence Corporation
G06Q30/0206G06Q10/0637G06Q30/0204G06Q30/0207G06Q30/0247G06Q30/0252
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,562,386
App. No.
15/795,821
Granted
Jan 24, 2023
Kind
B2
Abstract

A computer system and computer-implemented method for retail merchandise planning, including promotional product selection, price optimization and planning. According to an embodiment, the computer system for generating an electronic retail plan for a retailer comprises, a data staging module configured to input retail sensory data from one or more computer systems associated with the retailer; a data processing module configured to pre-process the inputted retail sensory data; a data warehouse module configured to store the inputted retail sensory data and the pre-processed retail sensory data; a state model module configured to generate a retailer state model for modeling operation of the retailer based on the retail sensory data; a calibration module configured to calibrate the state model module according to one or more control parameters; and an output module for generating an electronic retail plan for the retailer based on the retailer state model.

Claims (72)

1. A computer-implemented method, comprising:

storing, at a data warehouse server, a plurality of product records;

providing an intelligent agent module at a control server in network communication with the data warehouse server, the intelligent agent module comprising:

a retail promotional model comprising a control policy and one or more coefficients;

a simulation component,

a sensor input component,

a memory component comprising a long-term memory representative of a long-term frequency response, the long-term memory storing a plurality of long-term sensor data and a plurality of long-term prior actions, and a short-term memory representative of a short-term frequency response, the short-term memory storing a plurality of short-term sensor data and a plurality of short-term prior actions, and

determining, at the intelligent agent module, one or more candidate itemsets in the plurality of product records, each of the one or more candidate itemsets identifying two or more products in the plurality of product records;

executing a reinforcement learning algorithm to simulate, at the simulation component, based on the short-term memory of the memory component, the long-term memory of the memory component, a simulated state of an environment, and the retail promotional model, an expected reward for each of the one or more candidate itemsets for the one or more time periods, each expected reward simulating a sales metric of the plurality of product records based on the respective candidate itemset;

selecting, at the intelligent agent module, using the control policy and the expected reward for the one or more candidate itemsets, one or more selected itemsets from the one or more candidate itemsets, each of the one or more selected itemsets identifying two or more products in the plurality of product records;

generating a current action corresponding to the one or more selected itemsets from the one or more candidate itemsets and each corresponding expected reward;

applying the current action to the simulated state of the environment;

receiving, at the sensor input component, sensor data comprising a measured reward, wherein the sensor data is representative of a measured state of the environment;

storing the sensor data in the plurality of short-term sensor data of the short-term memory;

storing the sensor data in the plurality of long-term sensor data of the long-term memory; and

correcting the simulated state of the environment based on the sensor data comprising the measured reward.

2. The method of claim 1 , wherein:

the plurality of product records further comprises a product category hierarchy;

the simulating the expected reward is for a first level of the product category hierarchy; and

wherein a first product belongs to the first level of the product category hierarchy and a first candidate itemset in the one or more candidate itemsets comprises the first product.

3. The method of claim 2 , wherein:

the simulating the expected reward further comprises simulating a second level of the product category hierarchy, the second level at a lower level than the first level in the product category hierarchy; and

wherein a second product belongs to the second level of the product category hierarchy and the first candidate itemset in the one or more candidate itemsets comprises the second product.

4. The method of claim 3 , wherein the simulating the expected reward further comprises:

simulating a first product set in the one or more selected itemsets, the first product set having a first product in the first level of the product category hierarchy and a second product in the second level of the product category hierarchy.

5. The method of claim 4 , wherein the simulating the expected reward further comprises:

determining one or more solution increments for the one or more time periods; and

simulating an addition or removal of product records to the one or more selected itemsets.

6. The method of claim 5 , further comprising:

receiving, from a retailer system, retail data for a current time period; and

updating the retail promotional model based on the retail data for the current time period and the one or more selected itemsets.

7. The method of claim 6 wherein the reinforcement learning algorithm comprises a genetic algorithm.

8. A computer-implemented system, comprising:

a data warehouse server, the data warehouse server comprising:

a first memory, the first memory comprising:

a plurality of product records;

a control server in network communication with the data warehouse server, the control server comprising:

a second memory, the second memory comprising:

an intelligent agent module, the intelligent agent module comprising:

a retail promotional model, the retail promotional model comprising:

 a control policy,

 one or more coefficients;

a simulation component,

a sensor input component,

a memory component comprising a long-term memory representative of a long-term frequency response, the long-term memory storing a plurality of long-term sensor data and a plurality of long-term prior actions, and a short-term memory representative of a short-term frequency response, the short-term memory storing a plurality of short-term sensor data and a plurality of short-term prior actions;

a network device;

a processor, the processor executing the intelligent agent module configured to:

determine one or more candidate itemsets in the plurality of product records, each of the one or more candidate itemsets identifying two or more products in the plurality of product records;

executing a reinforcement learning algorithm to simulate, at the simulation component, based on the short-term memory of the memory component, the long-term memory of the memory component, a simulated state of an environment, and the retail promotional model, an expected reward for each of the one or more candidate itemsets for the one or more time periods, each expected reward simulating a sales metric of the plurality of product records based on the respective candidate itemset;

select, using the control policy and the expected reward for the one or more candidate itemsets, one or more selected itemsets from the one or more candidate itemsets, each of the one or more selected itemsets identifying two or more products in the plurality of product records;

generate a current action corresponding to the one or more selected itemsets from the one or more candidate itemsets and each corresponding expected reward;

apply the current action to the simulated state of the environment;

receive, at the sensor component, sensor data comprising a measured reward, wherein the sensor data is representative of a measured state of the environment;

storing the sensor data in the plurality of short-term sensor data of the short-term memory;

storing the sensor data in the plurality of long-term sensor data of the long-term memory; and

correct the simulated state of the environment based on the sensor data comprising the measured reward.

9. The system of claim 8 , wherein:

the plurality of product records further comprises a product category hierarchy;

the simulating the expected reward is for a first level of the product category hierarchy; and

wherein a first product belongs to the first level of the product category hierarchy and a first candidate itemset in the one or more candidate itemsets comprises the first product.

10. The system of claim 9 , wherein:

the simulating the expected reward further comprises simulating a second level of the product category hierarchy, the second level at a lower level than the first level in the product category hierarchy; and

wherein a second product belongs to the second level of the product category hierarchy and the first candidate itemset in the one or more candidate itemsets comprises the second product.

11. The system of claim 10 , wherein the processor is further configured to simulate the expected reward by:

simulate a first product set in the one or more selected itemsets, the first product set having a first product in the first level of the product category hierarchy and a second product in the second level of the product category hierarchy.

12. The system of claim 11 , wherein the processor is further configured to simulate the expected reward by:

determine one or more solution increments for the time period; and

simulate an addition or removal of product records to the one or more selected itemsets.

13. The system of claim 12 , wherein the processor is further configured to:

receive, from a retailer system using the network device, retail data for the time period; and

update the retail promotional model based on the retail data and the one or more selected itemsets.

14. The system of claim 13 wherein the using reinforcement learning algorithm comprises a genetic algorithm.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 30, 2024
From: DAISY INTELLIGENCE CORPORATION
To: HARRIS & PARTNERS INC.
Reel/Frame 067097/0872 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 30, 2024
From: HARRIS & PARTNERS INC.
To: DAISY INTEL INC.
Reel/Frame 067097/0890 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 3, 2018
From: SAARENVIRTA, KARI
To: DAISY INTELLIGENCE CORPORATION
Reel/Frame 045420/0629 →
Priority Claims (1)
CA 2982930 · Oct 18, 2017 · national
Continuity (1)
Related Publication 20190114655A1 · Apr 18, 2019