IP Library › Granted Patent US 12,481,956
Granted Patent B2
US 12,481,956 · App. 18/368,916 · Granted Nov 25, 2025

System and method of reinforced machine-learning retail allocation

Inventors: Ganesh Muthusamy (Hyderabad, IN); Sudhakar Jayapal (Hyderabad, IN); Karthik Kondapaneni (Hyderabad, IN); Rajneesh Kumar Agrawal (Bangalore, IN)
Assignee: Blue Yonder Group, Inc.
G06Q10/087G06N5/04G06N20/00G06Q10/06313G06Q10/067G06Q10/08G06Q30/0202
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,481,956
App. No.
18/368,916
Granted
Nov 25, 2025
Kind
B2
Abstract

A system and method for allocation planning comprise a server comprising a processor and memory and configured to calculate a reward for a historical allocation of a product to one or more stores associated with a retailer. Embodiments include simulating what-if scenarios for the historical allocation to identify an allocation having a greater reward than the historical allocation and allocating a quantity of a product for a current allocation to the one or more stores based, at least in part, on a distance calculation of one or more independent variables for the historical allocation and the current allocation and the identified allocation having the greater reward then the historical allocation.

Claims (47)

1 . A system of allocation planning, comprising:

a network, automated warehousing equipment and a server, comprising a processor and memory, the server operably connected over the network to the automated warehousing equipment, the server further configured to:

calculate, using a reward-penalty function as part of a reinforcement learning process, a reward for a historical allocation of a product to one or more stores associated with a retailer, wherein the reward-penalty function comprises product margin, inventory carrying cost and opportunity cost represented by Bellman's equation, and wherein each of one or more states in the reinforcement learning process comprise a state of inventory at different time points in an allocation cycle and for an allocation horizon comprising a predetermined number of states;

calculate a constrained allocation quantity based on the reward for the historical allocation of the product;

allocate a quantity of a product for a current allocation to the one or more stores based, at least in part, on the constrained allocation quantity; and

responsive to a difference between a current inventory level and the constrained allocation quantity, retrieve a quantity of the product equal to the difference between the current inventory level and the constrained allocation quantity for transportation to stores by sending instructions over the network to the automated warehousing equipment of one or more distribution centers to automatically retrieve the quantity of the product.

2 . The system of claim 1 , wherein the server is further configured to:

calculate a pre-season buy based at least in part on the constrained allocation quantity.

3 . The system of claim 1 , wherein the server is further configured to:

calculate the constrained allocation quantity for the one or more stores based, at least in part, on an overall profit across all stores of the retailer to reduce mark down impact and to reduce stock-outs.

4 . The system of claim 1 , wherein the server is further configured to:

retrieve the quantity of the product according to one or more of:

a minimum order quantity, a maximum order quantity, a discount, a step-size order quantity and one or more batch quantity rules.

5 . The system of claim 1 , wherein the one or more stores comprise one or more nodes of a multi-echelon supply chain.

6 . The system of claim 1 , wherein the reinforcement learning process further comprises a finite set of states, each state of the finite set of states comprising an amount of stock of the product at a particular time for one or more inventory locations.

7 . The system of claim 6 , wherein transitions between states occur when a demand leads to a sale and inventory increases by receipt of allocated products.

8 . A computer-implemented method of allocation planning, comprising:

networking a computer with automated warehousing equipment;

calculating, by the computer comprising a processor and memory, using a reward-penalty function as part of a reinforcement learning process, a reward for a historical allocation of a product to one or more stores associated with a retailer, wherein the reward-penalty function comprises product margin, inventory carrying cost and opportunity cost represented by Bellman's equation, and wherein each of one or more states in the reinforcement learning process comprise a state of inventory at different time points in an allocation cycle and for an allocation horizon comprising a predetermined number of states;

calculating, by the computer, a constrained allocation quantity based on the reward for the historical allocation of the product;

allocating, by the computer, a quantity of a product for a current allocation to the one or more stores based, at least in part, on the constrained allocation quantity; and

responsive to a difference between a current inventory level and the constrained allocation quantity, retrieve a quantity of the product equal to the difference between the current inventory level and the constrained allocation quantity for transportation to stores by sending instructions over the network to the automated warehousing equipment of one or more distribution centers to automatically retrieve the quantity of the product.

9 . The computer-implemented method of claim 8 , further comprising:

calculating, by the computer a pre-season buy based at least in part on the constrained allocation quantity.

10 . The computer-implemented method of claim 8 , further comprising:

calculating, by the computer, the constrained allocation quantity for the one or more stores based, at least in part, on an overall profit across all stores of the retailer to reduce mark down impact and to reduce stock-outs.

11 . The computer-implemented method of claim 8 , further comprising:

retrieving, by the computer, the quantity of the product according to one or more of:

a minimum order quantity, a maximum order quantity, a discount, a step-size order quantity and one or more batch quantity rules.

12 . The computer-implemented method of claim 8 , wherein the one or more stores comprise one or more nodes of a multi-echelon supply chain.

13 . The computer-implemented method of claim 8 , wherein the reinforcement learning process further comprises a finite set of states, each state of the finite set of states comprising an amount of stock of the product at a particular time for one or more inventory locations.

14 . The computer-implemented method of claim 13 , wherein transitions between states occur when a demand leads to a sale and inventory increases by receipt of allocated products.

15 . A non-transitory computer-readable medium embodied with software, the software when executed:

networks a computer with automated warehousing equipment;

calculates, using a reward-penalty function as part of a reinforcement learning process, a reward for a historical allocation of a product to one or more stores associated with a retailer, wherein the reward-penalty function comprises product margin, inventory carrying cost and opportunity cost represented by Bellman's equation, and wherein each of one or more states in the reinforcement learning process comprise a state of inventory at different time points in an allocation cycle and for an allocation horizon comprising a predetermined number of states;

calculates a constrained allocation quantity based on the reward for the historical allocation of the product;

allocates a quantity of a product for a current allocation to the one or more stores based, at least in part, on the constrained allocation quantity; and

responsive to a difference between a current inventory level and the constrained allocation quantity, retrieves a quantity of the product equal to the difference between the current inventory level and the constrained allocation quantity for transportation to stores by sending instructions over the network to the automated warehousing equipment of one or more distribution centers to automatically retrieve the quantity of the product.

16 . The non-transitory computer-readable medium of claim 15 , the software when executed further:

calculates a pre-season buy based at least in part on the constrained allocation quantity.

17 . The non-transitory computer-readable medium of claim 15 , the software when executed further:

calculates the constrained allocation quantity for the one or more stores based, at least in part, on an overall profit across all stores of the retailer to reduce mark down impact and to reduce stock-outs.

18 . The non-transitory computer-readable medium of claim 15 , the software when executed further:

retrieves the quantity of the product according to one or more of:

a minimum order quantity, a maximum order quantity, a discount, a step-size order quantity and one or more batch quantity rules.

19 . The non-transitory computer-readable medium of claim 15 , wherein the one or more stores comprise one or more nodes of a multi-echelon supply chain.

20 . The non-transitory computer-readable medium of claim 15 , wherein the reinforcement learning process further comprises a finite set of states, each state of the finite set of states comprising an amount of stock of the product at a particular time for one or more inventory locations.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 19, 2023
From: MUTHUSAMY, GANESH; JAYAPAL, SUDHAKAR; KONDAPANENI, KARTHIK; AGRAWAL, RAJNEESH KUMAR
To: BLUE YONDER GROUP, INC.
Reel/Frame 064955/0988 →
Continuity (3)
Continuation 17119591 · Dec 11, 2020
Provisional Application 62947971 · Dec 13, 2019
Related Publication 20240005237A1 · Jan 4, 2024
References Cited (9)
US 7092929B1 · Dvorak et al. · 2006 [cited by applicant]
US 10423923B2 · Harsha et al. · 2019 [cited by applicant]
US 11321650B2 · Anandan Kartha et al. · 2022 [cited by applicant]
US 20170185933A1 · Adulyasak · 2017 [cited by examiner]
US 20200126015A1 · Anandan Kartha · 2020 [cited by examiner]
WO WO2019001120A1 · 2019 [cited by examiner]
Hutse, et al. Reinforcement Learning for Inventory Optimisation in Multi-Echelon Supply Chains, Master in business engineering, Ghent University (2019) (Year: 2019). [cited by examiner]
Sui, et al., A Reinforcement Learning Approach for Inventory Replenishment in Vendor-Managed Inventory Systems With Consignment Inventory, 22 Engineering Management Journal 44 (2010) (Year: 2010). [cited by examiner]
Sui et al., “A Reinforcement Learning Approach for Inventory Replenishment in Vendor-Managed Inventory Systems With Consignment Inventory,” Engineering Management Journal, Dec. 2010, vol. 22 No. 4 pp. 44-53, (Year: 2010… [cited by applicant]