IP Library Granted Patent US 11,200,505
Granted Patent B2
US 11,200,505 · App. 15/653,138 · Granted Dec 14, 2021

System and method for calculating search term probability

Inventors: Varun Srivastava (Sunnyvale, CA); Yiye Ruan (Columbus, OH); Yan Zheng (San Jose, CA)
Assignee: WALMART APOLLO, LLC
G06N7/005G06F16/951G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,200,505
App. No.
15/653,138
Granted
Dec 14, 2021
Kind
B2
Abstract

A system and method for predicting search term popularity is disclosed herein. A database system may comprise a first database cluster H and a second database cluster L. A machine learning algorithm is trained to create a predictive model. Thereafter, for each record in a database system, the predictive model is used to calculate a probability of the record being accessed. If the calculated probability of the record being accessed is greater than a threshold value, then the record in the first database cluster H; otherwise, the record is placed in the second database cluster L. Training the machine learning algorithm comprises inputting a training feature vector associated with the record into the machine learning algorithm, inputting a cost vector into the machine learning algorithm, and iteratively operating the machine learning algorithm on each record in the set of records to create a predictive model. Other embodiments are also disclosed herein.

Claims (92)

1. A system comprising:

one or more processors; and

one or more non-transitory memory storage devices storing computing instructions configured to run on the one or more processors and perform:

for each record in a set of distinct records in a database system:

inputting a training feature vector associated with the record into a machine learning algorithm, the training feature vector associated with the record comprising a list of characteristics of the record; and

inputting a cost vector associated with the record into the machine learning algorithm, the cost vector associated with the record configured to train the machine learning algorithm to reduce a probability of a false negative prediction for the record;

iteratively operating the machine learning algorithm on each record in the set of distinct records to train the machine learning algorithm to create a predictive model;

for each record of the set of distinct records:

using the predictive model to calculate a probability of the record being accessed;

when the probability of the record being accessed, as calculated, is greater than a threshold value, placing the record in a first database cluster H; and

when the probability of the record being accessed, as calculated, is not greater than the threshold value, placing the record in a second database cluster L; and

receiving a request from a requester for at least one record of the set of distinct records.

2. The system of claim 1 , wherein:

the database system comprises the first database cluster H and the second database cluster L; and

the one or more non-transitory memory storage devices storing the computing instructions are further configured to run on the one or more processors and perform:

presenting the at least one record from the set of distinct records to the requester in response to the request.

3. The system of claim 2 , wherein:

the threshold value is determined such that at least approximately 99 percent of predicted accesses will access records of the set of distinct records placed in the first database cluster H; and

using the predictive model to calculate the probability of the record being accessed comprises:

for each record in the set of distinct records:

inputting a list of prediction feature vectors, each prediction feature vector of the list of prediction feature vectors comprising the list of characteristics of the record; and

using the predictive model to analyze the list of prediction feature vectors to calculate the probability of the record being accessed.

4. The system of claim 1 , wherein the machine learning algorithm comprises at least one of: a decision tree, a bagging technique, a logistic regression, a perceptron, a support vector machine, or a relevance vector machine.

5. The system of claim 1 , wherein iteratively operating the machine learning algorithm to train the machine learning algorithm to create the predictive model comprises:

operating the machine learning algorithm on a periodic basis; and

for each record of the set of distinct records:

reviewing historical access data associated with the record; and

comparing the probability of the record being accessed, as calculated, with the historical access data associated with the record.

6. The system of claim 1 , wherein:

the training feature vector further comprises a label configured to indicate, for each record in the set of distinct records, when the record has been accessed within a pre-defined time period; and

the cost vector associated with the record represents an estimate of cost incurred when the record is placed in an incorrect database cluster of the first database cluster H or the second database cluster L.

7. A method implemented via execution of computer instructions configured to run on one or more processors and configured to be stored on one or more non-transitory memory storage devices, the method comprising:

for each record in a set of distinct records in a database system:

inputting a training feature vector associated with the record into a machine learning algorithm, the training feature vector associated with the record comprising a list of characteristics of the record; and

inputting a cost vector associated with the record into the machine learning algorithm, the cost vector associated with the record configured to train the machine learning algorithm to reduce a probability of a false negative prediction for the record;

iteratively operating the machine learning algorithm on each record in the set of distinct records to train the machine learning algorithm to create a predictive model;

for each record of the set of distinct records:

using the predictive model to calculate a probability of the record being accessed;

when the probability of the record being accessed, as calculated, is greater than a threshold value, placing the record in a first database cluster H; and

when the probability of the record being accessed, as calculated, is not greater than the threshold value, placing the record in a second database cluster L; and

receiving a request from a requester for at least one record of the set of distinct records.

8. The method of claim 7 , wherein:

the database system comprises the first database cluster H and the second database cluster L; and

the method further comprises:

presenting the at least one record from the set of distinct records to the requester in response to the request.

9. The method of claim 8 , wherein:

the threshold value is determined such that at least approximately 99 percent of predicted accesses will access records of the set of distinct records placed in the first database cluster H; and

using the predictive model to calculate the probability of the record being accessed comprises:

for each record in the set of distinct records:

inputting a list of prediction feature vectors, each prediction feature vector of the list of prediction feature vectors comprising the list of characteristics of the record; and

using the predictive model to analyze the list of prediction feature vectors to calculate the probability of the record being accessed.

10. The method of claim 7 , wherein the machine learning algorithm comprises at least one of: a decision tree, a bagging technique, a logistic regression, a perceptron, a support vector machine, or a relevance vector machine.

11. The method of claim 7 , wherein iteratively operating the machine learning algorithm to train the machine learning algorithm to create the predictive model comprises:

operating the machine learning algorithm on a periodic basis; and

for each record of the set of distinct records:

reviewing historical access data associated with the record; and

comparing the probability of the record being accessed, as calculated, with the historical access data associated with the record.

12. The method of claim 7 , wherein:

the training feature vector further comprises a label configured to indicate, for each record in the set of distinct records, when the record has been accessed within a pre-defined time period; and

the cost vector associated with the record represents an estimate of cost incurred when the record is placed in an incorrect database cluster of the first database cluster H or the second database cluster L.

13. A system comprising:

one or more processors; and

one or more non-transitory memory storage devices storing computing instructions configured to run on the one or more processors and perform:

training a machine learning algorithm to create a predictive model;

receiving, from a requesting party, a request to analyze a probability that a record of a database will be requested within a predetermined time period;

retrieving a feature vector corresponding to the record;

calculating a prediction of the probability that the record will be requested within the predetermined time period, the prediction being based on the predictive model used in conjunction with the feature vector;

when the prediction of the probability is greater than a threshold value, placing the record in a first database cluster H;

when the prediction of the probability is greater than the threshold value, placing the record in a second database cluster L; and

presenting, to the requesting party, the prediction of the probability, as calculated, and a storage location of the record.

14. The system of claim 13 , wherein training the machine learning algorithm to create the predictive model comprises:

for each record in a set of distinct records in the database:

inputting a training feature vector associated with the record into the machine learning algorithm, the training feature vector comprising a list of characteristics of the record; and

inputting a cost vector associated with the record into the machine learning algorithm, the cost vector configured to train the machine learning algorithm to reduce a probability of a false negative prediction for the record; and

iteratively operating the machine learning algorithm on each record in the set of distinct records to create the predictive model.

15. The system of claim 13 , wherein the machine learning algorithm uses a MetaCost algorithm and a cost-insensitive machine learning algorithm used in conjunction with the MetaCost algorithm.

16. The system of claim 13 , wherein the feature vector corresponding to the record comprises a list of characteristics of the record.

17. A method implemented via execution of computer instructions configured to run on one or more processors and configured to be stored on one or more non-transitory memory storage devices, the method comprising:

training a machine learning algorithm to create a predictive model;

receiving, from a requesting party, a request to analyze a probability that a record of a database will be requested within a predetermined time period;

retrieving a feature vector corresponding to the record;

calculating a prediction of the probability that the record will be requested within the predetermined time period, the prediction being based on the predictive model used in conjunction with the feature vector;

when the prediction of the probability is greater than a threshold value, placing the record in a first database cluster H;

when the prediction of the probability is greater than the threshold value, placing the record in a second database cluster L; and

presenting, to the requesting party, the prediction of the probability, as calculated, and a storage location of the record.

18. The method of claim 17 , wherein training the machine learning algorithm to create the predictive model comprises:

for each record in a set of distinct records in the database:

inputting a training feature vector associated with the record into the machine learning algorithm, the training feature vector comprising a list of characteristics of the record; and

inputting a cost vector associated with the record into the machine learning algorithm, the cost vector configured to train the machine learning algorithm to reduce a probability of a false negative prediction for the record; and

iteratively operating the machine learning algorithm on each record in the set of distinct records to create the predictive model.

19. The method of claim 17 , wherein the machine learning algorithm uses a MetaCost algorithm and a cost-insensitive machine learning algorithm used in conjunction with the MetaCost algorithm.

20. The method of claim 17 , wherein the feature vector corresponding to the record comprises a list of characteristics of the record.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 2, 2018
From: WAL-MART STORES, INC.
To: WALMART APOLLO, LLC
Reel/Frame 045817/0115 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 19, 2017
From: SRIVASTAVA, VARUN; RUAN, YIYE; ZHENG, YAN
To: WAL-MART STORES, INC.
Reel/Frame 043048/0283 →
Continuity (2)
Continuation 14498170 · Sep 26, 2014
Related Publication 20170316334A1 · Nov 2, 2017