Systems and methods for determining substitutions
A system including one or more processors and one or more non-transitory computer-readable media storing computing instructions that, when executed on the one or more processors, cause the one or more processors to perform operations: training, using labeled training data and a list of substitutes for an item, a machine learning algorithm; determining, using the machine learning algorithm, as trained, a respective similarity score for each substitute of the list of substitutes; ranking each substitute of the list of substitutes based on its respective similarity score; and re-training the machine learning algorithm based on at least the labeled training data and a highest ranked substitute of the list of substitutes. Other embodiments are disclosed herein.
1 . A system comprising:
one or more processors; and
one or more non-transitory computer-readable media storing computing instructions that, when executed on the one or more processors, cause the one or more processors to perform operations comprising:
training, using labeled training data, a machine learning algorithm, wherein the labeled training data comprises positive and negative historical acceptance data, regarding interactions with a graphical user interface (GUI) via one or more devices, that includes at least one of:
a probability of a similarity between a historical title for a historical element of the GUI and a historical description of a historical substitute,
a taxonomy difference for the historical element and the historical substitute,
a difference between the historical element and the historical substitute, or
a normalized rank comparison for the historical element and the historical substitute;
using, based on current data and without requiring use of historical substitute data for an element, the machine learning algorithm, as trained, to determine machine learning outputs for a list of substitutes for the element; and
re-training the machine learning algorithm based on at least the labeled training data and a highest ranked substitute from a ranking of the list of substitutes based on the machine learning outputs.
2 . The system of claim 1 , wherein the computing instructions, when executed on the one or more processors, further cause the one or more processors to perform an operation comprising:
determining the list of substitutes comprising:
accessing an item taxonomy database comprising a respective item taxonomy for each item in a catalog of items;
identifying a specific item taxonomy for the item;
filtering out non-matching items of the catalog of items, wherein the non-matching items comprise a different item taxonomy than the specific item taxonomy of the item; and
adding matching items of the catalog of items to the list of substitutes, wherein the matching items of the catalog of items comprise the specific item taxonomy.
3 . The system of claim 1 , wherein the machine learning outputs are based on comparisons of a respective attribute for each substitute of the list of substitutes and an attribute of the element.
4 . The system of claim 1 , wherein the operations further comprise:
storing information, based on a selection of the highest ranked substitute, as additional training data for the labeled training data.
5 . The system of claim 1 , wherein the operations further comprise:
determining the ranking of the list of substitutes comprises by:
when two or more substitutes of the list of substitutes have approximately similar final scores, determining a quantity ratio comprising a ratio of a quantity for the element to a respective quantity for each substitute of the list of substitutes, wherein:
the quantity for the element comprises:
either (i) a weight of the element or (ii) a volume of the element, divided by a count of the element; and
the respective quantity for each substitute of the list of substitutes comprises:
either (i) a respective weight of each substitute of the list of substitutes or (ii) a respective volume of each substitute of the list of substitutes, divided by a respective count of each substitute of the list of substitutes.
6 . The system of claim 1 , wherein the computing instructions, when executed on the one or more processors, further cause the one or more processors to perform an operation comprising:
determining a respective historical substitution score comprising:
determining a respective number of successful substitutions for each substitute of the list of substitutes;
determining a respective number of unsuccessful substitutions for each substitute of the list of substitutes;
calculating a respective historical acceptance rate for each substitute of the list of substitutes using the respective number of successful substitutions for each substitute of the list of substitutes and the respective number of unsuccessful substitutions for each substitute of the list of substitutes; and
assigning the respective historical substitution score based on the respective historical acceptance rate for each substitution substitute of the list of substitutes and the respective number of successful substitutions for each substitute of the list of substitutes.
7 . The system of claim 1 , wherein the computing instructions, when executed on the one or more processors, further cause the one or more processors to perform an operation comprising:
inputting a respective similarity score for each substitute of the list of substitutes and a respective historical substitution score for each substitute of the list of substitutes into a feed-forward neural network comprising one or more rectifiers having ReLU non-linearity,
wherein the feed-forward neural network is trained without unsupervised pre-training.
8 . The system of claim 1 , wherein the computing instructions, when executed on the one or more processors, further cause the one or more processors to perform operations comprising:
determining respective qualities for each substitute of the list of substitutes;
facilitating a display, on a user interface of a user device, of the highest ranked substitute of the list of substitutes;
receiving, from the user interface of the user device, a selection of the highest ranked substitute;
after receiving the selection of the highest ranked substitute, substituting the highest ranked substitute;
receiving a respective title for each substitute of the list of substitutes; and
identifying a respective brand for each substitute of the list of substitutes using the respective title for each substitute of the list of substitutes.
9 . The system of claim 1 , wherein the computing instructions, when executed on the one or more processors, further cause the one or more processors to perform operations comprising:
when an item, corresponding to the element, is out of stock, comparing a respective dietary restriction for each substitute of the list of substitutes with a dietary restriction of the item; and
when a dietary restriction of a substitute of a list of substitutions does not match the dietary restriction of the item, removing the substitute of the list of substitutions from the list of substitutes.
10 . A method being implemented via execution of computing instructions configured to run on one or more processors and stored at one or more non-transitory computer-readable media, the method comprising:
training, using labeled training data, a machine learning algorithm, wherein the labeled training data comprises positive and negative historical acceptance data that includes at least one of:
a probability of a similarity between a historical title for a historical item and a historical description of a historical substitute,
a taxonomy difference for the historical item and the historical substitute,
a difference between the historical item and the historical substitute, or
a normalized rank comparison for the historical item and the historical substitute;
using, based on current data and without requiring use of historical substitute data for an item, the machine learning algorithm, as trained, to determine a ranking of a list of substitutes for the item; and
re-training the machine learning algorithm based on at least the labeled training data and a highest ranked substitute from the ranking of the list of substitutes.
11 . The method of claim 10 , further comprising:
determining the list of substitutes for the item comprising:
accessing an item taxonomy database comprising a respective item taxonomy for each item in a catalog of items;
identifying a specific item taxonomy for the item;
filtering out non-matching items of the catalog of items, wherein the non-matching items comprise a different item taxonomy than the specific item taxonomy of the item; and
adding matching items of the catalog of items to the list of substitutes, wherein the matching items of the catalog of items comprise the specific item taxonomy.
12 . The method of claim 10 , further comprising:
determining a respective title similarity score for each substitute of the list of substitutes by comparing:
a respective brand for each substitute of the list of substitutes;
a respective attribute for each substitute of the list of substitutes and an attribute of the item; and
a respective product for each substitute of the list of substitutes and a product of the item.
13 . The method of claim 10 , further comprising:
storing information, based on a selection of the highest ranked substitute, as additional training data for the labeled training data.
14 . The method of claim 10 , wherein ranking each substitute of the list of substitutes further comprises further comprising:
when two or more substitutes of the list of substitutes have approximately similar final scores, determining the ranking of the list of substitutes based on a quantity ratio comprising a ratio of a quantity for the item to a respective quantity for each substitute of the list of substitutes, wherein:
the quantity for the item comprises:
either (i) a weight of the item or (ii) a volume of the item, divided by a count of the item; and
the respective quantity for each substitute of the list of substitutes comprises:
either (i) a respective weight of each substitute of the list of substitutes or (ii) a respective volume of each substitute of the list of substitutes, divided by a respective count of each substitute of the list of substitutes.
15 . The method of claim 10 further comprising:
determining a respective historical substitution score comprising:
determining a respective number of successful substitutions for each substitute of the list of substitutes;
determining a respective number of unsuccessful substitutions for each substitute of the list of substitutes;
calculating a respective historical acceptance rate for each substitute of the list of substitutes using the respective number of successful substitutions for each substitute of the list of substitutes and the respective number of unsuccessful substitutions for each substitute of the list of substitutes; and
assigning the respective historical substitution score based on the respective historical acceptance rate for each substitution substitute of the list of substitutes and the respective number of successful substitutions for each substitute of the list of substitutes.
16 . The method of claim 10 , further comprising:
inputting a respective similarity score for each substitute of the list of substitutes and a respective historical substitution score for each substitute of the list of substitutes into a feed-forward neural network comprising one or more rectifiers having ReLU non-linearity, wherein the feed-forward neural network is trained without unsupervised pre-training.
17 . The method of claim 10 , further comprising:
facilitating a display, on a user interface of a user device, of information identifying the highest ranked substitute of the list of substitutes;
receiving, from the user interface of the user device, a selection of the highest ranked substitute of the list of substitutes; and
after receiving the selection of the highest ranked substitute, substituting the highest ranked substitute for the item.
18 . A non-transitory, computer-readable medium comprising instructions that, when executed by a processing resource, cause the processing resource to:
train, using labeled training data, a machine learning algorithm, wherein the labeled training data comprises positive and negative historical acceptance data that includes at least one of:
a probability of a similarity between a historical title for a historical item and a historical description of a historical substitute,
a taxonomy difference for the historical item and the historical substitute,
a difference between the historical item and the historical substitute, or
a normalized rank comparison for the historical item and the historical substitute;
use the machine learning algorithm, as trained, to determine a ranking of a list of substitutes for an item without requiring use of historical substitute data for the item; and
re-train the machine learning algorithm based on at least the labeled training data and a highest ranked substitute from the ranking of the list of substitutes.
19 . The non-transitory, computer-readable medium of claim 18 , wherein the instructions are further to cause the processing resource to:
determining the list of substitutes for the item by accessing an item taxonomy database comprising a respective item taxonomy for each item in a catalog of items.
20 . The non-transitory, computer-readable medium of claim 18 , wherein the positive and negative historical acceptance data includes the probability of the similarity between the historical title for the historical item and the historical description of the historical substitute.