System, method, and non-transitory machine-readable medium for self-guided sequence selection and extrapolation
Embodiments described herein provide systems and methods for training a sequential recommendation model. Methods include determining a difficulty and quality (DQ) score associated with user behavior sequences from a training dataset. User behavior sequences are sampled during training based on their DQ scores. A meta-extrapolator may also be trained based on user behavior sequences sampled according to DQ score. The meta-extrapolator may be trained with high quality low difficulty sequences. The meta-extrapolator may then be used with an input of high quality high difficulty sequences to generate synthetic user behavior sequences. The synthetic user behavior sequences may be used to augment the training dataset to fine-tune the sequential recommendation model, while continuing to sample user behavior sequences based on DQ score. As the DQ score is based on current model predictions, DQ scores iteratively update during the training process.
1 . A method for training a sequential recommendation model to automatically generate a sequential recommendation of multiple items in a sequence to a user, comprising:
receiving, via a communication interface, a training dataset comprising a plurality of user behavior sequences;
augmenting the training dataset with a plurality of synthetic user behavior sequences generated by a trained extrapolator based on a subset of user behavior sequences of the plurality of user behavior sequences;
determining, via a DQ score generator, for at least one user behavior sequence from the training dataset:
a difficulty score with an inverse relation to a probability of a sequential recommendation model with a first set of parameters correctly recommending an item in the at least one user behavior sequence, and
a quality score based on a first prediction generated by the sequential recommendation model with the first set of parameters from the at least one user behavior sequence;
determining, via the DQ score generator, a respective combined difficulty and quality (DQ) score for each of the plurality of user behavior sequences based on the respective difficulty and quality scores;
sampling one or more user behavior sequences from the plurality of user behavior sequences with a probability proportional to the respective DQ scores;
generating, by the sequential recommendation model, a second prediction from the sampled one or more user behavior sequences;
training a sequence encoder of the sequential recommendation model in multiple stages, including:
training the sequence encoder of the sequential recommendation model based on a training objective to minimize a loss computed via a noise contrastive estimation (NCE) module, wherein the loss is based on a comparison of the second prediction against a ground truth, corresponding to the sampled one or more user behavior sequences, and a noise distribution, resulting in a second set of parameters of the sequential recommendation model;
updating the respective combined DQ score for each of the plurality of user behavior sequences based on the sequential recommendation model with the second set of parameters; and
training the sequence encoder of the sequential recommendation model with the second set of parameters utilizing the updated respective combined DQ score for each of the plurality of user behavior sequences, resulting in a third set of parameters of the sequential recommendation model; and
generating, by the trained sequential recommendation model with the third set of parameters integrated at a recommender system, a next recommended item predicted based on an input user behavior sequence received via a user interface; and
displaying, via the user interface, the next recommended item.
2 . The method of claim 1 , wherein the difficulty score of the at least one user behavior sequence is based on an accuracy of next-item predictions of the sequential recommendation model for the at least one user behavior sequence.
3 . The method of claim 1 , wherein the difficulty score of the at least one user behavior sequence is combined with a previous difficulty score of the at least one user behavior sequence.
4 . The method of claim 1 , wherein the quality score of the at least one user behavior sequence is based on a measure of variance of prediction scores of the sequential recommendation model across items in the at least one user behavior sequence.
5 . The method of claim 1 , wherein the quality score of the at least one user behavior sequence is combined with a previous quality score of the at least one user behavior sequence.
6 . The method of claim 1 , wherein the determining the respective combined DQ score of the at least one user behavior sequence comprises summing a square of the difficulty score with a square of the quality score.
7 . The method of claim 6 , wherein the determining the respective combined DQ score of the at least one user behavior sequence further comprises raising the sum of the squares to a predetermined power.
8 . The method of claim 1 , wherein the determining the respective combined DQ score the at least one user behavior sequence comprises weighting the difficulty score and the quality score differently.
9 . The method of claim 1 , wherein the difficulty score and the quality score are iteratively updated based on the trained sequential recommendation model.
10 . A system for automatically generating a sequential recommendation of multiple items in a sequence to a user, comprising:
a memory that stores a sequential recommendation model;
a communication interface that receives a plurality of user behavior sequences; and
one or more hardware processors that:
receives, via the communication interface, a training dataset comprising a plurality of user behavior sequences;
augments the training dataset with a plurality of synthetic user behavior sequences generated by a trained extrapolator based on a subset of user behavior sequences of the plurality of user behavior sequences;
determines, via a DQ score generator for at least one user behavior sequence from the training dataset:
a difficulty score with an inverse relation to a probability of a sequential recommendation model with a first set of parameters correctly recommending an item in the at least one user behavior sequence, and
a quality score based on predictions of the sequential recommendation model with the first set of parameters;
determines, via the DQ score generator a respective combined difficulty and quality (DQ) score for each of the plurality of user behavior sequences based on the respective difficulty and quality scores;
samples one or more user behavior sequences from the plurality of user behavior sequences with a probability proportional to the respective DQ scores;
trains a sequence encoder of the sequential recommendation model in multiple stage, including:
training the sequence encoder of the sequential recommendation model based on a training objective to minimize a loss computed via a noise contrastive estimation (NCE) module, wherein the loss is based on a comparison of predictions generated by the sequential recommendation model against a ground-truth from the sampled one or more user behavior sequences, and a noise distribution, resulting in a second set of parameters of the sequential recommendation model,
updating the respective combined DQ score for each of the plurality of user behavior sequences based on the sequential recommendation model with the second set of parameters, and
training the sequence encoder of the sequential recommendation model with the second set of parameters utilizing the updated respective combined DQ score for each of the plurality of user behavior sequences, resulting in a third set of parameters of the sequential recommendation model;
generates by the trained sequential recommendation model with the third set of parameters integrated with a recommender system, a next recommended item predicted based on an input user behavior sequence received via a user interface; and
displaying, via the user interface, the next recommended item.
11 . The system of claim 10 , wherein the difficulty score of the at least one user behavior sequence is based on an accuracy of next-item predictions of the sequential recommendation model for the at least one user behavior sequence.
12 . The system of claim 10 , wherein the difficulty score of the at least one user behavior sequence is combined with a previous difficulty score of the at least one user behavior sequence.
13 . The system of claim 10 , wherein the quality score of the at least one user behavior sequence is based on a measure of variance of prediction scores of the sequential recommendation model across items in the at least one user behavior sequence.
14 . The system of claim 10 , wherein the quality score of the at least one user behavior sequence is combined with a previous quality score of the at least one user behavior sequence.
15 . The system of claim 10 , wherein the determining the respective combined DQ score of the at least one user behavior sequence comprises summing a square of the difficulty score with a square of the quality score.
16 . The system of claim 15 , wherein the determining the respective combined DQ score of the at least one user behavior sequence further comprises raising the sum of the squares to a predetermined power.
17 . The system of claim 10 , wherein the determining the respective combined DQ score the at least one user behavior sequence comprises weighting the difficulty score and the quality score differently.
18 . A non-transitory machine-readable medium comprising a plurality of machine-executable instructions which, when executed by one or more processors, are adapted to cause the one or more processors to perform operations comprising:
receiving, via a communication interface, a training dataset comprising a plurality of user behavior sequences;
determining, for at least one user behavior sequence from the training dataset:
a difficulty score with an inverse relation to a probability of a sequential recommendation model with a first set of parameters correctly recommending an item in the at least one user behavior sequence, and
a quality score based on a first prediction generated by the sequential recommendation model with the first set of parameters;
determining a first set of difficulty and quality (DQ) scores based on a combination of the difficulty score and the quality score, corresponding to the plurality of user behavior sequences, respectively;
training a sequence encoder of the sequential recommendation model in multiple stages, including:
updating a first set of parameters of the sequence encoder of the sequential recommendation model based on a training objective to minimize a loss computed via a noise contrastive estimation (NCE) module, wherein the loss is based on a comparison of predictions generated by the sequential recommendation model against a ground-truth from the plurality of user behavior sequences sampled based on respective DQ scores of the first set of DQ scores, and a noise distribution, resulting in a second set of parameters of the sequential recommendation model,
determining a second set of DQ scores based on the sequential recommendation model with the second set of parameters, and
updating the second set of parameters of the sequential recommendation utilizing the second set of DQ scores, resulting in a third set of parameters of the sequential recommendation model;
selecting a first subset of user behavior sequences and a second subset of user behavior sequences from the plurality of user behavior sequences based on the second set of DQ scores;
training an extrapolator using the first subset of user behavior sequences;
generating, by the trained extrapolator, a plurality of synthetic user behavior sequences based on the second subset of user behavior sequences; and
training the sequence encoder of the sequential recommendation model using user behavior sequences sampled from the training dataset with a probability proportional to respective DQ scores of the second set of DQ scores;
fine-tuning the sequence encoder of the sequential recommendation model using the plurality of synthetic user behavior sequences;
generating, by the fine-tuned recommendation model, a next recommended item predicted based on an input user behavior sequence received via a user interface; and
displaying, via the user interface, the next recommended item.
19 . The non-transitory machine-readable medium of claim 18 , wherein the first subset of user behavior sequences includes user behavior sequences with a respective quality score above a first predetermined threshold, and a respective difficulty score below a second predetermined threshold.
20 . The non-transitory machine-readable medium of claim 18 , wherein the second subset of user behavior sequences includes user behavior sequences with a respective quality score above a first predetermined threshold, and a respective difficulty score above a second predetermined threshold.