User portrait obtaining method, apparatus, and storage medium according to user behavior log records on features of articles
View Patent ↗User portrait obtaining method, apparatus, and storage medium are provided. The method includes: obtaining M training samples according to a user behavior log. An initialized user parameter matrix W m×k and an initialized tag parameter matrix H k×n are modified according to the M training samples by using a data fitting model, to obtain a final user parameter matrix W m×k and a final tag parameter matrix H k×n . Further, a user portrait matrix P m×n is obtained according to the final user parameter matrix W m×k and the final tag parameter matrix H k×n . In the present disclosure, a user and a tag are parameterized, and a user parameter matrix and a tag parameter matrix are modified by using the data fitting model, so as to fit a training sample.
1. A user portrait obtaining method, comprising:
recording behaviors of m users on articles, the behaviors performed by the m users on the articles including browsing, purchasing, adding to favorites, deleting, using, reposting, liking, disliking, or commenting, wherein the articles are each a physical object for work or life consumption;
generating a user behavior log of the m users on the articles according to the behaviors recorded of the m users on the articles;
obtaining M training samples according to the user behavior log, wherein the m users include a user u, the articles include an article i and an article j, the articles are represented by tags, and each of the tags includes a keyword describing an article feature, and wherein a training sample <u, i, j> of the M training samples reflects a preference degree of the user u for the article i different than a preference degree of the user u for the article j, and M being a positive integer, wherein obtaining the M training samples includes:
obtaining a user article matrix according to the user behavior log, wherein the user article matrix is represented by R m×h , with rows representing the m users and columns representing h articles, m being greater than or equal to 3, h being greater than or equal to 4, and an element R ui representing whether the user u has performed behavior on the article i; and
obtaining the M training samples according to the user article matrix;
modifying an initialized user parameter matrix W m×k and an initialized tag parameter matrix H k×n according to the M training samples, to obtain a final user parameter matrix W m×k and a final tag parameter matrix H k×n , m representing number of the users, k representing number of factors, n representing number of the tags, m being a positive integer, k being a positive integer, and n being an integer greater than 1;
obtaining a user portrait matrix P m×n according to the final user parameter matrix W m×k and the final tag parameter matrix H k×n , an element P ut in a u th row and a t th column in the user portrait matrix P m×n representing a preference degree of the user u for a tag t, u being an integer greater than or equal to 1 and smaller than or equal to m, and t being an integer greater than or equal to 1 and smaller than or equal to n;
identifying m−1 association users each in a social relationship with the user u, the m−1 association users and the user u together forming the m users, and generating a user similarity matrix S m×m reflecting similarity between every two users of the m users, wherein the user similarity matrix S m×m includes S uv representing a similarity between the user u and a user v, the user v is one of the m−1 association users, and normalization is performed on the S uv such that
∑
m
v
=
1
S
u
v
=
1
,
and wherein the similarity includes a correlation on communication frequency between the user u and the user v;
calculating a preference degree of the user u for a target article according to the user portrait matrix P m×n and tags of the target article and the user similarity matrix S m×m ; and
sending personalized recommendation to the user u of the target article according to the preference degree.
2. The user portrait obtaining method according to claim 1 , wherein the modifying an initialized user parameter matrix W m×k and an initialized tag parameter matrix H k×n according to the M training samples, to obtain a final user parameter matrix W m×k and a final tag parameter matrix H k×n comprises:
letting a=0, and calculating a preference degree of each of the m users for each of the n tags according to a user parameter matrix W m×k obtained after an a th round of modification and a tag parameter matrix H k×n obtained after the a th round of modification, wherein a is an integer greater than or equal to 0, a user parameter matrix W m×k obtained after the 0 th round of modification is the initialized user parameter matrix W m×k , and a tag parameter matrix H k×n obtained after the 0 th round of modification is the initialized tag parameter matrix H k×n ;
calculating a preference degree of each of the m users for each of h articles according to the preference degree of each of the m users for each of the n tags and an article tag matrix A h×n , wherein h represents the number of articles, and h is an integer greater than 1;
obtaining, according to the preference degree of each of the m users for each of the h articles, a probability corresponding to each of the M training samples, wherein a probability corresponding to the training sample <u, i, j> is a probability that a preference degree of the user u for the article i is greater than a preference degree for the article j;
using the probabilities respectively corresponding to the M training samples as input parameters of a data fitting model, and calculating an output result of the data fitting model;
when the output result satisfies a preset condition, determining the user parameter matrix W m×k obtained after the a th round of modification and the tag parameter matrix H k×n obtained after the a th round of modification as the final user parameter matrix W m×k and the final tag parameter matrix H k×n , respectively; and
when the output result dissatisfies the preset condition, modifying the user parameter matrix W m×k obtained after the a th round of modification to obtain a user parameter matrix W m×k obtained after an (a+1) th round of modification, modifying the tag parameter matrix H k×n obtained after the a th round of modification to obtain a tag parameter matrix H k×n obtained after the (a+1) th round of modification, letting a=a+1, and re-performing the steps from the step of calculating the preference degree of each of the m users for each of the n tags according to the user parameter matrix W m×k obtained after the a th round of modification and the tag parameter matrix H k×n obtained after the a th round of modification.
3. The user portrait obtaining method according to claim 1 , further comprising:
obtaining social network information of each of the m users;
generating the user similarity matrix S m×m according to the social network information of the m users; and
establishing a data fitting model according to the user similarity matrix S m×m .
4. The user portrait obtaining method according to claim 1 , wherein elements in the initialized user parameter matrix W m×k are normally distributed random numbers, and elements in the initialized tag parameter matrix H k×n are normally distributed random numbers.
5. The user portrait obtaining method according to claim 1 , wherein h represents number of the articles, and M is m×h×(h−1)/2.
6. The user portrait obtaining method according to claim 1 , wherein the element R ui in the user article matrix R m×h is either 0 or 1, with 0 representing the user u has not performed any behavior on the article i and 1 representing the user u has performed some behavior on the article i.
7. The user portrait obtaining method according to claim 1 , wherein the user u is user 3 and the article i is article 4 , the user article matrix R m×h includes user 1 , user 2 , the user 3 , article 1 , article 2 , article 3 , and the article 4 , training samples of the user 1 include <1, 2, 1>, <1, 2, 3>, and <1, 2, 4>, the training sample <1, 2, 1> represents the user 1 has a greater preference for the article 2 than the article 1 , the training sample <1, 2, 3> represents the user 1 has a greater preference for the article 2 than the article 3 , and the training sample <1, 2, 4> represents the user 1 has a greater preference for the article 2 than the article 4 .
8. The user portrait obtaining method according to claim 1 , wherein the user similarity matrix S m×m is generated by:
listing the m users respectively in m columns of the user similarity matrix S m×m ;
listing the m users respectively in m rows of the user similarity matrix S m×m ; and
setting a diagonal element Suu of the user similarity matrix to 0, the diagonal element Suu representing a similarity of the user u and the user u himself or herself.
9. The user portrait obtaining method according to claim 1 , wherein the user similarity matrix S m×m is generated by:
setting each similarity value of the user similarity matrix S m×m to be smaller than 1.
10. A user portrait obtaining apparatus, comprising: a memory, configured to store computer-readable instructions; and one or more processors, coupled to the memory and when the computer-readable instructions being executed, configured to:
record behaviors of m users on articles, the behaviors performed by the m users on the articles including browsing, purchasing, adding to favorites, deleting, using, reposting, liking, disliking, or commenting, wherein the articles are each a physical object for work or life consumption;
generate a user behavior log of the m users on the articles according to the behaviors recorded of the m users on the articles;
obtain M training samples according to the user behavior log, wherein the m users include a user u, the articles include an article i and an article j, the articles are represented by tags, and each of the tags includes a keyword describing an article feature, and wherein a training sample <u, i, j> of the M training samples reflects a preference degree of the user u for the article i different than a preference degree of the user u for the article j, and M being a positive integer, wherein to obtain the M training samples includes:
obtaining a user article matrix according to the user behavior log, wherein the user article matrix is represented by R m×h , with rows representing the m users and columns representing h articles, m being greater than or equal to 3, h being greater than or equal to 4, and an element R ui representing whether the user u has performed behavior on the article i; and
obtaining the M training samples according to the user article matrix;
modify an initialized user parameter matrix W m×k and an initialized tag parameter matrix H k×n according to the M training samples, to obtain a final user parameter matrix W m×k and a final tag parameter matrix H k×n , m representing number of the users, k representing number of factors, n representing number of the tags, m being a positive integer, k being a positive integer, and n being an integer greater than 1;
obtain a user portrait matrix P m×n according to the final user parameter matrix W m×k and the final tag parameter matrix H k×n , an element P ut in a u th row and a t th column in the user portrait matrix P m×n representing a preference degree of the user u for a tag t, u being an integer greater than or equal to 1 and smaller than or equal to m, and t being an integer greater than or equal to 1 and smaller than or equal to n;
identify m−1 association users each in a social relationship with the user u, the m−1 association users and the user u together forming the m users, and generate a user similarity matrix S m×m reflecting similarity between every two users of the m users, wherein the user similarity matrix S m×m includes S uv representing a similarity between the user u and a user v, the user v is one of the m−1 association users, and normalization is performed on the S uv such that
∑
v
=
1
m
S
uv
=
1
,
and wherein the similarity includes a correlation on communication frequency between the user u and the user v;
calculate a preference degree of the user u for a target article according to the user portrait matrix P m×n and tags of the target article and the user similarity matrix S m×m ; and
send personalized recommendation to the user u of the target article according to the preference degree.
11. The user portrait obtaining apparatus according to claim 10 , wherein the one or more processors are further configured to:
let a=0, and calculate a preference degree of each of the m users for each of the n tags according to a user parameter matrix W m×k obtained after an a th round of modification and a tag parameter matrix H k×n obtained after the a th round of modification, wherein a is an integer greater than or equal to 0, a user parameter matrix W m×k obtained after the 0 th round of modification is the initialized user parameter matrix W m×k , and a tag parameter matrix H k×n obtained after the 0 th round of modification is the initialized tag parameter matrix H k×n ;
calculate a preference degree of each of the m users for each of h articles according to the preference degree of each of the m users for each of the n tags and an article tag matrix A h×n , wherein h represents the number of articles, and h is an integer greater than 1;
obtain, according to the preference degree of each of the m users for each of the h articles, a probability corresponding to each of the M training samples, wherein a probability corresponding to the training sample <u, i, j> is a probability that a preference degree of the user u for the article i is greater than a preference degree for the article j;
use the probabilities respectively corresponding to the M training samples as input parameters of a data fitting model, and calculate an output result of the data fitting model;
when the output result satisfies a preset condition, determine the user parameter matrix W m×k obtained after the a th round of modification and the tag parameter matrix H k×n obtained after the a th round of modification as the final user parameter matrix W m×k and the final tag parameter matrix H k×n , respectively; and
when the output result dissatisfies the preset condition, modify the user parameter matrix W m×k obtained after the a th round of modification to obtain a user parameter matrix W m×k obtained after an (a+1) th round of modification, modify the tag parameter matrix H k×n obtained after the a th round of modification to obtain a tag parameter matrix H k×n obtained after the (a+1) th round of modification, let a=a+1, and re-perform the steps from the step of calculating the preference degree of each of the m users for each of the n tags according to the user parameter matrix W m×k obtained after the a th round of modification and the tag parameter matrix H k×n obtained after the a th round of modification.
12. The user portrait obtaining apparatus according to claim 10 , wherein the one or more processors are further configured to:
obtain social network information of each of the m users;
generate the user similarity matrix S m×m according to the social network information of the m users; and
establish a data fitting model according to the user similarity matrix S m×m .
13. The user portrait obtaining apparatus according to claim 10 , wherein elements in the initialized user parameter matrix W m×k are normally distributed random numbers, and elements in the initialized tag parameter matrix H k×n are normally distributed random numbers.
14. The user portrait obtaining apparatus according to claim 10 , wherein the one or more processors are further configured to:
calculate a preference degree of the user u for a target article according to the user portrait matrix P m×n and tags of the target article.
15. The user portrait obtaining apparatus according to claim 10 , wherein h represents number of the articles, and M is m×h×(h−1)/2.
16. A non-transitory computer-readable storage medium containing computer-executable program instructions for one or more processors to:
record behaviors of m users on articles, the behaviors performed by the m users on the articles including browsing, purchasing, adding to favorites, deleting, using, reposting, liking, disliking, or commenting, wherein the articles are each a physical object for work or life consumption;
generate a user behavior log of the m users on the articles according to the behaviors recorded of the m users on the articles;
obtain M training samples according to the user behavior log, wherein the m users include a user u, and the articles include an article i and an article j, the articles are represented by tags, and each of the tags includes a keyword describing an article feature, and wherein a training sample <u, i, j> of the M training samples reflects a preference degree of the user u for the article i different than a preference degree of the user u for the article j, and M being a positive integer, wherein to obtain the M training samples includes:
obtaining a user article matrix according to the user behavior log, wherein the user article matrix is represented by R m×h , with rows representing the m users and columns representing h articles, m being greater than or equal to 3, h being greater than or equal to 4, and an element R ui representing whether the user u has performed behavior on the article i; and
obtaining the M training samples according to the user article matrix;
modify an initialized user parameter matrix W m×k and an initialized tag parameter matrix H k×n according to the M training samples, to obtain a final user parameter matrix W m×k and a final tag parameter matrix H k×n , m representing number of the users, k representing number of factors, n representing number of the tags, m being a positive integer, k being a positive integer, and n being an integer greater than 1; and
obtain a user portrait matrix P m×n according to the final user parameter matrix W m×k and the final tag parameter matrix H k×n , an element P ut in a u th row and a t th column in the user portrait matrix P m×n representing a preference degree of the user u for a tag t, u being an integer greater than or equal to 1 and smaller than or equal to m, and t being an integer greater than or equal to 1 and smaller than or equal to n;
identify m−1 association users each in a social relationship with the user u, the m−1 association users and the user u together forming the m users, and generate a user similarity matrix S m×m reflecting similarity between every two users of the m users, wherein the user similarity matrix S m×m includes S uv representing a similarity between the user u and a user v, the user v is one of the m−1 association users, and normalization is performed on the S uv such that
∑
v
=
1
m
S
uv
=
1
,
and wherein the similarity includes a correlation on communication frequency between the user u and the user v;
calculate a preference degree of the user u for a target article according to the user portrait matrix P m×n and tags of the target article and the user similarity matrix S m×m ; and
send personalized recommendation to the user u of the target article according to the preference degree.
17. The non-transitory computer-readable storage medium according to claim 16 , wherein the one or more processors are further configured to:
let a=0, and calculate a preference degree of each of the m users for each of the n tags according to a user parameter matrix W m×k obtained after an a th round of modification and a tag parameter matrix H k×n obtained after the a th round of modification, wherein a is an integer greater than or equal to 0, a user parameter matrix W m×k obtained after the 0 th round of modification is the initialized user parameter matrix W m×k , and a tag parameter matrix H k×n obtained after the 0 th round of modification is the initialized tag parameter matrix H k×n ;
calculate a preference degree of each of the m users for each of h articles according to the preference degree of each of the m users for each of the n tags and an article tag matrix A h×n , wherein h represents the number of articles, and h is an integer greater than 1;
obtain, according to the preference degree of each of the m users for each of the h articles, a probability corresponding to each of the M training samples, wherein a probability corresponding to the training sample <u, i, j> is a probability that a preference degree of the user u for the article i is greater than a preference degree for the article j;
use the probabilities respectively corresponding to the M training samples as input parameters of a data fitting model, and calculate an output result of the data fitting model;
when the output result satisfies a preset condition, determine the user parameter matrix W m×k obtained after the a th round of modification and the tag parameter matrix H k×n obtained after the a th round of modification as the final user parameter matrix W m×k and the final tag parameter matrix H k×n , respectively; and
when the output result dissatisfies the preset condition, modify the user parameter matrix W m×k obtained after the a th round of modification to obtain a user parameter matrix W m×k obtained after an (a+1) th round of modification, modify the tag parameter matrix H k×n obtained after the a th round of modification to obtain a tag parameter matrix H k×n obtained after the (a+1) th round of modification, let a=a+1, and re-perform the steps from the step of calculating the preference degree of each of the m users for each of the n tags according to the user parameter matrix W m×k obtained after the a th round of modification and the tag parameter matrix H k×n obtained after the a th round of modification.
18. The non-transitory computer-readable storage medium according to claim 16 , wherein h represents number of the articles, and M is m×h×(h−1)/2.