IP Library › Granted Patent US 11,157,980
Granted Patent B2
US 11,157,980 · App. 15/856,883 · Granted Oct 26, 2021

Building and matching electronic user profiles using machine learning

Inventors: Swaminathan Balasubramanian (Troy, MI); Avijit Chatterjee (White Plains, NY); Rajiv Joshi (Yorktown Heights, NY); John J. Thomas (Fishkill, NY)
Assignee: International Business Machines Corporation
G06Q30/0623G06F16/9535G06N20/00G06Q30/0282G06Q30/0627H04L67/306
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,157,980
App. No.
15/856,883
Granted
Oct 26, 2021
Kind
B2
Abstract

Method and apparatus for generating profiles using machine learning and influencing online interactions are provided. The methods include receiving, from a first user of a plurality of users, a first set of electronic documents, where each electronic document in the first set of electronic documents corresponds to a respective user in the plurality of users. The methods also include identifying a plurality of user profiles, where each of the plurality of user profiles was generated by processing a corpus of electronic documents associated with each respective user using a first trained machine learning model. The methods include determining a plurality of match coefficients, based on comparing a plurality of user profiles associated with each respective user in the plurality of users, filtering the first set of electronic documents based on the plurality of match coefficients, and providing the filtered first set of electronic documents to the first user.

Claims (69)

1. A system, comprising:

a processor; and

a computer memory storing a program, which, when executed on the processor, performs an operation comprising:

receiving a first set of electronic documents, wherein each electronic document in the first set of electronic documents corresponds to a respective user in a plurality of users;

generating a first user profile comprising a plurality of principal attributes for a first user using a trained machine learning model, comprising:

generating a feature vector for an electronic document authored by the first user;

generating a numerical score for a first attribute of the plurality of principal attributes by processing the feature vector using the trained machine learning model, wherein the trained machine learning model was trained based on a plurality of documents, each respective document associated with a corresponding label indicating a respective score of the respective document with respect to the first attribute; and

updating the first user profile based on the generated numerical score;

identifying a new attribute using one or more unsupervised machine learning models, wherein the new attribute is not included in the plurality of principal attributes;

adding the new attribute to the plurality of principal attributes in the first user profile;

updating the first user profile to include a numerical score for the new attribute;

identifying a plurality of user profiles, wherein each of the plurality of user profiles was generated by processing a corpus of electronic documents associated with each respective user using the trained machine learning model, wherein each user profile specifies a plurality of attribute values for the plurality of principal attributes;

determining a plurality of match coefficients, one for each of the plurality of users, based on comparing the first user profile associated with the first user and the plurality of user profiles associated with each respective user in the plurality of users, wherein the plurality of match coefficients comprise numerical values indicating how closely matched the first user is with each respective user;

filtering the first set of electronic documents by removing at least one electronic document from the first set based on a match coefficient associated with a second user of the plurality of users, wherein the at least one electronic document corresponds to the second user; and

providing the filtered first set of electronic documents to the first user.

2. The system of claim 1 , wherein the at least one electronic document is removed from the first set of electronic documents based on determining that the match coefficient associated with the second user does not exceed a predefined threshold.

3. The system of claim 1 , wherein the at least one electronic document is removed from the first set of electronic documents based on determining that the match coefficient associated with the second user exceeds a predefined threshold.

4. The system of claim 1 , wherein each electronic document in the first set of electronic documents comprises a rating of a product or service.

5. The system of claim 4 , wherein determining the plurality of match coefficients comprises:

identifying one or more of the plurality of principal attributes that are relevant to the product or service; and

comparing only the one or more identified principal attributes.

6. The system of claim 4 , the operation further comprising:

calculating an updated rating for the product or service based on the filtered first set of electronic documents.

7. The system of claim 1 , the operation further comprising:

providing at least an indication of the plurality of match coefficients to the first user.

8. The system of claim 1 , the operation further comprising:

sorting the filtered first set of electronic documents.

9. A method, performed by one or more processors, comprising:

receiving a plurality of electronic documents, wherein each of the plurality of electronic documents was created by a respective user in a plurality of users;

generating a first user profile comprising a plurality of principal attributes for a first user using a trained machine learning model, comprising:

generating a feature vector for an electronic document authored by the first user;

generating a numerical score for a first attribute of the plurality of principal attributes by processing the feature vector using the trained machine learning model, wherein the trained machine learning model was trained based on a plurality of documents, each respective document associated with a corresponding label indicating a respective score of the respective document with respect to the first attribute; and

updating the first user profile based on the generated numerical score;

identifying a new attribute using one or more unsupervised machine learning models, wherein the new attribute is not included in the plurality of principal attributes;

adding the new attribute to the plurality of principal attributes in the first user profile;

updating the first user profile to include a numerical score for the new attribute;

determining a plurality of match coefficients for the plurality of principal attributes, one for each of the plurality of users, by comparing the first user profile with a respective user profile of each respective user in the plurality of users, wherein the plurality of match coefficients comprise numerical values indicating how closely matched the first user is with each respective user; and

filtering the plurality of electronic documents based at least in part on the determined match coefficients.

10. The method of claim 9 , wherein the plurality of principal attributes for each user are generated by processing a corpus of electronic documents associated with each respective user using a first trained machine learning model.

11. The method of claim 9 , wherein each of the plurality of electronic documents are associated with a first concept, and wherein determining the plurality of match coefficients comprises:

identifying one or more of the plurality of principal attributes that are relevant to the first concept; and

comparing only the one or more identified principal attributes.

12. The method of claim 9 , wherein each of the plurality of electronic documents comprises a rating of a product or service, the method further comprising:

calculating an updated rating for the product or service based on the filtered plurality of electronic documents.

13. A computer-readable storage medium having computer-readable program code embodied therewith, the computer-readable program code executable by one or more computer processors to perform an operation comprising:

receiving a first set of electronic documents, wherein each electronic document in the first set of electronic documents corresponds to a respective user in a plurality of users;

generating a first user profile comprising a plurality of principal attributes for a first user using a trained machine learning model, comprising:

generating a feature vector for an electronic document authored by the first user;

generating a numerical score for a first attribute of the plurality of principal attributes by processing the feature vector using the trained machine learning model, wherein the trained machine learning model was trained based on a plurality of documents, each respective document associated with a corresponding label indicating a respective score of the respective document with respect to the first attribute; and

updating the first user profile based on the generated numerical score;

identifying a new attribute using one or more unsupervised machine learning models, wherein the new attribute is not included in the plurality of principal attributes;

adding the new attribute to the plurality of principal attributes in the first user profile;

updating the first user profile to include a numerical score for the new attribute;

identifying a plurality of user profiles, wherein each of the plurality of user profiles was generated by processing a corpus of electronic documents associated with each respective user using the trained machine learning model, wherein each user profile specifies a plurality of attribute values for the plurality of principal attributes;

determining a plurality of match coefficients, one for each of the plurality of users, based on comparing the first user profile associated with the first user and the plurality of user profiles associated with each respective user in the plurality of users, wherein the plurality of match coefficients comprise numerical values indicating how closely matched the first user is with each respective user;

filtering the first set of electronic documents by removing at least one electronic document from the first set based on a match coefficient associated with a second user of the plurality of users, wherein the at least one electronic document corresponds to the second user; and

providing the filtered first set of electronic documents to the first user.

14. The computer-readable storage medium of claim 13 , wherein the at least one electronic document is removed from the first set of electronic documents based on determining that the match coefficient associated with the second user does not exceed a predefined threshold.

15. The computer-readable storage medium of claim 13 , wherein the at least one electronic document is removed from the first set of electronic documents based on determining that the match coefficient associated with the second user exceeds a predefined threshold.

16. The computer-readable storage medium of claim 13 , wherein each electronic document in the first set of electronic documents comprises a rating of a product or service.

17. The computer-readable storage medium of claim 16 , wherein determining the plurality of match coefficients comprises:

identifying one or more of the plurality of principal attributes that are relevant to the product or service; and

comparing only the one or more identified principal attributes.

18. The computer-readable storage medium of claim 16 , the operation further comprising:

calculating an updated rating for the product or service based on the filtered first set of electronic documents.

19. The computer-readable storage medium of claim 13 , the operation further comprising:

providing at least an indication of the plurality of match coefficients to the first user.

20. The computer-readable storage medium of claim 13 , the operation further comprising:

sorting the filtered first set of electronic documents.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 28, 2017
From: BALASBRAMANIAN, SWAMINATHAN; CHATTERJEE, AVIJIT; JOSHI, RAJIV; THOMAS, JOHN J
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 044501/0460 →
Continuity (1)
Related Publication 20190205950A1 · Jul 4, 2019