IP Library Granted Patent US 10,748,118
Granted Patent B2
US 10,748,118 · App. 15/091,077 · Granted Aug 18, 2020

Systems and methods to develop training set of data based on resume corpus

Inventor: Miaoqing Fang (Menlo Park, CA)
Assignee: Facebook, Inc.
G06Q10/1053G06N5/04G06N20/00G06Q10/105
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,748,118
App. No.
15/091,077
Granted
Aug 18, 2020
Kind
B2
Abstract

Systems, methods, and non-transitory computer readable media are configured to acquire a resume corpus. The resume corpus is processed to generate resume tokens. A machine learning model is trained based on the resume tokens. The machine learning model is applied to recommend a job classification based on evaluation data.

Claims (49)

1. A computer-implemented method comprising:

acquiring, by a computing system, a resume corpus;

processing, by the computing system, the resume corpus to generate resume tokens from the resume corpus, wherein the processing comprises:

determining a ratio based on co-occurrence of a first word and a second word of the resume corpus versus individual occurrence of the first word and the second word; and

determining, based on the ratio, the existence of a bigram including the first word and the second word to be used as training data;

training, by the computing system, a machine learning model to recommend a job classification based at least in part on the bigram; and

applying, by the computing system, the machine learning model to recommend a job classification based on evaluation data.

2. The computer-implemented method of claim 1 , wherein the resume corpus is based on textual data from a plurality of resumes.

3. The computer-implemented method of claim 1 , wherein the resume tokens include one or more unigrams and one or more bigrams.

4. The computer-implemented method of claim 1 , wherein the processing the resume corpus comprises:

removing stop words from the resume corpus.

5. The computer-implemented method of claim 4 , wherein the stop words include at least one of pronouns, prepositions, articles, and conjunctions.

6. The computer-implemented method of claim 1 , wherein the processing the resume corpus comprises:

modifying capitalized letters of words in the resume corpus to have lowercase letters.

7. The computer-implemented method of claim 1 , wherein the processing the resume corpus comprises:

generating a whitelist of bigrams constituting job titles parsed from the resume corpus,

wherein the training a machine learning model to recommend a job classification is further based at least in part on the whitelist of bigrams constituting job titles.

8. The computer-implemented method of claim 7 , wherein the processing the resume corpus further comprises:

including a bigram in the whitelist based on satisfaction of a threshold appearance value relating to the bigram.

9. The computer-implemented method of claim 1 , wherein the job classification includes a job title or a job pipeline.

10. The computer-implemented method of claim 1 , wherein the processing further comprises:

comparing the ratio to a threshold value; and

determining the existence of a bigram including the first word and the second word to be used as training data when the ratio satisfies the threshold value.

11. A system comprising:

at least one processor; and

a memory storing instructions that, when executed by the at least one processor, cause the system to perform:

acquiring a resume corpus;

processing the resume corpus to generate resume tokens from the resume corpus, wherein the processing comprises:

determining a ratio based on co-occurrence of a first word and a second word of the resume corpus versus individual occurrence of the first word and the second word; and

determining, based on the ratio, the existence of a bigram including the first word and the second word to be used as training data;

training a machine learning model to recommend a job classification based at least in part on the bigram; and

applying the machine learning model to recommend a job classification based on evaluation data.

12. The system of claim 11 , wherein the resume corpus is based on textual data from a plurality of resumes.

13. The system of claim 11 , wherein the resume tokens include one or more unigrams and one or more bigrams.

14. The system of claim 11 , wherein the processing the resume corpus comprises:

removing stop words from the resume corpus.

15. The system of claim 14 , wherein the stop words include at least one of pronouns, prepositions, articles, and conjunctions.

16. A non-transitory computer-readable storage medium including instructions that, when executed by at least one processor of a computing system, cause the computing system to perform a method comprising:

acquiring a resume corpus;

processing the resume corpus to generate resume tokens from the resume corpus, wherein the processing comprises:

determining a ratio based on co-occurrence of a first word and a second word of the resume corpus versus individual occurrence of the first word and the second word; and

determining, based on the ratio, the existence of a bigram including the first word and the second word to be used as training data;

training a machine learning model to recommend a job classification based at least in part on the bigram; and

applying the machine learning model to recommend a job classification based on evaluation data.

17. The non-transitory computer-readable storage medium of claim 16 , wherein the resume corpus is based on textual data from a plurality of resumes.

18. The non-transitory computer-readable storage medium of claim 16 , wherein the resume tokens include one or more unigrams and one or more bigrams.

19. The non-transitory computer-readable storage medium of claim 16 , wherein the processing the resume corpus comprises:

removing stop words from the resume corpus.

20. The non-transitory computer-readable storage medium of claim 19 , wherein the stop words include at least one of pronouns, prepositions, articles, and conjunctions.

Assignments (2)
CHANGE OF NAME Recorded Nov 24, 2021
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 058250/0283 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2016
From: FANG, MIAOQING
To: FACEBOOK, INC.
Reel/Frame 040143/0945 →
Continuity (1)
Related Publication 20170286914A1 · Oct 5, 2017
Cited By (3)
US 12,321,694 US 12,380,722 US 12,670,196