IP Library Granted Patent US 10,331,764
Granted Patent B2
US 10,331,764 · App. 14/723,404 · Granted Jun 25, 2019

Methods and system for automatically obtaining information from a resume to update an online profile

Inventors: Ashwin Rao (Palo Alto, CA); Gaurav Bubna (Kolkata, IN); Zubin Mehta (Mumbai, IN)
Assignee: Hired, Inc.
G06F17/21G06F17/2229G06F17/2705G06F17/2745G06F17/2785G06Q10/063112G06Q10/10G06Q10/105G06Q10/1053
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,331,764
App. No.
14/723,404
Granted
Jun 25, 2019
Kind
B2
Abstract

Techniques involving accessing a resume of a person; automatically parsing the resume at least in part by: identifying, based at least in part on formatting of the resume, a plurality of sections in the resume including a first section; identifying, based at least in part on content in the first section and formatting of the content, a plurality of subsections of the first section; and processing text in the plurality of subsections to identify a plurality of credentials and associated attributes; and updating a profile for the person to reflect the plurality of credentials and the associated attributes.

Claims (97)

1. A method comprising:

using at least one computer hardware processor to perform:

accessing an electronic version of a resume of a person;

automatically parsing the resume at least in part by:

identifying, based at least in part on formatting of the resume, a plurality of sections in the resume including a first section, wherein identifying the plurality of sections comprises identifying a plurality of section headings at least in part by:

identifying a plurality of section heading candidates at least in part by:

 identifying a first phrase in the resume as a first section heading candidate based, at least in part, on content of the first phrase, and

 identifying a second phrase in the resume as a second section heading candidate when at least a threshold number of formatting characteristics of the second phrase match those of the first phrase; and

selecting, based on one or more attributes of the plurality of section heading candidates, the plurality of section headings from the plurality of section heading candidates;

identifying, based at least in part on content in the first section and formatting of the content, a plurality of subsections of the first section including a first subsection and a second subsection, the identifying comprising:

generating one or more first formatting features and one or more first content features for each line of text of multiple lines of text in the first section,

clustering the multiple lines of text in the first section based on the one or more first formatting features and the one or more first content features to obtain a first plurality of clusters including a first cluster, the first cluster comprising at least one line of text from the first subsection of the plurality of subsections of the first section and at least one line of text from the second subsection of the plurality of subsections of the first section,

identifying, from the first plurality of clusters, a cluster containing beginning lines of text of multiple subsections of the first section, and

identifying the plurality of subsections of the first section based, at least in part, on the identified cluster containing the beginning lines of text of the multiple subsections of the first section; and

processing text in the plurality of subsections to identify a plurality of credentials and associated attributes; and

populating an online profile for the person to reflect the plurality of credentials and the associated attributes, wherein the online profile for the person is one of a plurality of online profiles associated with a plurality of users of an online service.

2. The method of claim 1 , wherein selecting, based on the one or more attributes of the plurality of section heading candidates, the plurality of section headings from the plurality of section heading candidates comprises:

identifying, based on the one or more attributes, a first set of section heading candidates of the plurality of section heading candidates that are unlikely to be section headings;

identifying a second set of section heading candidates of the plurality of section heading candidates that are likely to be section headings; and

selecting the second set of section heading candidates as the plurality of section headings in the resume.

3. The method of claim 1 , wherein the plurality of sections in the resume include a second section, and the method further comprises:

identifying a plurality of subsections of the second section at least in part by:

generating one or more second formatting features and one or more second content features for each line of text of multiple lines of text in the second section, the one or more second formatting features and the one or more second content features being different from the one or more first formatting features and the one or more first content features associated with the first section;

clustering the multiple lines of text in the second section based on the one or more second formatting features and the one or more second content features to obtain a second plurality of clusters including a second cluster, the second cluster comprising at least one line of text from a first subsection of the plurality of subsections of the second section and at least one line of text from a second subsection of the plurality of subsections of the second section;

identifying, from the second plurality of clusters, a cluster containing beginning lines of text of multiple subsections of the second section; and

identifying the plurality of subsections of the second section based, at least in part, on the identified cluster containing the beginning lines of text of the multiple subsections of the second section.

4. The method of claim 1 , wherein the one or more first formatting features comprises at least one member selected from the group consisting of: a feature indicating whether the line of text starts with a bullet, a font color of a first token or one or more subsequent tokens in the line of text, a name of a font of the first token or the one or more subsequent tokens in the line of text, a font size of the first token or the one or more subsequent tokens in the line of text, a feature indicating whether the line of text is aligned in a particular way, a feature indicating whether the first token is upper case or lower case, a feature indicating a vertical distance of the line of text from a nearest line of text above the line of text, and a feature indicating an amount of empty space between the first token and a last token in the line of text.

5. The method of claim 1 , wherein the first subsection includes text describing a first credential, wherein the second subsection includes text describing a second credential, and wherein processing the text in the plurality of subsections comprises:

identifying first text in the first subsection as a first attribute of the first credential; and

identifying second text in the second subsection as a second attribute of the second credential based on formatting of the first text.

6. The method of claim 1 , wherein the first plurality of clusters comprises the first cluster and a second cluster, the cluster identified as containing the beginning lines of text is the first cluster, and identifying the cluster containing the beginning lines of text comprises identifying that the first cluster has a smaller number of lines of text than the second cluster.

7. At least one non-transitory computer-readable storage medium storing processor executable instructions that, when executed using at least one computer hardware processor, cause the at least one computer hardware processor to perform a method comprising:

accessing an electronic version of a resume of a person;

automatically parsing the resume at least in part by:

identifying, based at least in part on formatting of the resume, a plurality of sections in the resume including a first section, wherein identifying the plurality of sections comprises identifying a plurality of section headings at least in part by:

identifying a plurality of section heading candidates at least in part by:

identifying a first phrase in the resume as a first section heading candidate based, at least in part, on content of the first phrase, and

identifying a second phrase in the resume as a second section heading candidate when at least a threshold number of formatting characteristics of the second phrase match those of the first phrase; and

selecting, based on one or more attributes of the plurality of section heading candidates, the plurality of section headings from the plurality of section heading candidates;

identifying, based at least in part on content in the first section and formatting of the content, a plurality of subsections of the first section including a first subsection and a second subsection, the identifying comprising:

generating one or more first formatting features and one or more first content features for each line of text of multiple lines of text in the first section,

clustering the multiple lines of text in the first section based on the one or more first formatting features and the one or more first content features to obtain a first plurality of clusters including a first cluster, the first cluster comprising at least one line of text from the first subsection of the plurality of subsections of the first section and at least one line of text from the second subsection of the plurality of subsections of the first section,

identifying, from the first plurality of clusters, a cluster containing beginning lines of text of multiple subsections of the first section, and

identifying the plurality of subsections of the first section based, at least in part, on the identified cluster containing the beginning lines of text of the multiple subsections of the first section; and

processing text in the plurality of subsections to identify a plurality of credentials and associated attributes; and

populating an online profile for the person to reflect the plurality of credentials and the associated attributes, wherein the online profile for the person is one of a plurality of online profiles associated with a plurality of users of an online service.

8. The at least one non-transitory computer-readable storage medium of claim 7 , wherein selecting, based on the one or more attributes of the plurality of section heading candidates, the plurality of section headings from the plurality of section heading candidates comprises:

identifying, based on the one or more attributes, a first set of section heading candidates of the plurality of section heading candidates that are unlikely to be section headings;

identifying a second set of section heading candidates of the plurality of section heading candidates that are likely to be section headings; and

selecting the second set of section heading candidates as the plurality of section headings in the resume.

9. The at least one non-transitory computer-readable storage medium of claim 7 , wherein the plurality of sections in the resume include a second section, and the method further comprises:

identifying a plurality of subsections of the second section at least in part by:

generating one or more second formatting features and one or more second content features for each line of text of multiple lines of text in the second section, the one or more second formatting features and the one or more second content features being different from the one or more first formatting features and the one or more first content features associated with the first section;

clustering the multiple lines of text in the second section based on the one or more second formatting features and one or more second content features to obtain a second plurality of clusters including a second cluster, the second cluster comprising at least one line of text from a first subsection of the plurality of subsections of the second section and at least one line of text from a second subsection of the plurality of subsections of the second section;

identifying, from the second plurality of clusters, a cluster containing beginning lines of text of multiple subsections of the second section; and

identifying the plurality of subsections of the second section based, at least in part, on the identified cluster containing the beginning lines of text of the multiple subsections of the second section.

10. The at least one non-transitory computer-readable storage medium of claim 7 , wherein the one or more first formatting features comprises at least one member selected from the group consisting of: a feature indicating whether the line of text starts with a bullet, a font color of a first token or one or more subsequent tokens in the line of text, a name of a font of the first token or the one or more subsequent tokens in the line of text, a font size of the first token or the one or more subsequent tokens in the line of text, a feature indicating whether the line of text is aligned in a particular way, a feature indicating whether the first token is upper case or lower case, a feature indicating a vertical distance of the line of text from a nearest line of text above the line of text, and a feature indicating an amount of empty space between the first token and a last token in the line of text.

11. The at least one non-transitory computer-readable storage medium of claim 7 , wherein the first subsection includes text describing a first credential, wherein the second subsection includes text describing a second credential, and wherein processing the text in the plurality of subsections comprises:

identifying first text in the first subsection as a first attribute of the first credential; and

identifying second text in the second subsection as a second attribute of the second credential based on formatting of the first text.

12. The at least one non-transitory computer-readable storage medium of claim 7 , wherein the first plurality of clusters comprises the first cluster and a second cluster, the cluster identified as containing the beginning lines of text is the first cluster, and identifying the cluster containing the beginning lines of text comprises identifying that the first cluster has a smaller number of lines of text than the second cluster.

13. A system comprising:

at least one computer hardware processor configured to perform:

accessing an electronic version of a resume of a person;

automatically parsing the resume at least in part by:

identifying, based at least in part on formatting of the resume, a plurality of sections in the resume including a first section, wherein identifying the plurality of sections comprises identifying a plurality of section headings at least in part by:

identifying a plurality of section heading candidates at least in part by:

 identifying a first phrase in the resume as a first section heading candidate based, at least in part, on content of the first phrase, and

 identifying a second phrase in the resume as a second section heading candidate when at least a threshold number of formatting characteristics of the second phrase match those of the first phrase; and

selecting, based on one or more attributes of the plurality of section heading candidates, the plurality of section headings from the plurality of section heading candidates;

identifying, based at least in part on content in the first section and formatting of the content, a plurality of subsections of the first section including a first subsection and a second subsection, the identifying comprising:

generating one or more first formatting features and one or more first content features for each line of text of multiple lines of text in the first section,

clustering the multiple lines of text in the first section based on the one or more first formatting features and the one or more first content features to obtain a first plurality of clusters including a first cluster, the first cluster comprising at least one line of text from the first subsection of the plurality of subsections of the first section and at least one line of text from the second subsection of the plurality of subsections of the first section,

identifying, from the first plurality of clusters, a cluster containing beginning lines of text of multiple subsections of the first section, and

identifying the plurality of subsections of the first section based, at least in part, on the identified cluster containing the beginning lines of text of the multiple subsections of the first section; and

processing text in the plurality of subsections to identify a plurality of credentials and associated attributes; and

populating an online profile for the person to reflect the plurality of credentials and the associated attributes, wherein the online profile for the person is one of a plurality of online profiles associated with a plurality of users of an online service.

14. The system of claim 13 , wherein selecting, based on the one or more attributes of the plurality of section heading candidates, the plurality of section headings from the plurality of section heading candidates comprises:

identifying, based on the one or more attributes, a first set of section heading candidates of the plurality of section heading candidates that are unlikely to be section headings;

identifying a second set of section heading candidates of the plurality of section heading candidates that are likely to be section headings; and

selecting the second set of section heading candidates as the plurality of section headings in the resume.

15. The system of claim 13 , wherein the plurality of sections in the resume include a second section, and the method further comprises:

identifying a plurality of subsections of the second section at least in part by:

generating one or more second formatting features and one or more second content features for each line of text of multiple lines of text in the second section, the one or more second formatting features and the one or more second content features being different from the one or more first formatting features and the one or more first content features associated with the first section;

clustering the multiple lines of text in the second section based on the one or more second formatting features and the one or more second content features to obtain a second plurality of clusters including a second cluster, the second cluster comprising at least one line of text from a first subsection of the plurality of subsections of the second section and at least one line of text from a second subsection of the plurality of subsections of the second section;

identifying, from the second plurality of clusters, a cluster containing beginning lines of text of multiple subsections of the second section; and

identifying the plurality of subsections of the second section based, at least in part, on the identified cluster containing the beginning lines of text of the multiple subsections of the second section.

16. The system of claim 13 , wherein the one or more first formatting features comprises at least one member selected from the group consisting of: a feature indicating whether the line of text starts with a bullet, a font color of a first token or one or more subsequent tokens in the line of text, a name of a font of the first token or the one or more subsequent tokens in the line of text, a font size of the first token or the one or more subsequent tokens in the line of text, a feature indicating whether the line of text is aligned in a particular way, a feature indicating whether the first token is upper case or lower case, a feature indicating a vertical distance of the line of text from a nearest line of text above the line of text, and a feature indicating an amount of empty space between the first token and a last token in the line of text.

17. The system of claim 13 , wherein the first subsection includes text describing a first credential, wherein the second subsection includes text describing a second credential, and wherein processing the text in the plurality of subsections comprises:

identifying first text in the first subsection as a first attribute of the first credential; and

identifying second text in the second subsection as a second attribute of the second credential based on formatting of the first text.

18. The method of claim 1 , wherein the one or more first content features comprises at least one member selected from the group consisting of: a feature indicating whether the line of text contains a job title, a feature indicating whether the line of text contains a date, a feature indicating whether a line of text contains a start date and an end date, a feature indicating whether a next line of text in the first section contains a date, a feature indicating whether the next line of text in the first section contains a start date and an end date, a feature indicating whether the line of text contains a name of a location, a feature indicating whether the line of text contains a name of a degree, a feature indicating whether the next line of text in the first section contains a name of a degree, a feature indicating whether the line of text contains a grade or grade point average, a feature indicating whether the next line of text in the first section contains a grade or grade point average, a feature indicating whether the line of text contains a field of a degree, a feature indicating whether the next line of text in the first section contains a field of a degree, a feature indicating whether the line of text contains a name of a school, and a feature indicating whether the next line of text in the first section contains a name of a school.

19. The at least one non-transitory computer-readable storage medium of claim 7 , wherein the one or more first content features comprises at least one member selected from the group consisting of: a feature indicating whether the line of text contains a job title, a feature indicating whether the line of text contains a date, a feature indicating whether a line of text contains a start date and an end date, a feature indicating whether a next line of text in the first section contains a date, a feature indicating whether the next line of text in the first section contains a start date and an end date, a feature indicating whether the line of text contains a name of a location, a feature indicating whether the line of text contains a name of a degree, a feature indicating whether the next line of text in the first section contains a name of a degree, a feature indicating whether the line of text contains a grade or grade point average, a feature indicating whether the next line of text in the first section contains a grade or grade point average, a feature indicating whether the line of text contains a field of a degree, a feature indicating whether the next line of text in the first section contains a field of a degree, a feature indicating whether the line of text contains a name of a school, and a feature indicating whether the next line of text in the first section contains a name of a school.

20. The system of claim 13 , wherein the one or more first content features comprises at least one member selected from the group consisting of: a feature indicating whether the line of text contains a job title, a feature indicating whether the line of text contains a date, a feature indicating whether a line of text contains a start date and an end date, a feature indicating whether a next line of text in the first section contains a date, a feature indicating whether the next line of text in the first section contains a start date and an end date, a feature indicating whether the line of text contains a name of a location, a feature indicating whether the line of text contains a name of a degree, a feature indicating whether the next line of text in the first section contains a name of a degree, a feature indicating whether the line of text contains a grade or grade point average, a feature indicating whether the next line of text in the first section contains a grade or grade point average, a feature indicating whether the line of text contains a field of a degree, a feature indicating whether the next line of text in the first section contains a field of a degree, a feature indicating whether the line of text contains a name of a school, and a feature indicating whether the next line of text in the first section contains a name of a school.

21. The method of claim 1 , wherein the one or more attributes of the plurality of section heading candidates comprises at least one member selected from the group consisting of: an attribute indicating whether a section heading candidate includes at least a threshold number of words, an attribute indicating whether a section heading candidate contains any character that is not a letter, an attribute indicating whether a section heading candidate includes a date, and an attribute indicating whether a section heading candidate occurs in a middle of other text on a same line.

22. The at least one non-transitory computer-readable storage medium of claim 7 , wherein the one or more attributes of the plurality of section heading candidates comprises at least member selected from the group consisting of: an attribute indicating whether a section heading candidate includes at least a threshold number of words, an attribute indicating whether a section heading candidate contains any character that is not a letter, an attribute indicating whether a section heading candidate includes a date, and an attribute indicating whether a section heading candidate occurs in a middle of other text on a same line.

23. The system of claim 13 , wherein the one or more attributes of the plurality of section heading candidates comprises at least one member selected from the group consisting of: an attribute indicating whether a section heading candidate includes at least a threshold number of words, an attribute indicating whether a section heading candidate contains any character that is not a letter, an attribute indicating whether a section heading candidate includes a date, and an attribute indicating whether a section heading candidate occurs in a middle of other text on a same line.

Assignments (6)
CHANGE OF NAME Recorded Oct 27, 2025
From: ADO PROFESSIONAL SOLUTIONS, INC.
To: LHH RECRUITMENT SOLUTIONS, INC.
Reel/Frame 073266/0314 →
MERGER Recorded Jan 13, 2025
From: VETTERY, INC.
To: ADO PROFESSIONAL SOLUTIONS, INC.
Reel/Frame 069831/0082 →
CORRECTIVE ASSIGNMENT TO CORRECT THE THE APPLICATION NUMBER PREVIOUSLY RECORDED AT REEL: 058334 FRAME: 0046. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jan 7, 2022
From: H RESTRUCTURING LLC
To: VETTERY, INC.
Reel/Frame 058664/0266 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 8, 2021
From: HIRED, INC.
To: H RESTRUCTURING LLC
Reel/Frame 058334/0005 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2016
From: ZLEMMA, INC.
To: HIRED, INC.
Reel/Frame 037638/0116 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2015
From: RAO, ASHWIN; BUBNA, GAURAV; MEHTA, ZUBIN
To: ZLEMMA, INC.
Reel/Frame 035908/0969 →
Continuity (2)
Continuation In Part 14269691 · May 5, 2014
Related Publication 20150317610A1 · Nov 5, 2015
Cited By (2)
US 12,321,694 US 12,670,196