IP Library › Granted Patent US 12,524,739
Granted Patent B2
US 12,524,739 · App. 17/729,939 · Granted Jan 13, 2026

Creating and using triplet representations to assess similarity between job description documents

Inventors: David G. George (Cary, NC); Sudhanshu S. Singh (New Delhi, IN); Joydeep Mondal (New Delhi, IN); Sarthak Ahuja (New Delhi, IN); John A. Medicke (Raleigh, NC); Amanda Klabzuba (Frisco, TX)
Assignee: International Business Machines Corporation
G06Q10/1053G06F7/026G06F40/205G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,739
App. No.
17/729,939
Granted
Jan 13, 2026
Kind
B2
Abstract

A method, system and computer program product for assessing similarity between two job description documents. Job description documents consist of sentences framed in a particular manner, where the sentences are represented as a set of actions, an object corresponding to each action and a set of attributes corresponding to the object. The two job description documents are parsed to generate a first and a second set of an action-object-attribute triplet representation, where the first set of the action-object-attribute triplet representation is associated with the first job description document and the second set of the action-object-attribute triplet representation is associated with the second job description document. A similarity score between the first and second sets of action-object-attribute triplet representations is then calculated by hierarchically matching the first and second sets of action-object-attribute triplet representations across the job description documents. In this manner, similar job positions/job descriptions may be more accurately identified.

Claims (49)

1 . A computer-implemented method for assessing similarity between two job description documents, the method comprising:

receiving, by a job description analyzer, a first and a second job description document, wherein each of said first and second job description documents comprises sentences represented as a set of actions, an object corresponding to each action and a set of attributes corresponding to said object;

parsing, by said job description analyzer, said first and second job description documents to generate a first and a second set of an action-object-attribute triplet representation, wherein a sentence tokenizer of said job description analyzer identifies a list of sentences in each of said first and second job description documents using natural language processing, wherein a word tokenizer of said job description analyzer identifies an action-object-attribute triplet representation for each sentence using dictionaries, language taxonomies and natural language processing, wherein said first set of action-object-attribute triplet representation is associated with said first job description document and said second set of action-object-attribute triplet representation is associated with said second job description document;

constructing, by said job description analyzer, an action similarity matrix containing semantic similarity scores between action terms represented in a triplet representation of a first document and action terms represented in a triplet representation of a second document, wherein said semantic similarity scores indicate how similar in meaning two terms are;

identifying, by said job description analyzer, assignments of actions with a highest semantic similarity score among job description documents based on said action similarity matrix;

calculating, by said job description analyzer, semantic similarity scores among objects and attributes of said action-object-attribute triplet representations corresponding to matched actions amongst said job description documents;

identifying, by said job description analyzer, assignments of objects with highest semantic similarity scores amongst said job description documents corresponding to matching pairs of actions using an object similarity matrix for action assignments which identifies a semantic similarity between object terms;

identifying, by said job description analyzer, assignments of attributes amongst said job description documents corresponding to matching pairs of objects using an attribute similarity matrix containing semantic similarity scores between attribute terms associated with object terms of said triplet representations of said job description documents;

hierarchically combining, by said job description analyzer, semantic similarity scores of said identified assignments of objects and said identified assignments of attributes to generate one matching score per pair of actions in said job description documents;

combining, by said job description analyzer, scores corresponding to all matching action pairs to create one document similarity score for said job description documents, wherein said document similarity score for said first and second job description documents is ((V A1 (1+N A1 (1+A A1 )/3)+(V A2 (1+(N A2 +N A3 +N A4 )/3))/2)/2), where V A1 is a semantic similarity score between action terms represented in triplet representations of said first document, where V A2 is a semantic similarity score between action terms represented in triplet representations of said second document, where N A1 is a highest similarity score in a first row of said object similarity matrix, where N A2 is a highest similarity score in a second row of said object similarity matrix, where N A3 is a highest similarity score in a third row of said object similarity matrix, where N A4 is a highest similarity score in a fourth row of said object similarity matrix, where A A1 is a semantic similarity score for said attribute terms associated with said object terms of said triplet representations of said first and second job description documents; and

determining, by said job description analyzer, a degree of similarity between said first and said second job description documents based on said document similarity score for said first and second description documents, wherein the higher a value of said document similarity score, the greater the degree of similarity between said first and second job description documents.

2 . The method as recited in claim 1 further comprising:

calculating semantic similarity scores among actions of said first and second sets of action-object-attribute triplet representations amongst said first and second job description documents.

3 . The method as recited in claim 1 further comprising:

identifying assignments of actions with highest semantic similarity scores amongst said first and second job description documents.

4 . The method as recited in claim 1 , wherein said sentence tokenizer identifies a beginning and an ending of sentences from said received first and second job descriptions using natural language processing or sentence boundary disambiguation, wherein after identifying said sentences from said received first and second job descriptions, said word tokenizer is used to identify a list of words in strings and to tag parts of speech, wherein said word tokenizer identifies actions, objects and attributes using said dictionaries and said language taxonomies.

5 . A computer program product for assessing similarity between two job description documents, the computer program product comprising one or more computer readable storage mediums having program code embodied therewith, the program code comprising programming instructions for:

receiving, by a job description analyzer, a first and a second job description document, wherein each of said first and second job description documents comprises sentences represented as a set of actions, an object corresponding to each action and a set of attributes corresponding to said object;

parsing, by said job description analyzer, said first and second job description documents to generate a first and a second set of an action-object-attribute triplet representation, wherein a sentence tokenizer of said job description analyzer identifies a list of sentences in each of said first and second job description documents using natural language processing, wherein a word tokenizer of said job description analyzer identifies an action-object-attribute triplet representation for each sentence using dictionaries, language taxonomies and natural language processing, wherein said first set of action-object-attribute triplet representation is associated with said first job description document and said second set of action-object-attribute triplet representation is associated with said second job description document;

constructing, by said job description analyzer, an action similarity matrix containing semantic similarity scores between action terms represented in a triplet representation of a first document and action terms represented in a triplet representation of a second document, wherein said semantic similarity scores indicate how similar in meaning two terms are;

identifying, by said job description analyzer, assignments of actions with a highest semantic similarity score among job description documents based on said action similarity matrix;

calculating, by said job description analyzer, semantic similarity scores among objects and attributes of said action-object-attribute triplet representations corresponding to matched actions amongst said job description documents;

identifying, by said job description analyzer, assignments of objects with highest semantic similarity scores amongst said job description documents corresponding to matching pairs of actions using an object similarity matrix for action assignments which identifies a semantic similarity between object terms;

identifying, by said job description analyzer, assignments of attributes amongst said job description documents corresponding to matching pairs of objects using an attribute similarity matrix containing semantic similarity scores between attribute terms associated with object terms of said triplet representations of said job description documents;

hierarchically combining, by said job description analyzer, semantic similarity scores of said identified assignments of objects and said identified assignments of attributes to generate one matching score per pair of actions in said job description documents;

combining, by said job description analyzer, scores corresponding to all matching action pairs to create one document similarity score for said job description documents, wherein said document similarity score for said first and second job description documents is ((V A1 (1+N A1 (1+A A1 )/3)+ (V A2 (1+ (N A2 +N A3 +N A4 )/3))/2)/2), where V A1 is a semantic similarity score between action terms represented in triplet representations of said first document, where V A2 is a semantic similarity score between action terms represented in triplet representations of said second document, where N A1 is a highest similarity score in a first row of said object similarity matrix, where N A2 is a highest similarity score in a second row of said object similarity matrix, where N A3 is a highest similarity score in a third row of said object similarity matrix, where N A4 is a highest similarity score in a fourth row of said object similarity matrix, where A A1 is a semantic similarity score for said attribute terms associated with said object terms of said triplet representations of said first and second job description documents; and

determining, by said job description analyzer, a degree of similarity between said first and said second job description documents based on said document similarity score for said first and second description documents, wherein the higher a value of said document similarity score, the greater the degree of similarity between said first and second job description documents.

6 . The computer program product as recited in claim 5 , wherein the program code further comprises the programming instructions for:

calculating semantic similarity scores among actions of said first and second sets of action-object-attribute triplet representations amongst said first and second job description documents.

7 . The computer program product as recited in claim 5 , wherein the program code further comprises the programming instructions for:

identifying assignments of actions with highest semantic similarity scores amongst said first and second job description documents.

8 . The computer program product as recited in claim 5 , wherein said sentence tokenizer identifies a beginning and an ending of sentences from said received first and second job descriptions using natural language processing or sentence boundary disambiguation, wherein after identifying said sentences from said received first and second job descriptions, said word tokenizer is used to identify a list of words in strings and to tag parts of speech, wherein said word tokenizer identifies actions, objects and attributes using said dictionaries and said language taxonomies.

9 . A job description analyzer, comprising:

a memory for storing a computer program for assessing similarity between two job description documents; and

a processor connected to said memory, wherein said processor is configured to execute program instructions of the computer program comprising:

receiving a first and a second job description document, wherein each of said first and second job description documents comprises sentences represented as a set of actions, an object corresponding to each action and a set of attributes corresponding to said object;

parsing said first and second job description documents to generate a first and a second set of an action-object-attribute triplet representation, wherein a sentence tokenizer of said job description analyzer identifies a list of sentences in each of said first and second job description documents using natural language processing, wherein a word tokenizer of said job description analyzer identifies an action-object-attribute triplet representation for each sentence using dictionaries, language taxonomies and natural language processing, wherein said first set of action-object-attribute triplet representation is associated with said first job description document and said second set of action-object-attribute triplet representation is associated with said second job description document;

constructing an action similarity matrix containing semantic similarity scores between action terms represented in a triplet representation of a first document and action terms represented in a triplet representation of a second document, wherein said semantic similarity scores indicate how similar in meaning two terms are;

identifying assignments of actions with a highest semantic similarity score among job description documents based on said action similarity matrix;

calculating semantic similarity scores among objects and attributes of said action-object-attribute triplet representations corresponding to matched actions amongst said job description documents;

identifying assignments of objects with highest semantic similarity scores amongst said job description documents corresponding to matching pairs of actions using an object similarity matrix for action assignments which identifies a semantic similarity between object terms;

identifying assignments of attributes amongst said job description documents corresponding to matching pairs of objects using an attribute similarity matrix containing semantic similarity scores between attribute terms associated with object terms of said triplet representations of said job description documents;

hierarchically combining semantic similarity scores of said identified assignments of objects and said identified assignments of attributes to generate one matching score per pair of actions in said job description documents;

combining scores corresponding to all matching action pairs to create one document similarity score for said job description documents, wherein said document similarity score for said first and second job description documents is ((V A1 (1+N A1 (1+A A1 )/3)+ (V A2 (1+ (N A2 +N A3 +N A4 )/3))/2)/2), where V A1 is a semantic similarity score between action terms represented in triplet representations of said first document, where V A2 is a semantic similarity score between action terms represented in triplet representations of said second document, where N A1 is a highest similarity score in a first row of said object similarity matrix, where N A2 is a highest similarity score in a second row of said object similarity matrix, where N A3 is a highest similarity score in a third row of said object similarity matrix, where N A4 is a highest similarity score in a fourth row of said object similarity matrix, where A A1 is a semantic similarity score for said attribute terms associated with said object terms of said triplet representations of said first and second job description documents; and

determining a degree of similarity between said first and said second job description documents based on said document similarity score for said first and second description documents, wherein the higher a value of said document similarity score, the greater the degree of similarity between said first and second job description documents.

10 . The job description analyzer as recited in claim 9 , wherein the program instructions of the computer program further comprise:

calculating semantic similarity scores among actions of said first and second sets of action-object-attribute triplet representations amongst said first and second job description documents.

11 . The job description analyzer as recited in claim 9 , wherein the program instructions of the computer program further comprise: identifying assignments of actions with highest semantic similarity scores amongst said first and second job description documents.

12 . The job description analyzer as recited in claim 9 , wherein said sentence tokenizer identifies a beginning and an ending of sentences from said received first and second job descriptions using natural language processing or sentence boundary disambiguation, wherein after identifying said sentences from said received first and second job descriptions, said word tokenizer is used to identify a list of words in strings and to tag parts of speech, wherein said word tokenizer identifies actions, objects and attributes using said dictionaries and said language taxonomies.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2022
From: GEORGE, DAVID G.; SINGH, SUDHANSHU S.; MONDAL, JOYDEEP; AHUJA, SARTHAK; MEDICKE, JOHN A.; KLABZUBA, AMANDA
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 059740/0361 →
Continuity (2)
Continuation 15854837 · Dec 27, 2017
Related Publication 20220261766A1 · Aug 18, 2022
References Cited (33)
US 6917952B1 · Dailey et al. · 2005 [cited by applicant]
US 7013264B2 · Dolan et al. · 2006 [cited by applicant]
US 7533094B2 · Zhang · 2009 [cited by examiner]
US 7610281B2 · Gandhi et al. · 2009 [cited by applicant]
US 7644047B2 · Assadian et al. · 2010 [cited by applicant]
US 8099415B2 · Luo · 2012 [cited by examiner]
US 8594239B2 · Manasse et al. · 2013 [cited by applicant]
US 8612457B2 · Brdiczka · 2013 [cited by examiner]
US 8650199B1 · Tong · 2014 [cited by applicant]
US 8688690B2 · Brdiczka et al. · 2014 [cited by applicant]
US 9201927B1 · Zhang · 2015 [cited by applicant]
US 9235624B2 · Zhou · 2016 [cited by applicant]
US 9355151B1 · Cranfill · 2016 [cited by examiner]
US 9514221B2 · Tropin et al. · 2016 [cited by applicant]
US 9710518B2 · Cheng et al. · 2017 [cited by applicant]
US 9892111B2 · Danielyan et al. · 2018 [cited by applicant]
US 11410130B2 · George et al. · 2022 [cited by applicant]
US 20050060643A1 · Glass et al. · 2005 [cited by applicant]
US 20070185871A1 · Canright · 2007 [cited by examiner]
US 20070288308A1 · Chen et al. · 2007 [cited by applicant]
US 20080065630A1 · Luo et al. · 2008 [cited by applicant]
US 20130031088A1 · Srikrishna et al. · 2013 [cited by applicant]
US 20160232160A1 · Buhrmann · 2016 [cited by examiner]
US 20170300563A1 · Kao · 2017 [cited by examiner]
CN 101079026A · 2007 [cited by examiner]
CN 105786781A · 2016 [cited by examiner]
CN 106250502A · 2016 [cited by applicant]
Guo, “personalized resume job matching system”. (Year: 2016). [cited by examiner]
Cunningham, “A Taxonomy of similarity mechanism for case based reasoning” (Year: 2009). [cited by examiner]
List of IBM Patents or Patent Applications Treated as Related, May 11, 2022, pp. 1-2. [cited by applicant]
Ahuja et al., “Similarity Computation Exploiting the Semantic and Syntactic Inherent Structure Among Job Titles,” CSOC 2017, LNCS 10601, 2017, pp. 3-18. [cited by applicant]
Samuel W.K. Chan, “Integrating Linguistic Primitives in Learning Context-Dependent Representation,” IEEE Transactions on Knowledge and Data Engineering, vol. 13, No. 2, Mar./Apr. 2001, pp. 157-175. [cited by applicant]
Padraig Cunningham, “A Taxonomy of Similarity Mechanisms for Case-Based Reasoning,” IEEE Transactions on Knowledge and Data Engineering, vol. 21, No. 11, Nov. 2009, pp. 1532-1543. [cited by applicant]