IP Library Granted Patent US 9,158,755
Granted Patent B2
US 9,158,755 · App. 13/663,563 · Granted Oct 13, 2015

Category-based lemmatizing of a phrase in a document

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,158,755
App. No.
13/663,563
Granted
Oct 13, 2015
Kind
B2
Abstract

A processor-implemented method, system, and/or computer program product lemmatizes a phrase for a specific category. An initial phrase, which is associated with a specific category, is received by a processor. The processor removes a last letter or set of letters from a word in the initial phrase to form an initial truncated version of the phrase, and then runs a term frequency-inverse document frequency (TF-IDF) algorithm on the initial truncated version of the phrase. The processor lemmatizes subsequent truncated versions of the initial phrase, and then runs the TF-IDF algorithm until a highest TF-IDF value is identified for a specific truncated version of the initial phrase when compared to TF-IDF values of other truncated versions of the initial phrase. The specific truncated version of the initial phrase that is associated with the highest TF-IDF value is then associated with the specific category.

Claims (40)

1. A processor-implemented method of lemmatizing a phrase for a specific category, the processor-implemented method comprising:

receiving, by a processor, a string of binary data that represents an initial phrase, wherein the initial phrase is an initial version of the phrase, wherein the phrase comprises multiple words, and wherein the phrase is associated with a specific category;

removing one or more letters from an end of a word in the initial phrase to form an initial truncated version of the phrase;

running, by the processor, a term frequency-inverse document frequency (TF-IDF) algorithm on the initial truncated version of the phrase;

the processor lemmatizing subsequent truncated versions of the initial phrase by recursively removing a remaining said one or more letters from the end of the word in a subsequent truncated version of the initial truncated version of the initial phrase;

the processor running the TF-IDF algorithm on subsequent truncated versions of the initial truncated version of the initial phrase until a highest TF-IDF value is identified for a specific truncated version of the initial phrase when compared to TF-IDF values of other truncated versions of the initial phrase;

assigning the specific truncated version of the initial phrase that is associated with the highest TF-IDF value to the specific category;

in response to receiving a request for the phrase within the specific category, returning the specific truncated version of the initial phrase that is associated with the highest TF-IDF value for said specific category; and

using the specific truncated version of the initial phrase that is associated with the highest TF-IDF value for said specific category to search a database that is dedicated to the specific category.

2. The processor-implemented method of claim 1 , wherein the specific category is a type of industry.

3. The processor-implemented method of claim 1 , wherein the specific truncated version of the initial phrase is a lemma for a lexeme, and wherein the specific category defines a breadth of the lemma.

4. The processor-implemented method of claim 1 , wherein the word in the initial truncated version whose last letter is removed is a last word in the initial phrase.

5. The processor-implemented method of claim 1 , wherein the word in the initial truncated version whose last letter is removed is not a last word in the initial phrase.

6. A computer program product for lemmatizing a phrase for a specific category, the computer program product comprising a tangible computer readable storage medium having program code embodied therewith, the program code readable and executable by a processor to perform a method comprising:

receiving a string of binary data that represents an initial phrase, wherein the initial phrase is an initial version of the phrase, wherein the phrase comprises multiple words, and wherein the phrase is associated with a specific category;

removing a last letter from a word in the initial phrase to form an initial truncated version of the phrase;

running a term frequency-inverse document frequency (TF-IDF) algorithm on the initial truncated version of the phrase;

lemmatizing subsequent truncated versions of the initial phrase by recursively removing a remaining last letter from the word in a subsequent truncated version of the initial truncated version of the initial phrase;

running the TF-IDF algorithm on subsequent truncated versions of the initial truncated version of the initial phrase until a highest TF-IDF value is identified for a specific truncated version of the initial phrase when compared to TF-IDF values of other truncated versions of the initial phrase;

assigning the specific truncated version of the initial phrase that is associated with the highest TF-IDF value to the specific category;

in response to receiving a request for the phrase within the specific category, returning the specific truncated version of the initial phrase that is associated with the highest TF-IDF value for said specific category; and

using the specific truncated version of the initial phrase that is associated with the highest TF-IDF value for said specific category to search a database that is dedicated to the specific category.

7. The computer program product of claim 6 , wherein the specific category is a type of industry.

8. The computer program product of claim 6 , wherein the specific category is an academic field of study.

9. The computer program product of claim 6 , wherein the word in the initial truncated version whose last letter is removed is a last word in the initial phrase.

10. The computer program product of claim 6 , wherein the word in the initial truncated version whose last letter is removed is not a last word in the initial phrase.

11. A computer system comprising:

a processor, a computer readable memory, and a computer readable storage medium;

first program instructions to receive a string of binary data that represents an initial phrase, wherein the initial phrase is an initial version of the phrase, wherein the phrase comprises multiple words, and wherein the phrase is associated with a specific category;

second program instructions to remove a last letter from a word in the initial phrase to form an initial truncated version of the phrase;

third program instructions to run a term frequency-inverse document frequency (TF-IDF) algorithm on the initial truncated version of the phrase;

fourth program instructions to lemmatize subsequent truncated versions of the initial phrase by recursively removing a remaining last letter from the word in a subsequent truncated version of the initial truncated version of the initial phrase;

fifth program instructions to run the TF-IDF algorithm on subsequent truncated versions of the initial truncated version of the initial phrase until a highest TF-IDF value is identified for a specific truncated version of the initial phrase when compared to TF-IDF values of other truncated versions of the initial phrase;

sixth program instructions to assign the specific truncated version of the initial phrase that is associated with the highest TF-IDF value to the specific category;

seventh program instructions to, in response to receiving a request for the phrase within the specific category, return the specific truncated version of the initial phrase that is associated with the highest TF-IDF value for said specific category; and

eighth program instructions to use the specific truncated version of the initial phrase that is associated with the highest TF-IDF value for said specific category to search a database that is dedicated to the specific category; and wherein

the first, second, third, fourth, fifth, sixth, seventh, and eighth program instructions are stored on the computer readable storage medium for execution by the processor via the computer readable memory.

12. The computer system of claim 11 , wherein the specific category is a type of industry.

13. The computer system of claim 11 , wherein the word in the initial truncated version whose last letter is removed is a last word in the initial phrase.

14. The computer system of claim 11 , wherein the word in the initial truncated version whose last letter is removed is not a last word in the initial phrase.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2021
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: AIRBNB, INC.
Reel/Frame 056427/0193 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 30, 2012
From: BOSTICK, JAMES E.; GANCI, JOHN M., JR.; KAEMMERER, JOHN P.; TRIM, CRAIG M.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 029209/0312 →