IP Library Granted Patent US 8,589,370
Granted Patent B2
US 8,589,370 · App. 13/382,094 · Granted Nov 19, 2013

Acronym extraction

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,589,370
App. No.
13/382,094
Granted
Nov 19, 2013
Kind
B2
Abstract

Disclosed is a system and computer-implemented method for extracting an acronym and one or more corresponding expansions of the acronym from a document represented in a markup language. The computer-implemented method comprises: identifying at least one acronym contained in the document; determining one or more expansions of the at least one identified acronym based on a portion of document located proximate the identified acronym; determining a ranking for each determined expansion based attributes of the document; and selecting one or more expansions for an identified acronym using the determined rankings.

Claims (48)

1. A computer-implemented method comprising:

examining a document represented in a markup language to identify acronyms contained therein, wherein the examining comprises:

performing a first stage comprising:

parsing the document using a markup language parser to identify a text node;

identifying an acronym in the text node; and

performing a second stage after the first stage, the second stage comprising:

transforming the document represented in the markup language into plain text by removing markup language tags from the document;

identifying a further acronym in the plain text, the further acronym being in addition to the acronym identified in the first stage;

retrieving portions of the document for the respective identified acronyms, wherein each of the portions is located within the document proximate the respective one of the identified acronyms;

determining one or more expansions of each of the identified acronyms based on the corresponding retrieved document portion;

determining a ranking for each expansion based on attributes of the document; and

selecting expansions for the identified acronyms using the determined rankings.

2. The method of claim 1 , wherein identifying the further acronym in the plain text comprises determining a grammatical classification of each term within the plain text to enable identification of an acronym.

3. The method of claim 2 , wherein determining the grammatical classification comprises examining a text term within the plain text and determining compliance of text term attributes with conditions in order to identify that the text term is an acronym, and wherein said conditions relate to at least one of:

a length of the text term being within a particular range;

an amount of capitalization within the text term;

the presence of a special character within the text term; and

the text term being a predetermined term excluded from consideration as an acronym.

4. The method of claim 1 , wherein the retrieving comprises selectively decomposing terms within at least one retrieved document portion into individual sub-terms to facilitate identification of an expansion within that retrieved document portion.

5. The method of claim 1 , wherein the determining one or more expansions comprises identifying terms within each retrieved document portion that include an initial portion containing a first character of the corresponding acronym and producing a corresponding set of terms for each identified term including said identified term and following terms within the retrieved document portion.

6. The method of claim 1 , wherein the determining the ranking utilizes an algorithm considering attributes of the markup language of the corresponding retrieved document portion and an indicator of the reliability of the document.

7. The method of claim 6 , wherein the algorithm determines compliance of attributes of the markup language with rules, and wherein said rules relate to at least one of:

the presence of predetermined markup tags;

the presence of one or more special characters;

the presence of predetermined text indicating an expansion of an acronym.

8. An apparatus comprising:

a non-transitory computer-readable medium storing a computer program; and

at least one processor, the computer program executable on the at least one processor to:

examine a document represented by a markup language to identify acronyms contained therein, the examining comprising:

performing a first stage comprising parsing the document using a markup language parser to identify a text node, and identifying an acronym in the text node; and

performing a second stage after the first stage, the second stage comprising transforming the document represented in the markup language into plain text by removing markup language tags from the document, and identifying a further acronym in the plain text, the further acronym being in addition to the acronym identified in the first stage;

retrieve portions of the document for the respective identified acronyms, wherein each of the portions is located within the document proximate the respective one of the identified acronyms;

determine one or more expansions of each of the identified acronyms based on the corresponding retrieved document portion;

determine a ranking for each expansion based on attributes of the document; and

select expansions for the identified acronyms using the determined rankings.

9. The apparatus of claim 8 , wherein the determining of the one or more expansions of each identified acronym comprises identifying terms within each retrieved document portion that include an initial portion containing a first character of the corresponding acronym and producing a corresponding set of terms for each identified term including said identified term and following terms within the retrieved document portion.

10. The apparatus of claim 8 , wherein the ranking utilizes an algorithm considering attributes of the markup language of the corresponding retrieved document portion and an indicator of the reliability of the document.

11. A non-transitory computer readable medium storing a computer program executable in a computer to:

examine a document represented in a markup language to identify acronyms contained in the document, the examining comprising:

performing a first stage comprising:

parsing the document using a markup language parser to identify a text node;

identifying an acronym in the text node; and

performing a second stage after the first stage, the second stage comprising:

transforming the document represented in the markup language into plain text by removing markup language tags from the document;

identifying a further acronym in the plain text, the further acronym being in addition to the acronym identified in the first stage;

retrieve portions of the document for the respective identified acronyms, wherein each of the portions is located within the document proximate the identified acronym;

determine one or more expansions of each of the identified acronyms based on the corresponding retrieved document portion; and

select expansions for the identified acronyms.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 9, 2021
From: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP; HEWLETT PACKARD ENTERPRISE COMPANY
To: THE CONUNDRUM IP LLC
Reel/Frame 057424/0170 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2015
From: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 037079/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 16, 2012
From: FENG, SHI-CONG; XIONG, YUHONG; LIU, WEI
To: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
Reel/Frame 027824/0980 →