IP Library Granted Patent US 11,507,743
Granted Patent B2
US 11,507,743 · App. 15/444,443 · Granted Nov 22, 2022

System and method for automatic key phrase extraction rule generation

Inventors: Inna Achlow (Kfar Sirkin, IL); Naomi Zeichner (Tel-Adashim, IL); Hila Kneller (Zufim, IL)
Assignee: NICE LTD.
G06F40/211G06F40/289
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,507,743
App. No.
15/444,443
Granted
Nov 22, 2022
Kind
B2
Abstract

A method, system, and non-transitory processor-readable storage medium for automatic key phrase rule generation for automatic key phrase extraction including: receiving a corpus sample including a plurality of documents containing text, receiving a plurality of identified key phrases which relate to a topic of the text of at least one corresponding document; assigning a part-of-speech to each word in the corpus sample; generating a part-of-speech pattern from each identified key phrase; and generating key phrase rules.

Claims (53)

1. A method for key phrase rule generation for key phrase extraction, comprising:

receiving, by a processor, a plurality of identified key phrases, each identified key phrase including at least a plurality of consecutive words from a text of at least one corresponding document of a corpus sample and relating to a topic of the text of the document corresponding to the key phrase, the corpus sample including a plurality of documents comprising text;

for each identified key phrase, generating, by the processor, a part-of-speech pattern from the identified key phrase, wherein the part-of-speech pattern is a sequence of parts-of-speech, where each part-of-speech in the sequence is assigned to a word in the identified key phrase;

determining, by the processor, how many times each generated part-of-speech pattern appears in at least a portion of the at least one corresponding document;

selecting, by the processor, a subset of generated part-of-speech patterns for extracting additional key phrases from an additional document based on how many times the part-of-speech pattern appears in at least the portion of the at least one document; and

generating, by the processor, a filter set of key phrases which do not indicate a topic of any document in the corpus sample by:

for each part-of-speech pattern from the subset of generated part-of-speech patterns, generating at least one generated key phrase from the corpus sample, wherein each generated key phrase comprises at least a plurality of consecutive words from the text of at least one document of the corpus sample, and each generated key phrase comprises a part-of-speech pattern from the subset of generated part-of-speech patterns; and

for each generated key phrase, adding, by the processor, the generated key phrase to the filter set if:

the generated key phrase appears in the corpus sample more than a second predetermined amount of times, and

the generated key phrase does not appear in the plurality of identified key phrases, and

removing, by the processor, the generated key phrase from the set of at least one generated key phrase if the generated key phrase was added to the filter set of key phrases.

2. The method according to claim 1 , wherein a part-of-speech pattern is selected for the subset of generated part-of-speech patterns if the part-of-speech pattern appears more than a predetermined amount of times in at least the portion of the at least one document.

3. The method according to claim 2 , wherein the predetermined amount of times is one time.

4. The method according to claim 1 , wherein the second predetermined amount of times is one time.

5. The method according to claim 1 , wherein the extracted additional set of at least one key phrase does not comprise a key phrase from the filter set of key phrases.

6. The method according to claim 1 , further comprising determining, by the processor, an accuracy score of the plurality of part-of-speech patterns.

7. The method according to claim 6 , further comprising, for each part-of-speech pattern in the subset of generated part-of-speech patterns:

determining, by the processor, an accuracy score of the plurality of part-of-speech patterns without the current part-of-speech pattern, and

if the accuracy score of the part-of-speech pattern is above the accuracy score of the plurality of part-of-speech patterns, removing, by the processor, the part-of-speech pattern from the subset of generated part-of-speech patterns.

8. The method according to claim 7 , wherein the determination of the accuracy score comprises determining a precision and a recall of the part-of-speech pattern.

9. The method according to claim 1 , wherein at least 1,500 identified key phrases are received.

10. A system for key phrase rule generation for key phrase extraction, comprising a memory; and

a processor, the processor configured to:

receive a plurality of identified key phrases, each identified key phrase including at least a plurality of consecutive words from a text of at least one corresponding document of a corpus sample and relating to a topic of the text of the document corresponding to the key phrase, the corpus sample including a plurality of documents containing text;

generate, for each identified key phrase, a part-of-speech pattern from the identified key phrase, wherein the part-of-speech pattern is a sequence of parts-of-speech, where each part-of-speech in the sequence is assigned to a word in the identified key phrase;

determine how many times each generated part-of-speech pattern appears in at least a portion of the at least one corresponding document;

select a subset of generated part-of-speech patterns for extracting additional key phrases from an additional document based on how many times the part-of-speech pattern appears in at least the portion of the at least one document;

for each part-of-speech pattern from the subset of generated part-of-speech patterns, generate at least one generated key phrase from the corpus sample, wherein each generated key phrase is at least a plurality of consecutive words from the text of at least one document of the corpus sample, and each generated key phrase comprises a part-of-speech pattern from the subset of generated part-of-speech patterns; and

for each generated key phrase, add the generated key phrase to a filter set of key phrases which do not indicate a topic of any document in the corpus sample if:

the generated key phrase is comprised in the corpus sample more than a second predetermined amount of times, and

the generated key phrase is not comprised in the plurality of identified key phrases, and

remove the generated key phrase from the set of at least one generated key phrase if the generated key phrase was added to the filter set of key phrases.

11. The system according to claim 10 , wherein the processor is configured to select a part-of-speech pattern for the subset of generated part-of-speech patterns if the part-of-speech pattern appears more than a predetermined amount of times in at least the portion of the at least one document.

12. The system according to claim 10 , wherein the processor is configured to, for each part-of-speech pattern in the subset of generated part-of-speech patterns:

determine an accuracy score of the plurality of part-of-speech patterns without the current part-of-speech pattern, and

if the accuracy score of the part-of-speech pattern is above the accuracy score of the plurality of part-of-speech patterns, remove the part-of-speech pattern from the subset of generated part-of-speech patterns.

13. The system according to claim 12 , wherein the processor determining the accuracy score comprises the processor determining a precision and a recall of the part-of-speech pattern.

14. A method, comprising:

receiving, by a processor, a plurality of documents, each document comprising text; receiving, by the processor, key phrases, each identified key phrase comprising at least a plurality of consecutive words from the corresponding document and each identified key phrase relating to a topic of the document corresponding to the keyphrase;

determining a part-of-speech for each word in the documents;

for each key phrase, creating, by the processor, a pattern from the key phrase, wherein the pattern is a series of parts-of-speech, where each part-of-speech in the series is assigned to a word in the key phrase; and

creating, by the processor, rules by:

determining, by the processor, how many appearances each generated pattern makes in at least a portion of the at least one document;

selecting, by the processor, a subset of created patterns for extracting key phrases from an additional document, wherein:

the selection is based on the appearances of the pattern in the document, and

extracting key phrases including a pattern from the subset of created patterns, wherein the extracted key phrase comprises at least a plurality of consecutive words from the additional document that relates to a topic of the additional document; and

generating, by the processor, a filter set of key phrases which do not indicate a topic of any document in the corpus sample by:

for each parts-of-speech pattern from the generated patterns, generating at least one generated key phrase from the documents, wherein each generated key phrase comprises at least a series of consecutive words from the text of at least one document from the plurality of documents, and each generated key phrase comprises a parts-of-speech pattern from the generated patterns; and

for each generated key phrase, adding, by the processor, the generated key phrase to the filter set if:

the generated key phrase appears in the plurality of documents more than a second predetermined amount of times, and

the generated key phrase does not appear in the plurality of identified key phrases, and

removing, by the processor, the generated key phrase from the set of at least one generated key phrase if the generated key phrase was added to the filter set of key phrases.

15. The method according to claim 14 , wherein generating key phrase rules comprises determining, by the processor, an accuracy score of the plurality of part-of-speech patterns.

Assignments (2)
SECURITY INTEREST Recorded Feb 26, 2026
From: NICE LTD; NICE SYSTEMS INC.; NICE SYSTEMS TECHNOLOGIES INC.; INCONTACT, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 074986/0208 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 29, 2017
From: ACHLOW, INNA; ZEICHNER, NAOMI; KNELLER, HILA
To: NICE LTD.
Reel/Frame 042521/0139 →
Continuity (1)
Related Publication 20180246872A1 · Aug 30, 2018