IP Library Granted Patent US 11,574,121
Granted Patent B2
US 11,574,121 · App. 17/156,864 · Granted Feb 7, 2023

Effective text parsing using machine learning

Inventors: Gautam K. Bhat (Mangalore, IN); Muniyandi Perumal Thevar (Madurai, IN); Nalini M (Chennai, IN); Sarita Lavania (Bengaluru, IN)
Assignee: Kyndryl, Inc.
G06F40/279G06F40/205G06F40/247G06F40/30G06N3/0445G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,574,121
App. No.
17/156,864
Granted
Feb 7, 2023
Kind
B2
Abstract

Techniques for data evaluation using machine learning are provided. A textual document is received, and the textual document is parsed using a recurrent neural network (RNN) to extract a plurality of keywords. A first subset of keywords which are known by a user and a second subset of keywords which are unknown by the user are each identified. A summary of the textual document is generated based on the second subset of keywords. The summary is output, comprising: outputting information related to a first keyword of the second subset of keywords and, upon determining that the first keyword is understood by the user, outputting information related to a second keyword of the second subset of keywords.

Claims (79)

1. A method, comprising:

receiving a textual document;

parsing the textual document using a recurrent neural network (RNN) to extract a plurality of keywords;

identifying a first subset of keywords, from the plurality of keywords, which are known by a first user;

identifying a second subset of keywords, from the plurality of keywords, which are unknown by the first user;

generating a summary of the textual document based on the second subset of keywords; and

outputting the summary, comprising:

outputting information related to a first keyword of the second subset of keywords; and

upon determining that the first keyword is understood by the first user, outputting information related to a second keyword of the second subset of keywords.

2. The method of claim 1 , wherein parsing the textual document to extract the plurality of keywords is based at least in part on (i) a set of keywords associated with an author of the textual document, and (ii) a genre of the textual document.

3. The method of claim 1 , wherein parsing the textual document to extract the plurality of keywords comprises:

identifying a third keyword for a first sentence of the textual document by processing the first sentence using the RNN;

identifying a fourth keyword for a second sentence of the textual document by processing the second sentence using the RNN; and

generating the first keyword for a first paragraph of the textual document by processing the third and fourth keywords using the RNN, wherein the first paragraph includes the first and second sentences.

4. The method of claim 1 , wherein parsing the textual document to extract the plurality of keywords comprises:

upon determining that the first keyword is not known to the first user, identifying one or more synonyms of the first keyword; and

upon determining that at least one of the one or more synonyms is known to the first user, mapping the first keyword and the at least one of the one or more synonyms.

5. The method of claim 1 , wherein parsing the textual document to extract the plurality of keywords comprises:

upon determining that the first keyword is not known to the first user, determining whether the first keyword has a regional meaning; and

upon determining that first keyword has a regional meaning, mapping the first keyword to the regional meaning.

6. The method of claim 1 , wherein parsing the textual document to extract the plurality of keywords comprises:

upon determining that the first keyword is not known to the first user, determining whether the first keyword has meaning in one or more social media platforms; and

upon determining that first keyword has a meaning in one or more social media platforms, mapping the first keyword to the meaning from the social media platforms.

7. The method of claim 1 , wherein parsing the textual document to extract the plurality of keywords comprises:

upon determining that the first keyword is not known to the first user, requesting a meaning for the first keyword; and

mapping the first keyword to the requested meaning.

8. One or more computer-readable storage media collectively containing computer program code that, when executed by operation of one or more computer processors, performs an operation comprising:

receiving a textual document;

parsing the textual document using a recurrent neural network (RNN) to extract a plurality of keywords;

identifying a first subset of keywords, from the plurality of keywords, which are known by a first user;

identifying a second subset of keywords, from the plurality of keywords, which are unknown by the first user;

generating a summary of the textual document based on the second subset of keywords; and

outputting the summary, comprising:

outputting information related to a first keyword of the second subset of keywords; and

upon determining that the first keyword is understood by the first user, outputting information related to a second keyword of the second subset of keywords.

9. The computer-readable storage media of claim 8 , wherein parsing the textual document to extract the plurality of keywords is based at least in part on (i) a set of keywords associated with an author of the textual document, and (ii) a genre of the textual document.

10. The computer-readable storage media of claim 8 , wherein parsing the textual document to extract the plurality of keywords comprises:

identifying a third keyword for a first sentence of the textual document by processing the first sentence using the RNN;

identifying a fourth keyword for a second sentence of the textual document by processing the second sentence using the RNN; and

generating the first keyword for a first paragraph of the textual document by processing the third and fourth keywords using the RNN, wherein the first paragraph includes the first and second sentences.

11. The computer-readable storage media of claim 8 , wherein parsing the textual document to extract the plurality of keywords comprises:

upon determining that the first keyword is not known to the first user, identifying one or more synonyms of the first keyword; and

upon determining that at least one of the one or more synonyms is known to the first user, mapping the first keyword and the at least one of the one or more synonyms.

12. The computer-readable storage media of claim 8 , wherein parsing the textual document to extract the plurality of keywords comprises:

upon determining that the first keyword is not known to the first user, determining whether the first keyword has a regional meaning; and

upon determining that first keyword has a regional meaning, mapping the first keyword to the regional meaning.

13. The computer-readable storage media of claim 8 , wherein parsing the textual document to extract the plurality of keywords comprises:

upon determining that the first keyword is not known to the first user, determining whether the first keyword has meaning in one or more social media platforms; and

upon determining that first keyword has a meaning in one or more social media platforms, mapping the first keyword to the meaning from the social media platforms.

14. The computer-readable storage media of claim 8 , wherein parsing the textual document to extract the plurality of keywords comprises:

upon determining that the first keyword is not known to the first user, requesting a meaning for the first keyword; and

mapping the first keyword to the requested meaning.

15. A system comprising:

one or more computer processors; and

one or more memories collectively containing one or more programs which when executed by the one or more computer processors performs an operation, the operation comprising:

receiving a textual document;

parsing the textual document using a recurrent neural network (RNN) to extract a plurality of keywords;

identifying a first subset of keywords, from the plurality of keywords, which are known by a first user;

identifying a second subset of keywords, from the plurality of keywords, which are unknown by the first user;

generating a summary of the textual document based on the second subset of keywords; and

outputting the summary, comprising:

outputting information related to a first keyword of the second subset of keywords; and

upon determining that the first keyword is understood by the first user, outputting information related to a second keyword of the second subset of keywords.

16. The system of claim 15 , wherein parsing the textual document to extract the plurality of keywords comprises:

identifying a third keyword for a first sentence of the textual document by processing the first sentence using the RNN;

identifying a fourth keyword for a second sentence of the textual document by processing the second sentence using the RNN; and

generating the first keyword for a first paragraph of the textual document by processing the third and fourth keywords using the RNN, wherein the first paragraph includes the first and second sentences.

17. The system of claim 15 , wherein parsing the textual document to extract the plurality of keywords comprises:

upon determining that the first keyword is not known to the first user, identifying one or more synonyms of the first keyword; and

upon determining that at least one of the one or more synonyms is known to the first user, mapping the first keyword and the at least one of the one or more synonyms.

18. The system of claim 15 , wherein parsing the textual document to extract the plurality of keywords comprises:

upon determining that the first keyword is not known to the first user, determining whether the first keyword has a regional meaning; and

upon determining that first keyword has a regional meaning, mapping the first keyword to the regional meaning.

19. The system of claim 15 , wherein parsing the textual document to extract the plurality of keywords comprises:

upon determining that the first keyword is not known to the first user, determining whether the first keyword has meaning in one or more social media platforms; and

upon determining that first keyword has a meaning in one or more social media platforms, mapping the first keyword to the meaning from the social media platforms.

20. The system of claim 15 , wherein parsing the textual document to extract the plurality of keywords comprises:

upon determining that the first keyword is not known to the first user, requesting a meaning for the first keyword; and

mapping the first keyword to the requested meaning.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2021
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: KYNDRYL, INC.
Reel/Frame 058213/0912 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2021
From: BHAT, GAUTAM K.; THEVAR, MUNIYANDI PERUMAL; M, NALINI; LAVANIA, SARITA
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 055030/0897 →
Continuity (1)
Related Publication 20220237375A1 · Jul 28, 2022