IP Library › Granted Patent US 11,556,711
Granted Patent B2
US 11,556,711 · App. 16/552,471 · Granted Jan 17, 2023

Analyzing documents using machine learning

Inventors: Kishore Gopalan (Jersey City, NJ); Satish Chaduvu (Hyderabad, IN); Thomas J. Kuchcicki (Farmingville, NY)
Assignee: Bank of America Corporation
G06F40/30G06F16/353G06F40/103G06F40/117G06F40/10G06F40/106G06K9/6215G06K9/6267G06K9/6268G06N20/00G06V30/414
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,556,711
App. No.
16/552,471
Granted
Jan 17, 2023
Kind
B2
Abstract

A document analysis device that includes a memory operable to store a machine learning model configured to receive a sentence as an input and to output a classification identifier that is associated with a sentence type for the received sentence. The device further includes an artificial intelligence (AI) processing engine configured to receive a document comprising text, to sentences within the document, and to classify the sentences using the machine learning model. The AI processing engine is further configured to identify tagging rules for the document and to annotate one or more sentences from the document with a sentence type that matches a sentence type that is identified by the tagging rules for the document.

Claims (77)

1. A document analysis device, comprising:

a memory operable to store:

a machine learning model configured to:

receive a sentence as an input; and

output a classification identifier for the received sentence, wherein the classification identifier identifies a sentence type; and

an artificial intelligence (AI) processing engine implemented by a processor operably coupled to the memory, configured to:

receive a document comprising text;

identify a plurality of sentences within the document;

classify the plurality of sentences using the machine learning model, wherein classifying the plurality of sentences associates each sentence of the plurality of sentences with a respective classification identifier that identifies a respective sentence type;

identify a document type of the document;

identify tagging rules for the document based on the identified document type, wherein the tagging rules identify one or more sentence types;

determine one or more sentences from the plurality of sentences, wherein the respective sentence types of the determined sentences match the sentence types identified by the tagging rules; and

annotate the determined sentences within the document, wherein annotating the determined sentences changes a format of the determined sentences.

2. The device of claim 1 , wherein:

the AI processing engine is configured to:

receive a user input that modifies an annotated sentence;

add the modified annotated sentence to a set of sentences for training the machine learning model; and

retrain the machine learning model using the set of sentences for training the machine learning model.

3. The device of claim 1 , wherein the AI processing engine is configured to:

compute similarity scores between an annotated sentence and a plurality of previously classified sentences, wherein a similarity score is a numeric value that represents how similar a pair of sentences are to each other based on the text within the pair of sentences;

identify a sentence from the plurality of previously classified sentences that corresponds with a similarity score that exceeds a similarity score threshold value, wherein the similarity score threshold value indicates a minimum similarity score for a pair of sentences to be considered alternatives of each other; and

output the identified sentence as an alternative sentence for the annotated sentence.

4. The device of claim 3 , wherein the plurality of previously classified sentences are each associated with a sentence type that is different from a sentence type associated with the annotated sentence.

5. The device of claim 1 , wherein classifying the plurality of sentences is based at least in part on verb tenses used in the plurality of sentences.

6. The device of claim 1 , wherein annotating the one or more identified sentences within the document comprises identifying a sentence type for the one or more identified sentences.

7. The device of claim 1 , wherein the AI processing engine is further configured to:

display the text from the document in a first application window; and

display the one or more annotated sentences in a second application window, wherein the second application window is different from the first application window.

8. A document analysis method, comprising:

receiving a document comprising text;

identifying a plurality of sentences within the document;

classifying the plurality of sentences using a machine learning model, wherein:

the machine learning model configured to:

receive a sentence as an input; and

output a classification identifier for the received sentence, wherein the classification identifier identifies a sentence type; and

classifying the plurality of sentences associates each sentence of the plurality of sentences with a respective classification identifier that identifies a respective sentence type;

identifying a document type of the document;

identifying tagging rules for the document based on the identified document type, wherein the tagging rules identify one or more sentence types;

determining one or more sentences from the plurality of sentences, wherein the respective sentence types of the determined sentences match the sentence types identified by the tagging rules; and

annotating the determined sentences within the document, wherein annotating the determined sentences changes a format of the determined sentences.

9. The method of claim 8 , further comprising:

receiving a user input that modifies an annotated sentence;

adding the modified annotated sentence to a set of sentences for training the machine learning model; and

retraining the machine learning model using the set of sentences for training the machine learning model.

10. The method of claim 8 , further comprising:

computing similarity scores between an annotated sentence and a plurality of previously classified sentences;

identifying a sentence from the plurality of previously classified sentences that corresponds with a similarity score that exceeds a similarity score threshold value, wherein the similarity score threshold value indicates a minimum similarity score for a pair of sentences to be considered alternatives of each other; and

outputting the identified sentence as an alternative sentence for the annotated sentence.

11. The method of claim 10 , wherein the plurality of previously classified sentences are each associated with a sentence type that is different from a sentence type associated with the annotated sentence.

12. The method of claim 8 , wherein classifying the plurality of sentences is based at least in part on verb tenses used in the plurality of sentences.

13. The method of claim 8 , wherein annotating the one or more identified sentences within the document comprises identifying a sentence type for the one or more identified sentences.

14. The method of claim 8 , wherein further comprising:

displaying the text from the document in a first application window; and

displaying the one or more annotated sentences in a second application window, wherein the second application window is different from the first application window.

15. A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to:

receive a document comprising text;

identify a plurality of sentences within the document;

classify the plurality of sentences using a machine learning model, wherein:

the machine learning model configured to:

receive a sentence as an input; and

output a classification identifier for the received sentence, wherein the classification identifier identifies a sentence type; and

classifying the plurality of sentences associates each sentence of the plurality of sentences with a respective classification identifier that identifies a respective sentence type;

identify a document type of the document;

identify tagging rules for the document based on the identified document type, wherein the tagging rules identify one or more sentence types;

determine one or more sentences from the plurality of sentences, wherein the respective sentence types of the determined sentences match the one or more sentence types identified by the tagging rules; and

annotate the determined sentences within the document, wherein annotating the determined sentences changes a format of the determined sentences.

16. The non-transitory computer-readable medium of claim 15 , further comprising instructions that when executed by the processor causes the processor to:

receive a user input that modifies an annotated sentence;

add the modified annotated sentence to a set of sentences for training the machine learning model; and

retrain the machine learning model using the set of sentences for training the machine learning model.

17. The non-transitory computer-readable medium of claim 15 , further comprising instructions that when executed by the processor causes the processor to:

compute similarity scores between an annotated sentence and a plurality of previously classified sentences, wherein a similarity score is a numeric value that represents how similar a pair of sentences are to each other based on the text within the pair of sentences;

identify a sentence from the plurality of previously classified sentences that corresponds with a similarity score that exceeds a similarity score threshold value, wherein the similarity score threshold value indicates a minimum similarity score value for a pair of sentences to be considered alternatives of each other; and

output the identified sentence as an alternative sentence for the annotated sentence.

18. The non-transitory computer-readable medium of claim 17 , wherein the plurality of previously classified sentences are each associated with a sentence type that is different from a sentence type associated with the annotated sentence.

19. The non-transitory computer-readable medium of claim 15 , wherein classifying the plurality of sentences is based at least in part on verb tenses used in the plurality of sentences.

20. The non-transitory computer-readable medium of claim 15 , wherein annotating the one or more identified sentences within the document comprises identifying a sentence type for the one or more identified sentences.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 27, 2019
From: GOPALAN, KISHORE; CHADUVU, SATISH; KUCHCICKI, THOMAS J.
To: BANK OF AMERICA CORPORATION
Reel/Frame 050183/0573 →
Continuity (1)
Related Publication 20210065041A1 · Mar 4, 2021