IP Library › Granted Patent US 12,154,362
Granted Patent B2
US 12,154,362 · App. 17/725,313 · Granted Nov 26, 2024

Framework for document layout and information extraction

Inventors: Mark Salacinski (Warren, NJ); Christian Joseph Merrill (Bordentown, NJ); Octavian Florin Filoti (Portsmouth, NH); Irene Yue-Ling Pak (New Brunswick, NJ)
Assignee: Bristol-Myers Squibb Company
G06V30/414G06V10/82G06V30/1448G06V30/147G06V30/19093
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,154,362
App. No.
17/725,313
Granted
Nov 26, 2024
Kind
B2
Abstract

Provided herein are system, apparatus, device, method, and/or computer program product embodiments, and/or combinations and sub-combinations thereof, for extracting data from a file. Embodiments described herein provide a framework to merge outputs of various models comprising extracted information from a file with its location information and annotated regions of interest into an output file ingestible by a database or knowledge base.

Claims (92)

1. A method for extracting information, the method comprising:

receiving, by a processor, a file comprising information and a plurality of regions of interest (ROIs), wherein the file is in a first format;

converting, by the processor, the file into an image;

generating, by the processor, using a first model, a first output comprising a first set of information extracted from the image and a first set of coordinates of the first set of information in the image;

generating, by the processor, using a second model, a second output comprising a second set of coordinates for each ROI of the plurality of ROIs in the image;

generating, by the processor, using a third model, a third output comprising a second set of information extracted from the image and a third set of coordinates of the second set of information the image;

merging, by the processor, the first output and the third output to generate the information included in the file and a plurality of coordinates, wherein the plurality of coordinates comprise coordinates for the information in the image;

generating, by the processor, an output file of a second format comprising a plurality of sections using the second output, wherein each section of the plurality of sections corresponds to an ROI of the plurality of ROIs, and each section of the plurality of sections is included in the output file based on coordinates of an ROI in the second set of coordinates corresponding to the respective section; and

populating, by the processor, each section of the plurality of sections in the output file with a portion of the information determined to correspond with the respective section based on coordinates corresponding with the portion of the information and the coordinates of the respective section,

wherein the second format allows the information in the output file to be searchable while it is rendered on a graphical user interface (GUI) or stored in a data storage device.

2. The method of claim 1 , wherein the information comprises one or more words and generating the first output or third output comprises generating a bounding box around each of the one or more words in the image.

3. The method of claim 1 , wherein the third set of coordinates form a bounding box around each ROI of the plurality of ROIs.

4. The method of claim 1 , wherein the first output is generated using optical character recognition (OCR).

5. The method of claim 1 , wherein the third output is generated using a neural network.

6. The method of claim 1 , wherein one or more machine-learning models use the output file to generate a knowledge base.

7. The method of claim 1 , wherein the information is selectable in the output file.

8. The method of claim 1 , further comprising, causing display, by the processor, of the output file and the image.

9. The method of claim 1 , wherein the information comprises words and/or images.

10. The method of claim 9 , wherein merging the first output with the third output to generate the information, comprises retaining the images included in the first output.

11. The method of claim 9 , further comprising:

identifying, by the processor, one or more words in the first output that share same coordinates with one or more words in the third output; and

determining, by the processor, a similarity level between the one or more words in the first output and the one or more words in the third output.

12. The method of claim 11 , further comprising:

assigning, by the processor, a first priority value to the first output and a second priority value to the third output;

including, by the processor, the one or more words from the third output in the plurality of words based on the second priority value of the third output and based on the similarity level being more than a predetermined threshold; and

excluding, by the processor, the one or more words from the first output in the plurality of words based on the first priority value of the first output and based on the similarity level being more than the predetermined threshold.

13. The method of claim 11 , further comprising:

identifying, by the processor, a first data type of the one or more words in the first output and a second data type of the one or more words in the third output;

including, by the processor, the one or more words from the first output in the plurality of words based on the first data type of the one or more words in the first output; and

excluding, by the processor, the one or more words from the third output in the plurality of words based on the second data type of the one or more words in the third output.

14. A system for extracting information, the system comprising:

a memory comprising instructions; and

a processor coupled to the memory, wherein the processor is configured to execute the instructions, and the instructions, when executed, cause the processor to:

receive a file of a first format comprising information and a plurality of regions of interest (ROIs);

convert the file into an image;

generate, using a first model, a first output comprising a first set of information extracted from the image and a first set of coordinates of the first set of information in the image;

generate, using a second model, a second output comprising a second set of coordinates for each ROI of the plurality of ROIs in the image;

generate, using a third model, a third output comprising a second set of information extracted from the image and a third set of coordinates of the second set of information the image;

merge the first output and the third output to generate the information included in the file and a plurality of coordinates, wherein the plurality of coordinates comprise coordinates for the information in the image;

generate an output file of a second format comprising a plurality of sections using the second output, wherein each section of the plurality of sections corresponds to an ROI of the plurality of ROIs, and each section of the plurality of sections is included in the output file based on coordinates of an ROI in the second set of coordinates corresponding to the respective section; and

populate each section of the plurality of sections in the output file with a portion of the information determined to correspond with the respective section based on coordinates corresponding with the portion of the information and the coordinates of the respective section,

wherein the second format allows the information in the output file to be searchable while it is rendered on a graphical user interface (GUI) or stored in a data storage device.

15. The system of claim 14 , wherein the information comprises one or more words and generating the first output or third output comprises generating a bounding box around each of the one or more words in the image.

16. The system of claim 14 , wherein the third set of coordinates form a bounding box around each ROI of the plurality of ROIs.

17. The system of claim 14 , wherein the first output is generated using optical character recognition (OCR).

18. The system of claim 14 , wherein the third output is generated using a neural network.

19. The system of claim 14 , wherein the output file is used by one or more machine-learning models to generate a knowledge base.

20. The system of claim 14 , wherein the information is selectable in the output file.

21. The system of claim 14 , wherein the instructions, when executed, further cause the processor to cause display of the output file and the image.

22. The system of claim 14 , wherein the information comprises words and/or images.

23. The system of claim 22 , wherein merging the first output with the third output to generate the information, comprises retaining the images included in the first output.

24. The system of claim 22 , wherein the instructions, when executed, further cause the processor to:

identify one or more words in the first output that share same coordinates with one or more words in the third output; and

determine a similarity level between the one or more words in the first output and the one or more words in the third output.

25. The system of claim 24 , wherein the instructions, when executed, further cause the processor to:

assign a first priority value to the first output and a second priority value to the third output;

include the one or more words from the third output in the plurality of words based on the second priority value of the third output and based on the similarity level being more than a predetermined threshold; and

exclude the one or more words from the first output in the plurality of words based on the first priority value of the first output and based on the similarity level being more than the predetermined threshold.

26. The system of claim 24 , wherein the instructions, when executed, further cause the processor to:

identify a first data type of the one or more words in the first output and a second data type of the one or more words in the third output;

include the one or more words from the first output in the plurality of words based on the first data type of the one or more words in the first output; and

exclude the one or more words from the third output in the plurality of words based on the second data type of the one or more words in the third output.

27. A non-transitory machine-readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising, the operations comprising:

receiving a file of a first format comprising information and a plurality of regions of interest (ROIs);

converting the file into an image;

generating, using a first model, a first output comprising a first set of information extracted from the image and a first set of coordinates of the first set of information in the image;

generating, using a second model, a second output comprising a second set of coordinates for each ROI of the plurality of ROIs in the image;

generating, using a third model, a third output comprising a second set of information extracted in the image and a third set of coordinates of the second set of information the image;

merging the first output and the third output to generate the information included in the file and a plurality of coordinates, wherein the plurality of coordinates comprises coordinates for the information in the image;

generating an output file of a second format comprising a plurality of sections using the second output, wherein each section of the plurality of sections corresponds to an ROI of the plurality of ROIs, and each section of the plurality of sections is included in the output file based on coordinates of an ROI in the second set of coordinates corresponding to the respective section; and

populating each section of the plurality of sections in the output file with a portion of the information determined to correspond with the respective section based on coordinates corresponding with the portion of the information and the coordinates of the respective section,

wherein the second format allows the information in the output file to be searchable while it is rendered on a graphical user interface (GUI) or stored in a data storage device.

28. The non-transitory machine-readable medium of claim 27 , wherein the information comprises one or more words and generating the first output or third output comprises generating a bounding box around each of the one or more words in the image.

29. The non-transitory machine-readable medium of claim 27 , wherein the third set of coordinates form a bounding box around each ROI of the plurality of ROIs.

30. The non-transitory machine-readable medium of claim 27 , wherein the first output is generated using optical character recognition (OCR).

31. The non-transitory machine-readable medium of claim 27 , wherein the third output is generated using a neural network.

32. The non-transitory machine-readable medium of claim 27 , wherein the output file is used by one or more machine-learning models to generate a knowledge base.

33. The non-transitory machine-readable medium of claim 27 , wherein the information is selectable in the output file.

34. The non-transitory machine-readable medium of claim 27 , wherein the operations further comprise causing display of the output file and the image.

35. The non-transitory machine-readable medium of claim 27 , wherein the information comprises words and/or images.

36. The non-transitory machine-readable medium of claim 35 , wherein merging the first output with the third output to generate the information, comprises retaining the images included in the first output.

37. The non-transitory machine-readable medium of claim 35 , wherein the operations further comprise:

identifying one or more words in the first output that share same coordinates with one or more words in the third output; and

determining a similarity level between the one or more words in the first output and the one or more words in the third output.

38. The non-transitory machine-readable medium of claim 37 , wherein the operations further comprise:

assigning a first priority value to the first output and a second priority value to the third output;

including the one or more words from the third output in the plurality of words based on the second priority value of the third output and based on the similarity level being more than a predetermined threshold; and

excluding the one or more words from the first output in the plurality of words based on the first priority value of the first output and based on the similarity level being more than the predetermined threshold.

39. The non-transitory machine-readable medium of claim 37 , wherein the operations further comprise:

identifying a first data type of the one or more words in the first output and a second data type of the one or more words in the third output;

including the one or more words from the first output in the plurality of words based on the first data type of the one or more words in the first output; and

excluding the one or more words from the third output in the plurality of words based on the second data type of the one or more words in the third output.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 15, 2024
From: FILOTI, OCTAVIAN FLORIN; SALACINSKI, MARK; PAK, IRENE YUE-LING; MERRILL, CHRISTIAN JOSEPH
To: BRISTOL-MYERS SQUIBB COMPANY
Reel/Frame 067415/0521 →
Continuity (1)
Related Publication 20230343126A1 · Oct 26, 2023