IP Library Granted Patent US 10,867,127
Granted Patent B2
US 10,867,127 · App. 16/408,046 · Granted Dec 15, 2020

Systems and methods for generating tables from print-ready digital source documents

Inventors: Mark Stephen Kyre (Winston-Salem, NC); Jeffrey Lucas Eldridge (Greensboro, NC); Austin Alexander Spears (Greensboro, NC); Samuel Allen Hudock (Greensboro, NC)
Assignee: DATAWATCH CORPORATION
G06F40/177G06F16/254G06F40/131G06K15/1814G06F16/313G06F16/951G06F40/103G06F40/106G06F40/14G06K9/00463
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,867,127
App. No.
16/408,046
Granted
Dec 15, 2020
Kind
B2
Abstract

Systems and methods are provided for generating tables from print-ready digital source documents. A document is received and one or more text fragments are identified on a rendered page of the document. A wrapping region collection is generated, comprising one or more wrapping regions. A tabular, narrative and label score is generated for each wrapping region. A block type is assigned to each wrapping region based on the scores. A wrapping region group and a block set are generated. One or more tables are generated based on text fragments corresponding to one of the one or more blocks. The text fragments are organized into corresponding fields of the one or more tables.

Claims (49)

1. A method comprising:

receiving, by a computing system, a print-ready digital source document, the digital source document comprising at least one rendered page;

generating, by the computing system, one or more wrapping regions based on the digital source document, wherein each wrapping region comprises one or more fragment runs, and wherein each fragment run comprises one or more text fragments of the digital source document that are adjacent to one another and within a horizontal separation threshold and a vertical separation threshold;

classifying, by the computing system, each of the one or more wrapping regions, wherein classifying each of the one or more wrapping regions comprises calculating, for each of the one or more wrapping regions, a tabular score, a narrative score, and a label score, wherein each of the tabular score, the narrative score, and the label score is a measure of qualification of a wrapping region as a tabular block type, as a narrative block type, and as a label block type, respectively;

generating, by the computing system, one or more blocks of wrapping regions based on the classifications of the one or more wrapping regions;

generating, by the computing system, one or more tables based on one or more blocks of wrapping regions; and

providing, by the computing system, an electronic document comprising the one or more tables.

2. The method of claim 1 , wherein the print-ready digital source document is at least one of an XPS document, an RTF document, or a PDF document.

3. The method of claim 1 , wherein the print-ready digital source document comprises a fixed layout file.

4. The method of claim 1 , wherein at least one of the tabular score, the narrative score or the label score of each of the one or more wrapping regions are calculated based on a normalization ratio, wherein the normalization ratio is a number of normalized text fragments divided by the total number of text fragments within a wrapping region.

5. The method of claim 1 , wherein at least one of the tabular score, the narrative score or the label score of each of the one or more wrapping regions are calculated based on a density ratio, wherein the density ratio is a percentage of a bounding box area of a wrapping region occupied by text fragment bounding boxes within the wrapping region.

6. The method of claim 1 , wherein at least one of the tabular score, the narrative score or the label score of each of the one or more wrapping regions are calculated based on an alignment ratio, wherein the alignment ratio is a number of normalized text fragments that fit within at least one alignment group divided by a total number of normalized text fragments within a wrapping region.

7. The method of claim 1 , wherein at least one of the tabular score, the narrative score or the label score of each of the one or more wrapping regions are calculated based on a capital or non-alphabetic ratio, wherein the capital or non-alphabetic ratio is a number of normalized text fragments that start with either a capital letter or a non-alphabetic character divided by a number of normalized text fragments within a wrapping region.

8. The method of claim 1 , wherein at least one of the tabular score, the narrative score or the label score of each of the one or more wrapping regions are calculated based on a text fragment quantity, wherein the text fragment quantity is a number of text fragments within a wrapping region.

9. The method of claim 1 , wherein at least one of the tabular score, the narrative score or the label score of each of the one or more wrapping regions are calculated based on a bold count, wherein the bold count is the number of text fragments with bolded text within a wrapping region.

10. A system comprising:

one or more processors;

one or more non-transitory computer-readable media including one or more sequences of instructions which, when executed by the one or more processors, causes:

receiving, by a computing system, a print-ready digital source document, the digital source document comprising at least one rendered page;

generating, by the computing system, one or more wrapping regions based on the digital source document, wherein each wrapping region comprises one or more fragment runs, and wherein each fragment run comprises one or more text fragments of the digital source document that are adjacent to one another and within a horizontal separation threshold and a vertical separation threshold;

classifying, by the computing system, each of the one or more wrapping regions, wherein classifying each of the one or more wrapping regions comprises calculating, for each of the one or more wrapping regions, a tabular score, a narrative score, and a label score, wherein each of the tabular score, the narrative score, and the label score is a measure of qualification of a wrapping region as a tabular block type, as a narrative block type, and as a label block type, respectively;

generating, by the computing system, one or more blocks of wrapping regions based on the classifications of the one or more wrapping regions;

generating, by the computing system, one or more tables based on one or more blocks of wrapping regions; and

providing, by the computing system, an electronic document comprising the one or more tables.

11. The system of claim 10 , wherein the print-ready digital source document is at least one of an XPS document, an RTF document, or a PDF document.

12. The system of claim 10 , wherein the print-ready digital source document comprises a fixed layout file.

13. The system of claim 10 , wherein at least one of the tabular score, the narrative score or the label score of each of the one or more wrapping regions are calculated based on a normalization ratio, wherein the normalization ratio is a number of normalized text fragments divided by the total number of text fragments within a wrapping region.

14. The system of claim 10 , wherein at least one of the tabular score, the narrative score or the label score of each of the one or more wrapping regions are calculated based on a density ratio, wherein the density ratio is a percentage of a bounding box area of a wrapping region occupied by text fragment bounding boxes within the wrapping region.

15. The system of claim 10 , wherein at least one of the tabular score, the narrative score or the label score of each of the one or more wrapping regions are calculated based on an alignment ratio, wherein the alignment ratio is a number of normalized text fragments that fit within at least one alignment group divided by a total number of normalized text fragments within a wrapping region.

16. The system of claim 10 , wherein at least one of the tabular score, the narrative score or the label score of each of the one or more wrapping regions are calculated based on a capital or non-alphabetic ratio, wherein the capital or non-alphabetic ratio is a number of normalized text fragments that start with either a capital letter or a non-alphabetic character divided by a number of normalized text fragments within a wrapping region.

17. The system of claim 10 , wherein at least one of the tabular score, the narrative score or the label score of each of the one or more wrapping regions are calculated based on a text fragment quantity, wherein the text fragment quantity is a number of text fragments within a wrapping region.

18. The system of claim 10 , wherein at least one of the tabular score, the narrative score or the label score of each of the one or more wrapping regions are calculated based on a bold count, wherein the bold count is the number of text fragments with bolded text within a wrapping region.

19. A non-transitory computer-readable medium including one or more sequences of instructions which, when executed by one or more processors, causes:

one or more processors;

one or more non-transitory computer-readable media including one or more sequences of instructions which, when executed by the one or more processors, causes:

receiving, by a computing system, a print-ready digital source document, the digital source document comprising at least one rendered page;

generating, by the computing system, one or more wrapping regions based on the digital source document, wherein each wrapping region comprises one or more fragment runs, and wherein each fragment run comprises one or more text fragments of the digital source document that are adjacent to one another and within a horizontal separation threshold and a vertical separation threshold;

classifying, by the computing system, each of the one or more wrapping regions, wherein classifying each of the one or more wrapping regions comprises calculating, for each of the one or more wrapping regions, a tabular score, a narrative score, and a label score, wherein each of the tabular score, the narrative score, and the label score is a measure of qualification of a wrapping region as a tabular block type, as a narrative block type, and as a label block type, respectively;

generating, by the computing system, one or more blocks of wrapping regions based on the classifications of the one or more wrapping regions;

generating, by the computing system, one or more tables based on one or more blocks of wrapping regions; and

providing, by the computing system, an electronic document comprising the one or more tables.

20. The non-transitory computer-readable medium of claim 19 , wherein the print-ready digital source document is at least one of an XPS document, an RTF document, or a PDF document.

21. The non-transitory computer-readable medium of claim 19 , wherein the print-ready digital source document comprises a fixed layout file.

22. The non-transitory computer-readable medium of claim 19 , wherein at least one of the tabular score, the narrative score or the label score of each of the one or more wrapping regions are calculated based on a normalization ratio, wherein the normalization ratio is a number of normalized text fragments divided by the total number of text fragments within a wrapping region.

23. The non-transitory computer-readable medium of claim 19 , wherein at least one of the tabular score, the narrative score or the label score of each of the one or more wrapping regions are calculated based on a density ratio, wherein the density ratio is a percentage of a bounding box area of a wrapping region occupied by text fragment bounding boxes within the wrapping region.

24. The non-transitory computer-readable medium of claim 19 , wherein at least one of the tabular score, the narrative score or the label score of each of the one or more wrapping regions are calculated based on an alignment ratio, wherein the alignment ratio is a number of normalized text fragments that fit within at least one alignment group divided by a total number of normalized text fragments within a wrapping region.

25. The non-transitory computer-readable medium of claim 19 , wherein at least one of the tabular score, the narrative score or the label score of each of the one or more wrapping regions are calculated based on a capital or non-alphabetic ratio, wherein the capital or non-alphabetic ratio is a number of normalized text fragments that start with either a capital letter or a non-alphabetic character divided by a number of normalized text fragments within a wrapping region.

26. The non-transitory computer-readable medium of claim 19 , wherein at least one of the tabular score, the narrative score or the label score of each of the one or more wrapping regions are calculated based on a text fragment quantity, wherein the text fragment quantity is a number of text fragments within a wrapping region.

27. The non-transitory computer-readable medium of claim 19 , wherein at least one of the tabular score, the narrative score or the label score of each of the one or more wrapping regions are calculated based on a bold count, wherein the bold count is the number of text fragments with bolded text within a wrapping region.

Assignments (3)
MERGER Recorded Feb 4, 2026
From: ALTAIR ENGINEERING INC.
To: SIEMENS INDUSTRY SOFTWARE INC.
Reel/Frame 074348/0312 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 21, 2024
From: DATAWATCH CORPORATION
To: ALTAIR ENGINEERING, INC.
Reel/Frame 066860/0606 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 9, 2019
From: KYRE, MARK STEPHEN; ELDRIDGE, JEFFREY LUCAS; SPEARS, AUSTIN ALEXANDER; HUDOCK, SAMUEL ALLEN
To: DATAWATCH CORPORATION
Reel/Frame 049133/0073 →
Continuity (3)
Continuation 15612979 · Jun 2, 2017
Continuation 14993988 · Jan 12, 2016
Related Publication 20190266233A1 · Aug 29, 2019
Cited By (2)
US 12,248,839 US 12,481,848