IP Library Granted Patent US 9,098,471
Granted Patent B2
US 9,098,471 · App. 13/543,445 · Granted Aug 4, 2015

Document content reconstruction

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,098,471
App. No.
13/543,445
Granted
Aug 4, 2015
Kind
B2
Abstract

A method, a storage medium and a system for document content reconstruction are provided in a digital content delivery and online education services platform to enable delivery of textbooks and other copyrighted material to multi-platform web browser applications. The method comprises ingesting a document page in an unstructured document format. The method further comprises extracting one or more images and metadata associated with the images and text and fonts associated with the texts from the document page. In addition, the method comprises coalescing text into paragraphs and creating a structured document page in a markup language format using the extracted images, text and fonts rendered with layout fidelity to the original ingested document page.

Claims (74)

1. A method comprising:

ingesting a document page in an unstructured document format, having a specified layout;

extracting one or more images and metadata associated with the images from the document page;

extracting text, and fonts associated with the texts, from the document page, wherein extracting the texts comprises:

determining Unicode mappings of the texts in the document page by:

mapping a first glyph of a letter to a standard code number in a Unicode chart table, and

mapping each of the one or more glyphs of the same letter to a unique code number in Private Use Area (PUA) of the Unicode, wherein each code number in PUA is obtained by masking the standard code number of the first glyph with a unique highest byte offset,

extracting all text characters and glyphs from the document page,

identifying horizontal and vertical positions of the extracted text characters and glyphs, and

extracting fonts associated with every character extracted;

coalescing the text into paragraphs; and

creating a structured document page in a markup language format using the extracted images, the text and the fonts rendered in accordance with the specified layout of the original ingested document page.

2. The method of claim 1 , wherein extracting one or more images comprises:

identifying graphical operations within the page;

determining one or more bounding boxes for the identified graphical operations and intersections of the bounding boxes;

combining intersecting bounding boxes; and

extracting images within each of the combined bounding boxes.

3. The method of claim 1 , wherein the metadata associated with the images includes at least one of: resolution, position, and caption of the images.

4. The method of claim 1 , further comprising restoring semantics of the letter by masking off the unique highest byte offset after displaying a corresponding glyph in an eReading browser application.

5. The method of claim 1 , wherein coalescing text into paragraphs comprises:

assembling the extracted text characters into individual words;

assembling words into lines, and lines into paragraphs; and

assembling paragraphs into respective bounding boxes or regions.

6. The method of claim 5 , wherein coalescing text into paragraphs is based on spacing and semantic analysis.

7. The method of claim 1 , further comprising:

repeating document reconstruction and recreation for each page of a plurality of pages in a document; and

constructing a table of contents for the document.

8. The method of claim 7 , wherein constructing a table of contents comprises:

searching for chapter headings within the document;

retrieving chapter level indices;

searching for sub-chapter headings within the document to retrieve sub-chapter level indices; and

updating dynamically the table of contents indices when new content is added.

9. A method for extracting texts within a document page comprising:

extracting text with custom fonts within the document page, the custom fonts having non-standard encodings, the extracting comprising determining Unicode mappings of the non-standard encodings by:

mapping a first glyph of a letter to a standard code number in a Unicode chart table; and

mapping one or more glyphs of the same letter to a unique code number in Private Use Area (PUA) of the Unicode, each code number in PUA obtained by masking the standard code number of the first glyph with a unique highest byte offset; and

restoring semantics of the letter by masking off the unique highest byte offset after displaying a corresponding glyph in an eReading browser application.

10. A non-transitory computer-readable storage medium storing executable computer program instructions for document content reconstruction, the computer program instructions comprising instructions for:

ingesting a document page in an unstructured document format, having a specified layout;

extracting one or more images and metadata associated with the images from the document page;

extracting text, and fonts associated with the texts, from the document page, wherein extracting the texts comprises:

determining Unicode mappings of the texts in the document page by:

mapping a first glyph of a letter to a standard code number in a Unicode chart table, and

mapping each of the one or more glyphs of the same letter to a unique code number in Private Use Area (PUA) of the Unicode, wherein each code number in PUA is obtained by masking the standard code number of the first glyph with a unique highest byte offset,

extracting all text characters and glyphs from the document page,

identifying horizontal and vertical positions of the extracted text characters and glyphs, and

extracting fonts associated with every character extracted;

coalescing the text into paragraphs; and

creating a structured document page in a markup language format using the extracted images, the text and the fonts rendered in accordance with the specified layout of the original ingested document page.

11. The storage medium of claim 10 , wherein extracting one or more images comprises:

identifying graphical operations within the page;

determining one or more bounding boxes for the identified graphical operations and intersections of the bounding boxes;

combining intersecting bounding boxes; and

extracting images within each of the combined bounding boxes.

12. The storage medium of claim 10 , wherein the metadata associated with the images includes at least one of: resolution, position, and caption of the images.

13. The storage medium of claim 10 , further comprising restoring semantics of the letter by masking off the unique highest byte offset after displaying a corresponding glyph in an eReading browser application.

14. The storage medium of claim 10 , wherein coalescing text into paragraphs comprises:

assembling the extracted text characters into individual words;

assembling words into lines, and lines into paragraphs; and

assembling paragraphs into respective bounding boxes or regions.

15. The storage medium of claim 14 , wherein coalescing text into paragraphs is based on spacing and semantic analysis.

16. The storage medium of claim 10 , further comprising:

repeating document reconstruction and recreation for each page of a plurality of pages in a document; and

constructing a table of contents for the document.

17. The storage medium of claim 16 , wherein constructing a table of contents comprises:

searching for chapter headings within the document;

retrieving chapter level indices;

searching for sub-chapter headings within the document to retrieve sub-chapter level indices; and

updating dynamically the table of contents indices when new content is added.

18. A non-transitory computer readable storage medium storing executable computer program instructions for extracting texts within a document page, the computer program instructions comprising instructions for:

extracting text with custom fonts within the document page, the custom fonts having non-standard encodings, the extracting comprising determining Unicode mappings of the non-standard encodings by:

mapping a first glyph of a letter to a standard code number in a Unicode chart table; and

mapping one or more glyphs of the same letter to a unique code number in Private Use Area (PUA) of the Unicode, each code number in PUA obtained by masking the standard code number of the first glyph with a unique highest byte offset; and

restoring semantics of the letter by masking off the unique highest byte offset after displaying a corresponding glyph in an eReading browser application.

Assignments (4)
SECURITY INTEREST Recorded Sep 22, 2016
From: CHEGG, INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION
Reel/Frame 039837/0859 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Sep 15, 2016
From: BANK OF AMERICA, N.A., AS LENDER
To: CHEGG, INC.
Reel/Frame 040043/0426 →
NOTICE OF GRANT OF SECURITY INTEREST IN PATENTS Recorded Aug 13, 2013
From: CHEGG, INC.
To: BANK OF AMERICA, N.A., AS LENDER
Reel/Frame 031006/0973 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 7, 2012
From: RICHARDSON, JOSHUA; LE CHEVALIER, VINCENT; JOSHI, ASHIT; ECKENBERG, DAX; DESAI, RAHUL RAVINDRA-MUTALIK; TWORETZKY, BRENT S.; GEIGER, CHARLES F.
To: CHEGG, INC.
Reel/Frame 029257/0363 →