IP Library › Granted Patent US 12,248,747
Granted Patent B2
US 12,248,747 · App. 18/537,378 · Granted Mar 11, 2025

Device dependent rendering of PDF content

Inventors: Erik Allan Juhl (København, DK); Anders Peter Fugmann (Værløse, DK)
Assignee: ISSUU, INC.
G06F40/109G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,248,747
App. No.
18/537,378
Filed
Dec 12, 2023
Granted
Mar 11, 2025
Kind
B2
Art Unit
2178
USPC
715/256
Abstract

The technology disclosed relates to systems and methods for device-dependent display of an article from a PDF file. The article can have multiple columns. The system can use a library to render the article from the PDF file. The rendering can include bounding boxes positioned at on-page coordinates that can include one or more images and multiple text blocks of glyphs. The system can partition the text blocks and images in two or more columns using dynamically adjusted valleys between columns. The system can set a reading order of the article after rendering. The system can merge and split text blocks to form paragraphs of text. The system includes logic to infer semantic information about typographic roles of the paragraphs from at least the font information. The system can cause display of the article in a device-dependent format using the semantic information and the reading order.

Claims (36)

1. A method of device-dependent display of an article from a PDF file that has multiple columns in at least parts of the article, the method including:

using a library to render the article from the PDF file, including rendering of a plurality of bounding boxes, positioned at on-page coordinates, that contain one or more images and multiple text blocks of glyphs, with font information for the glyphs;

setting a reading order of the article after the rendering, including pulling out text blocks spanning more than half of a width of a page and pulling out images, then reflowing the text blocks to produce the reading order;

merging the text blocks as they appear in the reading order into one or more paragraphs of text using the font information and using starting and ending positions of horizontally arranged text elements in the text blocks to delimit the paragraphs;

inferring semantic information about typographic roles of the paragraphs in the merged text blocks from at least the font information, including font name and font size distribution for sequences of the glyphs; and

causing display of the article in a device-dependent format, including the merged text blocks, using the semantic information and the reading order.

2. The method of claim 1 , wherein the font information of the glyphs further includes offsets and advances for characters indicating locations of characters in the bounding box.

3. The method of claim 1 , further including dividing some of the text blocks into separate paragraphs.

4. The method of claim 1 , further including using variation in vertical spacing between the horizontally arranged text elements during the merging of the text blocks into the paragraphs.

5. The method of claim 1 , wherein inferring the semantic information about the typographic roles of the paragraphs further includes distinguishing among at least headline text, subtitle text, body text, byline text, image attribution text, and image caption text and annotating the paragraphs with the inferred semantic information.

6. The method of claim 1 , further including displaying the images with the article.

7. The method of claim 1 , further including partitioning the text blocks and images into two or more columns and one or more sections of the article using dynamically adjusted valleys between the columns, wherein the dynamically adjusted valleys have sizes responsive to the font information for the glyphs in adjoining text blocks, and wherein dynamically adjusting the dynamically adjusted valleys further includes increasing a threshold for valley width, used when detecting a valley separating characters, responsive to large font sizes.

8. A system including a memory and one or more processors coupled to the memory, the memory loaded with computer instructions to generate device-dependent display of an article from a PDF file that has multiple columns in at least parts of the article, which computer instructions, when executed on the processors, implement actions comprising:

using a library to render the article from the PDF file, including rendering of a plurality of bounding boxes, positioned at on-page coordinates, that contain one or more images and multiple text blocks of glyphs, with font information for the glyphs;

setting a reading order of the article after the rendering, including pulling out text blocks spanning more than half of a width of a page and pulling out images, then reflowing the text blocks to produce the reading order;

merging the text blocks as they appear in the reading order into one or more paragraphs of text using the font information and using starting and ending positions of horizontally arranged text elements in the text blocks to delimit the paragraphs;

inferring semantic information about typographic roles of the paragraphs in the merged text blocks from at least the font information, including font name and font size distribution for sequences of the glyphs; and

causing display of the article in a device-dependent format, including the merged text blocks, using the semantic information and the reading order.

9. The system of claim 8 , wherein the font information of the glyphs further includes offsets and advances for characters indicating locations of characters in the bounding boxes.

10. The system of claim 8 , further implementing actions comprising, dividing some of the text blocks into separate paragraphs.

11. The system of claim 8 , further implementing actions comprising, using variation in vertical spacing between the horizontally arranged text elements during the merging of the text blocks into the one or more paragraphs.

12. The system of claim 8 , wherein inferring the semantic information about the typographic roles of the one or more paragraphs further includes distinguishing among at least headline text, subtitle text, body text, byline text, image attribution text, and image caption text and annotating the one or more paragraphs with the inferred semantic information.

13. The system of claim 8 , further implementing actions comprising, displaying the images with the article.

14. The system of claim 8 , further including using dynamically adjusted valleys between the columns, partitioning the text blocks and images into two or more columns and one or more sections of the article, wherein the dynamically adjusted valleys have sizes responsive to the font information for the glyphs in adjoining text blocks, and wherein dynamically adjusting the dynamically adjusted valleys further includes increasing a threshold for valley width, used when detecting a valley separating characters, responsive to large font sizes.

15. A non-transitory computer readable storage medium impressed with computer program instructions to generate device-dependent display of an article from a PDF file that has multiple columns in at least parts of the article, which computer program instructions, when executed on a processor, implement a method comprising:

using a library to render the article from the PDF file, including rendering of a plurality of bounding boxes, positioned at on-page coordinates, that contain one or more images and multiple text blocks of glyphs, with font information for the glyphs;

setting a reading order of the article after the rendering, including pulling out text blocks spanning more than half of a width of a page and pulling out images, then reflowing the text blocks to produce the reading order;

merging the text blocks as they appear in the reading order into one or more paragraphs of text using the font information and using starting and ending positions of horizontally arranged text elements in the text blocks to delimit the paragraphs;

inferring semantic information about typographic roles of the paragraphs in the merged text blocks from at least the font information, including font name and font size distribution for sequences of the glyphs; and

causing display of the article in a device-dependent format, including the merged text blocks, using the semantic information and the reading order.

16. The non-transitory computer readable storage medium of claim 15 , wherein the font information of the glyphs further includes offsets and advances for characters indicating locations of characters in the bounding box.

17. The non-transitory computer readable storage medium of claim 15 , implementing the method further comprising, dividing some of the text blocks into separate paragraphs.

18. The non-transitory computer readable storage medium of claim 15 , implementing the method further comprising, using variation in vertical spacing between the horizontally arranged text elements during the merging of the text blocks into the one or more paragraphs.

19. The non-transitory computer readable storage medium of claim 15 , wherein inferring the semantic information about the typographic roles of the paragraphs further includes distinguishing among at least headline text, subtitle text, body text, byline text, image attribution text, and image caption text and annotating the paragraphs with the inferred semantic information.

20. The non-transitory computer readable storage medium of claim 15 , implementing the method further comprising, displaying the images with the article.

21. The non-transitory computer readable storage medium of claim 15 , further including using dynamically adjusted valleys between the columns, partitioning the text blocks and images into two or more columns and one or more sections of the article, wherein the dynamically adjusted valleys have sizes responsive to the font information for the glyphs in adjoining text blocks, and wherein dynamically adjusting the dynamically adjusted valleys further includes increasing a threshold for valley width, used when detecting a valley separating characters, responsive to large font sizes.

Assignments (3)
MERGER Recorded Apr 24, 2026
From: ISSUU, INC.
To: BENDING SPOONS US INC.
Reel/Frame 074469/0246 →
PATENT SECURITY AGREEMENT Recorded Dec 30, 2024
From: ISSUU, INC.
To: INTESA SANPAOLO S.P.A., AS SECURITY AGENT
Reel/Frame 069813/0611 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 14, 2024
From: JUHL, ERIK ALLAN; FUGMANN, ANDERS PETER
To: ISSUU, INC.
Reel/Frame 066778/0041 →
Continuity (3)
Continuation 17888367 · Aug 15, 2022
Continuation 17099441 · Nov 16, 2020
Related Publication 20240119218A1 · Apr 11, 2024
References Cited (75)
US 5737599A · Rowe et al. · 1998 [cited by applicant]
US 5781785A · Rowe et al. · 1998 [cited by applicant]
US 5819301A · Rowe et al. · 1998 [cited by applicant]
US 7146566B1 · Hohensee · 2006 [cited by examiner]
US 7580038B2 · Chik et al. · 2009 [cited by applicant]
US 7853866B2 · Tanaka · 2010 [cited by applicant]
US 7908284B1 · Mathes et al. · 2011 [cited by applicant]
US 7979785B1 · Wang et al. · 2011 [cited by applicant]
US 8209600B1 · Koh · 2012 [cited by examiner]
US 8397155B1 · Szabo · 2013 [cited by examiner]
US 8539342B1 · Lewis · 2013 [cited by examiner]
US 9098471B2 · Richardson · 2015 [cited by examiner]
US 9224041B2 · Dejean et al. · 2015 [cited by applicant]
US 10108695B1 · Yeturu et al. · 2018 [cited by applicant]
US 10643022B2 · Karlapudi · 2020 [cited by examiner]
US 10839161B2 · Galitsky · 2020 [cited by applicant]
US 11295061B2 · Liu · 2022 [cited by examiner]
US 11615635B2 · Pellinen · 2023 [cited by examiner]
US 20010014900A1 · Brauer · 2001 [cited by examiner]
US 20060242166A1 · Larcheveque et al. · 2006 [cited by applicant]
US 20060248070A1 · Dejean et al. · 2006 [cited by applicant]
US 20070055518A1 · Doi · 2007 [cited by applicant]
US 20070196015A1 · Meunier et al. · 2007 [cited by applicant]
US 20080263023A1 · Vailaya et al. · 2008 [cited by applicant]
US 20080263032A1 · Vailaya et al. · 2008 [cited by applicant]
US 20080263033A1 · Vailaya et al. · 2008 [cited by applicant]
US 20090144277A1 · Trutner et al. · 2009 [cited by applicant]
US 20090144614A1 · Dresevic et al. · 2009 [cited by applicant]
US 20100174985A1 · Levy et al. · 2010 [cited by applicant]
US 20100251104A1 · Massand · 2010 [cited by examiner]
US 20120102388A1 · Fan · 2012 [cited by applicant]
US 20120159313A1 · Dejean · 2012 [cited by applicant]
US 20140099038A1 · Okada et al. · 2014 [cited by applicant]
US 20140208191A1 · Zaric et al. · 2014 [cited by applicant]
US 20140225928A1 · Konnola et al. · 2014 [cited by applicant]
US 20140258851A1 · Sesum et al. · 2014 [cited by applicant]
US 20140281939A1 · Agrawal · 2014 [cited by applicant]
US 20150154308A1 · Hagg et al. · 2015 [cited by applicant]
US 20160044196A1 · Baba · 2016 [cited by examiner]
US 20160140086A1 · O'Connor · 2016 [cited by applicant]
US 20170053424A1 · Wang et al. · 2017 [cited by applicant]
US 20170212870A1 · Thomsen · 2017 [cited by examiner]
US 20180366013A1 · Arvindam · 2018 [cited by applicant]
US 20190108204A1 · Ghosh et al. · 2019 [cited by applicant]
US 20190251163A1 · Bellert · 2019 [cited by applicant]
US 20200175101A1 · Holmberg-Nielsen et al. · 2020 [cited by applicant]
US 20200265225A1 · Hosabettu · 2020 [cited by examiner]
US 20200293607A1 · Nelson et al. · 2020 [cited by applicant]
US 20210019366A1 · Markey et al. · 2021 [cited by applicant]
US 20210065569A1 · Arvindam · 2021 [cited by applicant]
US 20220188503A1 · Arora et al. · 2022 [cited by applicant]
US 20220222420A1 · Jain et al. · 2022 [cited by applicant]
US 20220284175A1 · Schwiebert et al. · 2022 [cited by applicant]
CN 101876967B · 2012 [cited by examiner]
CN 112069771B · 2024 [cited by examiner]
WO WO2007117932A2 · 2007 [cited by examiner]
WO 2008130501A1 · 2008 [cited by applicant]
WO WO2021108038A1 · 2021 [cited by examiner]
Ishitani, Yasuto, “Document Transformatoin System from Papers to XML Data Based on Pivot XML Document Method”, Seventh International Conference on Document Analysis and Recoginition, 2003, 6 pages (ISSU 1006-1). [cited by applicant]
U.S. Appl. No. 17/099,449—Notice of Allowance dated Jan. 28, 2021, 9 pages. [cited by applicant]
Ramanathan, Chandrashekar et al., “Challenges in Generating Bookmarks from TOC Entries in e-Books”, In Proceedings of the 2012 ACM symposium on Document engineering, Association for Computing Machinery, New York, NY, US… [cited by applicant]
Marinai, Simone et al, “Table of Contents Recognition for Converting PDF documents in E-book Formats”, Proceedings of the 10th ACM Symposium on Document Engineering, Association for Computing Machinery, New York, NY, US… [cited by applicant]
U.S. Appl. No. 17/099,441—Notice of Allowance dated Dec. 8, 2021, 17 pages. [cited by applicant]
Bast, Hannah, et al, “A Benchmark and Evaluation for Text Extraction from PDF”, 2017 ACM/IEEE Joint Conference on Digital Libraries (JCDL), DOI: 10.1109/JCDL.2017.799164, pp. 1-10, 2017 (Year: 2017). [cited by applicant]
U.S. Appl. No. 17/340,511—Non-Final Office Action dated Jan. 11, 2022, 11 pages. [cited by applicant]
U.S. Appl. No. 17/099,441—Office Action dated Jul. 23, 2021, 13 pages. [cited by applicant]
U.S. Appl. No. 17/099,441—Response to Office Action dated Jul. 23, 2021 filed Oct. 22, 2021, 21 pages. [cited by applicant]
U.S. Appl. No. 17/340,511—Response to Non-Final Office Action dated Jan. 11, 2022 filed May 5, 2022, 34 pages. [cited by applicant]
U.S. Appl. No. 17/340,511—Notice of Allowance dated May 17, 2022, 14 pages. [cited by applicant]
U.S. Appl. No. 17/099,441—Notice of Allowance dated Dec. 8, 2021, 7 pages. [cited by applicant]
U.S. Appl. No. 17/099,441—Notice of Allowance dated Apr. 20, 2022, 16 pages. [cited by applicant]
U.S. Appl. No. 17/099,449, filed Nov. 16, 2020, U.S. Pat. No. 11,030,387, Jun. 8, 2021, Granted. [cited by applicant]
U.S. Appl. No. 17/340,511, filed Jun. 7, 2021, U.S. Pat. No. 11,449,663, Sep. 20, 2022, Granted. [cited by applicant]
U.S. Appl. No. 17/942,985, filed Sep. 12, 2022, U.S. Pat. No. 11,775,733, Oct. 3, 2023, Granted. [cited by applicant]
U.S. Appl. No. 18/374,565, filed Sep. 28, 2023, 20240104290, Mar. 28, 2024, Allowed. [cited by applicant]