IP Library Granted Patent US 12,242,806
Granted Patent B2
US 12,242,806 · App. 18/451,153 · Granted Mar 4, 2025

Systems and methods for structure and header extraction

Inventor: Richard Anthony Pito (Toronto, CA)
Assignee: Thomson Reuters Enterprise Centre GmbH
G06F40/279G06F3/0481G06F40/109G06F40/137G06F40/166G06F40/232G06F40/242G06F40/258G06F40/284G06F40/289G06V30/416
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,242,806
App. No.
18/451,153
Granted
Mar 4, 2025
Kind
B2
Abstract

The present disclosure is directed towards systems and methods for extracting structure and headers from a body of text. This computational extraction is based on the visual and logical similarities between portions of text. Structure is derived from a programmatic and methodic computation of similarities between header pairs.

Claims (41)

1. A system comprising memory and one or more processors communicatively coupled to the memory, the one or more processors configured to:

identify an orthography feature related to orthography characteristics for a respective header from a plurality of headers included in a document, wherein the orthography feature is determined based on capitalization formatting for a string of characters in a data chunk that corresponds to the respective header included in the document;

identify two or more typography features related to typography characteristics for the respective header from the plurality of headers included in the document, wherein the two or more typography features comprise at least (a) a first typography feature configured with a first binary value or a second binary value based on a font setting for at least one character in the data chunk and (b) a second typography feature configured based on a page layout setting for the document;

generate a graph representation of the plurality of headers based at least in part on the orthography feature and the two or more typography features, wherein respective vertices of the graph representation correspond to the respective headers;

compare the graph representation of the plurality of headers to a predefined graph representation to determine a performance metric for the graph representation; and

perform one or more computer-implemented processing tasks with respect to the document based at least in part on the graph representation and the performance metric.

2. The system of claim 1 , wherein the one or more processors are further configured to:

determine a similarity between pairs of headers from the plurality of headers based at least in part on the orthography feature and the two or more typography features; and

transform the graph representation of the plurality of headers into a transformed graph representation of the plurality of headers that segments the headers into groups of one or more similar adjacent headers based on orthography similarities and typography similarities between adjacent headers in the plurality of headers.

3. The system of claim 2 , wherein the one or more processors are further configured to:

match non-adjacent groups of similar adjacent headers based on the groups of the one or more similar adjacent headers associated with the transformed graph representation of the plurality of headers.

4. The system of claim 1 , wherein the two or more typography features further comprise a third typography feature configured with a non-binary value based on the font setting for the at least one character in the data chunk.

5. The system of claim 1 , wherein a path through the graph representation of the plurality of headers represents a sequence of headers included in the document.

6. The system of claim 1 , wherein the graph representation is configured as a tree-structured hierarchy of the plurality of headers included in the document.

7. A method, comprising:

identifying an orthography feature related to orthography characteristics for a respective header from a plurality of headers included in a document, wherein the orthography feature is determined based on capitalization formatting for a string of characters in a data chunk that corresponds to the respective header included in the document;

identifying two or more typography features related to typography characteristics for the respective header from the plurality of headers included in the document, wherein the two or more typography features comprise at least (a) a first typography feature configured with a first binary value or a second binary value based on a font setting for at least one character in the data chunk and (b) a second typography feature configured based on a page layout setting for the document;

generating a graph representation of the plurality of headers based at least in part on the orthography feature and the two or more typography features, wherein respective vertices of the graph representation correspond to the respective headers;

comparing the graph representation of the plurality of headers to a predefined graph representation to determine a performance metric for the graph representation; and

performing one or more computer-implemented processing tasks with respect to the document based at least in part on the graph representation and the performance metric.

8. The method of claim 7 , further comprising:

determining a similarity between pairs of headers from the plurality of headers based at least in part on the orthography feature and the two or more typography features; and

transforming the graph representation of the plurality of headers into a transformed graph representation of the plurality of headers that segments the headers into groups of one or more similar adjacent headers based on orthography similarities and typography similarities between adjacent headers in the plurality of headers.

9. The method of claim 8 , further comprising:

matching non-adjacent groups of similar adjacent headers based on the groups of the one or more similar adjacent headers associated with the transformed graph representation of the plurality of headers.

10. The method of claim 7 , wherein the two or more typography features further comprise a third typography feature configured with a non-binary value based on the font setting for the at least one character in the data chunk.

11. The method of claim 7 , wherein a path through the graph representation of the plurality of headers represents a sequence of headers included in the document.

12. The method of claim 7 , wherein the graph representation is configured as a tree-structured hierarchy of the plurality of headers included in the document.

13. A computer program product, stored on a non-transitory computer readable storage medium, comprising instructions that when executed by one or more processors cause the one or more processors to:

identify an orthography feature related to orthography characteristics for a respective header from a plurality of headers included in a document, wherein the orthography feature is determined based on capitalization formatting for a string of characters in a data chunk that corresponds to the respective header included in the document;

identify two or more typography features related to typography characteristics for the respective header from the plurality of headers included in the document, wherein the two or more typography features comprise at least (a) a first typography feature configured with a first binary value or a second binary value based on a font setting for at least one character in the data chunk and (b) a second typography feature configured based on a page layout setting for the document;

generate a graph representation of the plurality of headers based at least in part on the orthography feature and the two or more typography features, wherein respective vertices of the graph representation correspond to the respective headers;

compare the graph representation of the plurality of headers to a predefined graph representation to determine a performance metric for the graph representation; and

perform one or more computer-implemented processing tasks with respect to the document based at least in part on the graph representation and the performance metric.

14. The computer program product of claim 13 , wherein the one or more processors are further configured to:

determine a similarity between pairs of headers from the plurality of headers based at least in part on the orthography feature and the two or more typography features; and

transform the graph representation of the plurality of headers into a transformed graph representation of the plurality of headers that segments the headers into groups of one or more similar adjacent headers based on orthography similarities and typography similarities between adjacent headers in the plurality of headers.

15. The computer program product of claim 14 , wherein the one or more processors are further configured to:

match non-adjacent groups of similar adjacent headers based on the groups of the one or more similar adjacent headers associated with the transformed graph representation of the plurality of headers.

16. The computer program product of claim 13 , wherein the two or more typography features further comprise a third typography feature configured with a non-binary value based on the font setting for the at least one character in the data chunk.

17. The computer program product of claim 13 , wherein a path through the graph representation of the plurality of headers represents a sequence of headers included in the document.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2023
From: THOMSON REUTERS CANADA LIMITED
To: THOMSON REUTERS ENTERPRISE CENTRE GMBH
Reel/Frame 065305/0353 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2023
From: PITO, RICHARD ANTHONY
To: THOMSON REUTERS CANADA LIMITED
Reel/Frame 065305/0408 →
Continuity (6)
Continuation 17156546 · Jan 23, 2021
Provisional Application 62975514 · Feb 12, 2020
Provisional Application 62965516 · Jan 24, 2020
Provisional Application 62965523 · Jan 24, 2020
Provisional Application 62965520 · Jan 24, 2020
Related Publication 20240054286A1 · Feb 15, 2024
References Cited (99)
US 6298357B1 · Wexler · 2001 [cited by applicant]
US 7672022B1 · Fan · 2010 [cited by applicant]
US 7788580B1 · Goodwin · 2010 [cited by applicant]
US 7797622B2 · Dejean · 2010 [cited by applicant]
US 7937653B2 · Dejean · 2011 [cited by applicant]
US 8413048B1 · Goodwin · 2013 [cited by applicant]
US 8631097B1 · Seo · 2014 [cited by applicant]
US 8898296B2 · Zeng · 2014 [cited by applicant]
US 8914720B2 · Harrington · 2014 [cited by applicant]
US 9218326B2 · Dejean · 2015 [cited by applicant]
US 9336202B2 · Khan · 2016 [cited by applicant]
US 9483463B2 · Galle et al. · 2016 [cited by applicant]
US 9514499B1 · Kogut-O'Connell et al. · 2016 [cited by applicant]
US 10019488B2 · Levy · 2018 [cited by applicant]
US 10049270B1 · Agarwalla · 2018 [cited by applicant]
US 10318568B2 · Yasue · 2019 [cited by applicant]
US 10460162B2 · Gelosi · 2019 [cited by applicant]
US 10614113B2 · Peled · 2020 [cited by applicant]
US 10726198B2 · Gelosi · 2020 [cited by applicant]
US 10878195B2 · Duta · 2020 [cited by applicant]
US 10885282B2 · Ilic · 2021 [cited by applicant]
US 11017426B1 · Garg · 2021 [cited by applicant]
US 11023675B1 · Neervannan · 2021 [cited by applicant]
US 11042555B1 · Kane · 2021 [cited by applicant]
US 11048711B1 · Fleming · 2021 [cited by applicant]
US 11170759B2 · Sim · 2021 [cited by applicant]
US 11205043B1 · Neervannan · 2021 [cited by applicant]
US 11205044B1 · Neervannan · 2021 [cited by applicant]
US 11232114B1 · Fleming · 2022 [cited by applicant]
US 11256856B2 · Gelosi · 2022 [cited by applicant]
US 11312956B2 · Geng · 2022 [cited by applicant]
US 11475209B2 · Gelosi · 2022 [cited by applicant]
US 11727708B2 · Geng · 2023 [cited by applicant]
US 11763079B2 · Pito · 2023 [cited by examiner]
US 20040049462A1 · Wang · 2004 [cited by applicant]
US 20040162827A1 · Nakano · 2004 [cited by applicant]
US 20060156226A1 · Dejean · 2006 [cited by applicant]
US 20080077588A1 · Zhang · 2008 [cited by applicant]
US 20080114757A1 · Dejean · 2008 [cited by applicant]
US 20110029952A1 · Harrington · 2011 [cited by applicant]
US 20110055206A1 · Martin et al. · 2011 [cited by applicant]
US 20110075932A1 · Komaki · 2011 [cited by applicant]
US 20110145701A1 · Dejean · 2011 [cited by applicant]
US 20110197121A1 · Kletter · 2011 [cited by applicant]
US 20110216975A1 · Rother · 2011 [cited by applicant]
US 20120278321A1 · Traub · 2012 [cited by applicant]
US 20120297025A1 · Zeng · 2012 [cited by applicant]
US 20130174029A1 · O'Sullivan et al. · 2013 [cited by applicant]
US 20130191366A1 · Jovanovic · 2013 [cited by applicant]
US 20130311169A1 · Khan · 2013 [cited by applicant]
US 20140074455A1 · Galle et al. · 2014 [cited by applicant]
US 20140101456A1 · Meunier et al. · 2014 [cited by applicant]
US 20140337719A1 · Xu · 2014 [cited by applicant]
US 20150067476A1 · Song · 2015 [cited by applicant]
US 20150100308A1 · Bedrax-Weiss et al. · 2015 [cited by applicant]
US 20150161102A1 · Gidney · 2015 [cited by applicant]
US 20150169676A1 · Bohra · 2015 [cited by applicant]
US 20160048520A1 · Levy · 2016 [cited by applicant]
US 20160224662A1 · King et al. · 2016 [cited by applicant]
US 20170011313A1 · Pochert et al. · 2017 [cited by applicant]
US 20170017641A1 · Gidney · 2017 [cited by applicant]
US 20170052934A1 · Hatsutori · 2017 [cited by applicant]
US 20170103466A1 · Syed · 2017 [cited by applicant]
US 20170329846A1 · Dole · 2017 [cited by applicant]
US 20170351688A1 · Yasue · 2017 [cited by applicant]
US 20180039907A1 · Kraley · 2018 [cited by applicant]
US 20180096060A1 · Peled · 2018 [cited by applicant]
US 20180260378A1 · Theodore et al. · 2018 [cited by applicant]
US 20180268506A1 · Wodetzki et al. · 2018 [cited by applicant]
US 20180300315A1 · Leal · 2018 [cited by applicant]
US 20190114479A1 · Gelosi · 2019 [cited by applicant]
US 20190155944A1 · Mahata et al. · 2019 [cited by applicant]
US 20190220503A1 · Gelosi · 2019 [cited by applicant]
US 20190272421A1 · Sugaya · 2019 [cited by applicant]
US 20190278853A1 · Chen · 2019 [cited by applicant]
US 20190340240A1 · Duta · 2019 [cited by applicant]
US 20190347284A1 · Roman et al. · 2019 [cited by applicant]
US 20200026916A1 · Wood et al. · 2020 [cited by applicant]
US 20200043113A1 · DePalma et al. · 2020 [cited by applicant]
US 20200097759A1 · Nadim · 2020 [cited by applicant]
US 20200104957A1 · Guo et al. · 2020 [cited by applicant]
US 20200110800A1 · Astigarraga et al. · 2020 [cited by applicant]
US 20200184013A1 · Ilic · 2020 [cited by applicant]
US 20200219481A1 · Sim · 2020 [cited by applicant]
US 20200226510A1 · Gupta · 2020 [cited by applicant]
US 20200311412A1 · Prebble · 2020 [cited by applicant]
US 20200327151A1 · Coquard et al. · 2020 [cited by applicant]
US 20200349199A1 · Jayaraman · 2020 [cited by applicant]
US 20200364291A1 · Bentabet · 2020 [cited by applicant]
US 20210117667A1 · Mehra · 2021 [cited by applicant]
US 20210150128A1 · Gelosi · 2021 [cited by applicant]
US 20210201013A1 · Makhija et al. · 2021 [cited by applicant]
US 20210279238A1 · Kane · 2021 [cited by applicant]
US 20220230465A1 · Geng · 2022 [cited by applicant]
US 20220277140A1 · Rhim · 2022 [cited by applicant]
U.S. Appl. No. 17/156,546, filed Jan. 23, 2021, 2021/0319177, Patented. [cited by applicant]
Cai et al., “ [cited by applicant]
Pomikalek, Jan, “Removing Boilerplate and Duplicate Content from Web Corpora.” 108 pages, (2011), https://is.muni.cz/th/45523/f_d/phdthesis.pdf. [cited by applicant]
Qin Liu, “ [cited by applicant]