IP Library › Granted Patent US 12,412,023
Granted Patent B2
US 12,412,023 · App. 18/650,928 · Granted Sep 9, 2025

Identifying and formatting headers for text content

Inventors: Sagar Gollamudi (San Diego, CA); Vishank Bhatia (Sunnyvale, CA); Xu Zhong (Vermont South, AU); Thanh Long Duong (Point Cook, AU); Mark Johnson (Castle Cove, AU); Srinivasa Phani Kumar Gadde (Fremont, CA); Vishal Vishnoi (Redwood City, CA)
Assignee: Oracle International Corporation
G06F40/103
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,412,023
App. No.
18/650,928
Granted
Sep 9, 2025
Kind
B2
Abstract

A data corpus is partitioned into text strings for header classification. A group characteristic is computed for a text string, and whether the group characteristic satisfies a group characteristic criterion is determined. The text string may be disqualified from header classification if the group characteristic criterion is not satisfied, or one or more font characteristics may be determined for the text string if the group characteristic criterion is satisfied. A font characteristic that meets one or more prevalence criteria may be identified and evaluated to determine whether the font characteristic meets at least one font characteristic criterion. The text string may be disqualified from header classification if the font characteristic criterion is not satisfied, or if the font characteristic meets the font characteristic criterion, the text string is classified as a header, and tagged content is generated by applying a header tag to the text string.

Claims (108)

1. One or more non-transitory computer-readable media comprising computer-executable instructions that, when executed by one or more processors, cause performance of operations, comprising:

identifying, in a data corpus, a first candidate text string for header classification;

determining a first font of the first candidate text string;

classifying the first candidate text string as a first header based at least in part on a first evaluation of the first font of the first candidate text string relative to one or more additional text strings in the data corpus;

applying, to the first header, a first header tag at least in part responsive to classifying the first candidate text string as the first header;

rendering the first header for display as a heading for a first set of related text strings in the data corpus.

2. The one or more non-transitory computer-readable media of claim 1 , wherein classifying the first candidate text string as the first header based at least in part on the first evaluation of the first font comprises:

evaluating a first prominence of the first font of the first candidate text string relative to one or more additional text strings in the data corpus;

classifying the first candidate text string as the first header based at least in part on the first prominence of the first font of the first candidate text string relative to one or more additional text strings in the data corpus.

3. The one or more non-transitory computer-readable media of claim 2 , wherein the operations further comprise:

identifying, in the data corpus, a second candidate text string for header classification;

classifying the second candidate text string as a second header based at least in part on a second prominence of a second font of the second candidate text string;

assigning, to the first header, a first header format based at least in part on the first prominence of the first font relative to the second prominence of the second font;

selecting the first header tag for the first header based at least in part on the first header format corresponding to the first header.

4. The one or more non-transitory computer-readable media of claim 3 , wherein the operations further comprise:

determining, for the first header, a first position in a font prominence hierarchy based at least in part on the first font;

determining, for the second header, a second position in the font prominence hierarchy based at least in part on the second font;

assigning to the second header, a second header format based at least in part on the second position in the font prominence hierarchy relative to the first position in the font prominence hierarchy.

5. The one or more non-transitory computer-readable media of claim 4 , wherein the operations further comprise:

based at least on a difference in prominence between the first candidate text string and the second candidate text string, determining that the first candidate text string is a first heading and determining that the second candidate text string is a subheading under the first heading;

applying, to the second header, a second header tag corresponding to the second header format;

rendering the second header for display as the subheading for a second set of related text strings in the data corpus.

6. The one or more non-transitory computer-readable media of claim 4 , wherein the operations further comprise:

assigning to the first header, the first header format from a header formatting schedule based at least in part on the first position in the font prominence hierarchy relative to the second position in the font prominence hierarchy,

wherein the first header format corresponds to a particular header level in the header formatting schedule, and wherein the second header format corresponds to an incremental step from the particular header level in the header formatting schedule.

7. The one or more non-transitory computer-readable media of claim 4 , wherein the font prominence hierarchy comprises an order of prominence for at least two of: a bold font style, an italic font style, a first font size, a second font size, a first font type, a second font type, a first font case, a second font case, a first font color, or a second font color.

8. The one or more non-transitory computer-readable media of claim 2 , wherein the operations further comprise:

assigning, to the first header, a first header format based at least in part on the first prominence of the first font;

selecting the first header tag for the first header based at least in part on the first header format corresponding to the first header, wherein the first header format comprises one or more font characteristics that differ from the first font.

9. The one or more non-transitory computer-readable media of claim 2 , wherein the operations further comprise:

identifying, in the data corpus, a second candidate text string for header classification;

computing, for the second candidate text string, a group characteristic comprising a number of elements in the second candidate text string;

determining that the group characteristic meets a set of one or more group characteristic criteria;

classifying the second candidate text string as a second header at least in part responsive to determining that the group characteristic meets the set of one or more group characteristic criteria.

10. The one or more non-transitory computer-readable media of claim 9 , wherein the operations further comprise:

responsive to determining that the group characteristic meets the set of one or more group characteristic criteria:

identifying, for the second candidate text string, a font characteristic that meets a set of one or more prevalence criteria;

determining that the font characteristic meets a set of one or more font characteristic criteria;

classifying the second candidate text string as the second header at least in part responsive to determining that the font characteristic meets the set of one or more font characteristic criteria.

11. The one or more non-transitory computer-readable media of claim 10 ,

wherein the set of one or more prevalence criteria comprises a prevalence threshold for the font characteristic,

wherein the prevalence threshold comprises a threshold number of elements that exhibit the font characteristic.

12. The one or more non-transitory computer-readable media of claim 10 , wherein the font characteristic comprises at least one of: bold characters, italic characters, underline characters, enlarged characters relative to a size threshold, a distinctive font type according to a distinctive font schedule, uppercase characters, capitalized characters, color characters, or a distinctive color according to a difference in Euclidean distance between color tuples.

13. The one or more non-transitory computer-readable media of claim 9 , wherein determining that the group characteristic meets the set of one or more group characteristic criteria comprises at least one of:

determining that a first number of elements in the second candidate text string meets a threshold number of elements, or

determining that a kind of content in the second candidate text string meets a set of one or more content criteria.

14. The one or more non-transitory computer-readable media of claim 1 , wherein the operations further comprise:

identifying, in the data corpus, a second candidate text string for header classification;

computing a group characteristic for the second candidate text string;

determining that a set of one or more group characteristic criteria are unmet by the group characteristic;

disqualifying the second candidate text string for header classification responsive to determining that the set of one or more group characteristic criteria are unmet by the group characteristic.

15. The one or more non-transitory computer-readable media of claim 14 , wherein the operations further comprise:

identifying, in the data corpus, a third candidate text string for header classification;

identifying, for the third candidate text string, a font characteristic that meets a set of one or more prevalence criteria;

determining that a set of one or more font characteristic criteria are unmet by the font characteristic;

disqualifying the third candidate text string for header classification responsive to determining that the set of one or more font characteristic criteria are unmet by the font characteristic.

16. The one or more non-transitory computer-readable media of claim 1 , wherein the operations further comprise:

partitioning the data corpus into a plurality of candidate text strings based on a set of one or more lexical rules, wherein the set of one or more lexical rules comprises: at least one context-free lexical rule, or at least one context-sensitive lexical rule.

17. The one or more non-transitory computer-readable media of claim 16 , wherein partitioning the data corpus into the plurality of candidate text strings comprises:

parsing content configured according to a static page description language to identify the plurality of candidate text strings; and

generating a data structure comprising the plurality of candidate text strings.

18. The one or more non-transitory computer-readable media of claim 1 , wherein the operations further comprise:

receiving a query,

identifying in the data corpus, based at least in part on the query, the first header and the first set of related text strings;

responsive at least in part to identifying the first header and the first set of related text strings in the data corpus, rendering the first header and the first set of related text strings for display on a user interface device.

19. A method, comprising:

identifying, in a data corpus, a first candidate text string for header classification;

determining a first font of the first candidate text string;

classifying the first candidate text string as a first header based at least in part on a first evaluation of the first font of the first candidate text string relative to one or more additional text strings in the data corpus;

applying, to the first header, a first header tag at least in part responsive to classifying the first candidate text string as the first header;

rendering the first header for display as a heading for a first set of related text strings in the data corpus;

wherein the method is performed by at least one device including a hardware processor.

20. A system, comprising:

at least one hardware processor;

wherein the system is configured to execute operations, using the at least one hardware processor, the operations comprising:

identifying, in a data corpus, a first candidate text string for header classification;

determining a first font of the first candidate text string;

classifying the first candidate text string as a first header based at least in part on a first evaluation of the first font of the first candidate text string relative to one or more additional text strings in the data corpus;

applying, to the first header, a first header tag at least in part responsive to classifying the first candidate text string as the first header;

rendering the first header for display as a heading for a first set of related text strings in the data corpus.

21. The one or more non-transitory computer-readable media of claim 1 , wherein classifying the first candidate text string as the first header comprises:

determining that the first candidate text string comprises one or more attributes that are not included in the one or more additional text strings, the one or more attributes of the first candidate text string comprising one or more of:

bold characters, italic characters, underline characters, enlarged characters relative to a size threshold, a distinctive font type according to a distinctive font schedule, uppercase characters, capitalized characters, color characters, or a distinctive color according to a distinctive color schedule.

22. The one or more non-transitory computer-readable media of claim 1 , wherein the operations further comprise:

assigning to the first header, a first header format from a header formatting schedule based at least in part on the first header having a first position in a font prominence hierarchy.

23. The one or more non-transitory computer-readable media of claim 22 , wherein the first header format corresponds to a particular header level in the header formatting schedule.

24. The one or more non-transitory computer-readable media of claim 22 , wherein the operations further comprise:

assigning to a second header, a second header format from the header formatting schedule based at least in part on the second header having a second position in the font prominence hierarchy.

25. The one or more non-transitory computer-readable media of claim 24 , wherein the second header format corresponds to an incremental step from the first header format in the header formatting schedule.

26. The one or more non-transitory computer-readable media of claim 1 , wherein the operations further comprise:

identifying, for the first candidate text string, a first font characteristic that meets a set of one or more prevalence criteria;

determining that the first font characteristic meets a set of one or more font characteristic criteria;

classifying the first candidate text string as the first header at least in part responsive to determining that the first font characteristic, that meets the set of one or more prevalence criteria, meets the set of one or more font characteristic criteria;

wherein the set of one or more prevalence criteria comprises a threshold number of elements that exhibit the first font characteristic.

27. The one or more non-transitory computer-readable media of claim 26 , wherein the operations further comprise:

identifying, in the data corpus, a second candidate text string for header classification;

identifying, for the second candidate text string, a second font characteristic that meets the set of one or more prevalence criteria;

determining that the set of one or more font characteristic criteria are unmet by the second font characteristic;

disqualifying the second candidate text string for header classification responsive to determining that the set of one or more font characteristic criteria are unmet by the second font characteristic.

28. The one or more non-transitory computer-readable media of claim 1 , wherein the operations further comprise:

computing, for the first candidate text string, a group characteristic comprising a number of elements in the first candidate text string;

determining that the group characteristic meets a set of one or more group characteristic criteria;

classifying the first candidate text string as the first header at least in part responsive to determining that the group characteristic meets the set of one or more group characteristic criteria;

wherein the set of one or more group characteristic criteria comprises a first number of elements in the first candidate text string meeting a threshold number of elements.

29. The one or more non-transitory computer-readable media of claim 1 , wherein the operations further comprise:

partitioning the data corpus into a plurality of candidate text strings based at least in part on a context-free lexical rule.

30. The one or more non-transitory computer-readable media of claim 1 , wherein the operations further comprise:

partitioning the data corpus into a plurality of candidate text strings based at least in part on a context-sensitive lexical rule.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2024
From: GOLLAMUDI, SAGAR; BHATIA, VISHANK; ZHONG, XU; DUONG, THANH LONG; JOHNSON, MARK; GADDE, SRINIVASA PHANI KUMAR; VISHNOI, VISHAL
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 067410/0933 →
Continuity (2)
Continuation 18334238 · Jun 13, 2023
Related Publication 20240419886A1 · Dec 19, 2024
References Cited (29)
US 7725499B1 · von Lepel · 2010 [cited by examiner]
US 8910036B1 · Cromwell · 2014 [cited by examiner]
US 9805073B1 · Davis · 2017 [cited by applicant]
US 9952763B1 · Bi · 2018 [cited by applicant]
US 10467338B2 · Hasan et al. · 2019 [cited by applicant]
US 10909360B2 · Omiya et al. · 2021 [cited by applicant]
US 11327646B2 · Tran et al. · 2022 [cited by applicant]
US 11347945B2 · Sato · 2022 [cited by applicant]
US 20150254530A1 · Gulwani · 2015 [cited by examiner]
US 20150294591A1 · Kullok et al. · 2015 [cited by applicant]
US 20170091151A1 · Jones et al. · 2017 [cited by applicant]
US 20170147202A1 · Donohue · 2017 [cited by applicant]
US 20190317993A1 · Toda · 2019 [cited by applicant]
US 20200151244A1 · Rastogi et al. · 2020 [cited by applicant]
US 20210110153A1 · Gupta et al. · 2021 [cited by applicant]
US 20210206481A1 · Brion et al. · 2021 [cited by applicant]
US 20220101060A1 · Wang · 2022 [cited by examiner]
US 20220366131A1 · Ekron · 2022 [cited by examiner]
US 20230040725A1 · Javeri · 2023 [cited by examiner]
US 20230334242A1 · Dadoo et al. · 2023 [cited by applicant]
CN 106951400A · 2017 [cited by applicant]
Aliyu et al., “SED: An Algorithm for Automatic Identification of Section and Subsection Headings in Text Documents”, IJCSI International Journal of Computer Science Issues, vol. 17, No. 6, Nov. 2020, pp. 40-47. [cited by applicant]
Bruijn L., “Extracting headers and paragraphs from pdf using PyMuPDF”, Retrieved from https://towardsdatascience.com/extracting-headers-and-paragraphs-from-pdf-using-pymupdf-676e8421c467, Apr. 9, 2020, pp. 1-8. [cited by applicant]
Budhiraja et al., “A Supervised Learning Approach For Heading Detection”, Sep. 2018, pp. 1-20. [cited by applicant]
Could clustering be used to parse pdf documents to get headings and titles?, Retrieved from https://ai.stackexchange.com/questions/20352/could-clustering-be-used-to-parse-pdf-documents-to-get-headings-and-titles, Retrie… [cited by applicant]
Hofer C., “Development of a structure-aware PDF parser”, Retrieved from https://medium.com/@_chriz_/development-of-a-structure-aware-pdf-parser-7285f3fe41a9, September 6. 2020, pp. 1-9. [cited by applicant]
IText-PDF reading issue on heading levels ( h1-h6 ), Retrieved from https://stackoverflow.com/questions/30001953/itext-pdf-reading-issue-on-heading-levels-h1-h6, Retrieved on Jan. 13, 2023, pp. 1-5. [cited by applicant]
Knowledge Extraction, Kore.ai Documentation v7.1, Retrieved from https://developer.kore.ai/v7-1/docs/bots/bot-builder-tool/knowledge-task/knowledge-extraction-service/, Retrieved on Jan. 13, 2023, pp. 1-6. [cited by applicant]
Vanderbeck S. et al., “A Machine Learning Approach to Identifying Sections in Legal Briefs”, Midwest Artificial Intelligence and Cognitive Science Conference, 2011, pp. 7. [cited by applicant]