IP Library Granted Patent US 10,885,282
Granted Patent B2
US 10,885,282 · App. 16/212,907 · Granted Jan 5, 2021

Document heading detection

Inventors: Andreja Ilić (Belgrade, RS); Katarina Jovanović (Belgrade, RS); Milo{hacek over (s)} Ra{hacek over (s)}ković (Belgrade, RS); Vladimir Ranković (Belgrade, RS)
Assignee: MICROSOFT TECHNOLOGY LICENSING, LLC
G06F40/30G06F40/211
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,885,282
App. No.
16/212,907
Granted
Jan 5, 2021
Kind
B2
Abstract

Document heading detection includes performing a classification on each of a plurality of paragraphs of a document to identify each paragraph as either a heading or non-heading paragraph. The classification is based on one or more pre-established values corresponding to one or more pre-established formatting features that are indicative of a heading paragraph relative to currently established values for each of the one or more pre-established formatting features in each of the plurality of paragraphs. Document heading detection further includes determining a strength of each of the one or more heading paragraphs by performing a linear regression on each heading paragraph and assigning each of the one or more heading paragraphs a heading level within a hierarchy of heading levels based on the determined strength.

Claims (37)

1. A method for detecting document headings:

receiving a plurality of paragraphs in a document created with an authoring application, wherein each of the plurality of paragraphs includes content;

performing a binary classification analysis on the content of each of the plurality of paragraphs to identify each paragraph as either a heading paragraph or a non-heading paragraph, the binary classification analysis assessing the content of each of the plurality of paragraphs through use of a boosted decision tree containing a plurality of content features having been pre-determined from historical content as indicative of a heading paragraph;

performing a heading level classification for at least two paragraphs identified as heading paragraphs as being at one of a plurality of heading levels in a heading hierarchy, the heading level classification including:

determining a strength value for the at least two paragraph identified as heading paragraphs based on a subset of the plurality of content features; and

assigning the at least two heading paragraphs to one of the heading levels based on the determined strength value.

2. The method of claim 1 , wherein the plurality of paragraphs are received in real time and wherein performing the binary classification analysis and the heading level classification are updated upon receipt of each new paragraph in real time.

3. The method of claim 1 , wherein the plurality of content features includes a direct formatting feature and a relative formatting feature.

4. The method of claim 3 , wherein the plurality of content features includes a syntactical feature and a semantical feature.

5. The method of claim 3 , wherein the direct formatting feature includes one or more of: bold, italic, underline, uppercase, font size, indentation, outline level, or alignment.

6. The method of claim 5 , wherein the relative formatting feature includes one or more of: different color than next, font size relative to next, normalized font size, indentation compared to next, indentation compared to previous, normalized indentation, distance to neighbors, or followed by bulleted list.

7. The method of claim 4 , wherein the syntactical feature includes one or more of: part of a bulleted list, starts with a number, sentence count, word count, ends with a colon, percentage of non-alphanumeric characters, number of tabs, number of empty paragraphs before, number of empty paragraphs after, ends with punctuation, text length, text length compared to previous, text length compared to next.

8. The method of claim 4 , wherein the semantical feature is determined using a term frequency-inverse document frequency analysis.

9. The method of claim 1 , where the subset of the plurality of content features includes: bold, italic, underline, uppercase, font size and indentation.

10. The method of claim 1 , wherein one or more of the plurality of paragraphs is associated with a paragraph style that is predefined by the authoring application and wherein performing the binary classification analysis and the heading level classification are performed without regard to the predefined paragraph style.

11. The method of claim 1 , wherein the document has a format associated with the authoring application and wherein the method further comprises: using the one or more heading paragraphs and their assigned heading levels to perform one or more of: converting the document from the format to a different format associated with a different authoring application, generating a table of contents for the document, generating an outline for the document or generating a navigational map for the document.

12. The method of claim 1 , wherein:

determining a strength value for each of the at least two paragraphs identified as heading paragraphs includes performing a linear regression on the at least two paragraphs based on the subset of the plurality of content features.

13. The method of claim 1 , wherein assigning each of the at least two heading paragraphs to one of the heading levels based on the determined strength values includes dividing the at least two heading paragraphs, using thresholding of the determined strength values, into a plurality of clusters with each of the plurality of clusters corresponding to one of the heading levels.

14. The method of claim 13 , wherein thresholding of the determined strength values comprises a minimal sum variance thresholding of the determined strength values.

15. The method of claim 1 , wherein each of at least a portion of the plurality of content features of the boosted decision tree are associated with a pre-determined strength value, the pre-determined strength values used in determining the strength value for each of the at least two paragraphs identified as heading paragraphs.

16. A system to detect document headings, the system comprising:

a memory storing executable instructions; and a processor, wherein when executing the executable instructions, the processor is caused to:

receive a plurality of paragraphs in a document created with an authoring application, wherein each of the plurality of paragraphs includes content;

perform a binary classification analysis on the content of each of the plurality of paragraphs to identify each paragraph as either a heading paragraph or a non-heading paragraph, the binary classification analysis assessing the content of each of the plurality of paragraphs through use of a boosted decision tree containing a plurality of content features having been pre-determined from historical content as indicative of a heading paragraph;

perform a heading level classification for at least two paragraphs identified as heading paragraphs as being at one of a plurality of heading levels in a heading hierarchy, wherein performance of the heading level classification includes:

determination of a strength value for the at least two paragraphs identified as heading paragraphs based on a subset of the plurality of content features; and

assignment of the at least two heading paragraphs to one of the heading levels based on the determined strength value.

17. The system of claim 16 , wherein the plurality of paragraphs are received in real time and wherein performance of the binary classification analysis and the heading level classification is updated upon receipt of each new paragraph in real time.

18. The system of claim 16 , wherein the plurality of content features includes one or more of a direct formatting feature, a relative formatting feature, a syntactical feature and a semantical feature.

19. The system of claim 16 , wherein the subset of the plurality of features comprises a plurality of direct formatting features.

20. A computer storage media that stores computer-executable instructions, the instructions direct a computer to:

receive a plurality of paragraphs in a document created with an authoring application, wherein each of the plurality of paragraphs includes content;

perform a binary classification analysis on the content of each of the plurality of paragraphs to identify each paragraph as either a heading paragraph or a non-heading paragraph, the binary classification analysis assessing the content of each of the Plurality of paragraphs through use of a boosted decision tree containing a plurality of content features having been pre-determined from historical content as indicative of a heading paragraph;

perform a heading level classification for at least two paragraphs identified as heading paragraphs as being at one of a plurality of heading levels in a heading hierarchy, wherein performance of the heading level classification includes:

determination of a strength value for the at least two Paragraphs identified as heading paragraphs based on a subset of the plurality of content features; and

assign the at least two heading paragraphs to one of the heading levels based on the determined strength.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2018
From: ILIC, ANDREJA
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 047724/0933 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 7, 2018
From: RANKOVIC, VLADIMIR; JOVANOVIC, KATARINA; RA?KOVIC, MILO?
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 047703/0607 →
Continuity (1)
Related Publication 20200184013A1 · Jun 11, 2020
Cited By (5)
US 12,190,059 US 12,242,806 US 12,462,111 US 12,626,058 US 12,705,420