IP Library › Granted Patent US 12,260,659
Granted Patent B2
US 12,260,659 · App. 17/660,639 · Granted Mar 25, 2025

Font attribute detection

Inventors: Ophir Azulai (Tivon, IL); Daniel Nechemia Rotman (Haifa, IL); Udi Barzelay (Haifa, IL)
Assignee: International Business Machines Corporation
G06V30/245G06F40/30G06V30/1444G06V30/153
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,659
App. No.
17/660,639
Granted
Mar 25, 2025
Kind
B2
Abstract

Described are techniques for font attribute detection. The techniques include receiving a document having different font attributes amongst a plurality of words respectively comprised of at least one character. The techniques further include generating a dense image document from the document by setting the plurality of words to a predefined size, removing blank spaces from the document, and altering an order of characters relative to the document. The techniques further include determining characteristics of the characters in the dense image document and aggregating the characteristics for at least one word. The techniques further include annotating the at least one word with a font attribute based on the aggregated characteristics.

Claims (57)

1. A computer-implemented method comprising:

receiving a document having different font attributes amongst a plurality of words respectively comprised of at least one character;

generating a dense image document from the document by:

setting the plurality of words to a predefined size;

removing blank spaces from the document; and

altering an order of characters relative to the document by placing the characters in a random order relative to the document;

determining characteristics of the characters in the dense image document;

aggregating the characteristics for at least one word; and

annotating the at least one word with a font attribute based on the aggregated characteristics.

2. The method of claim 1 , wherein the different font attributes are selected from a group consisting of: bold, italic, and underline.

3. The method of claim 1 , wherein receiving the document further comprises:

applying bounding boxes to respective words of the plurality of words in the document; and

applying bounding boxes to respective characters in the respective words of the plurality of words.

4. The method of claim 3 , wherein setting the plurality of words to the predefined size is performed by modifying a height the plurality of words while maintaining an aspect ratio of bounding boxes corresponding to the respective words of the plurality of words.

5. The method of claim 1 , wherein placing the characters in the random order causes at least one character with the font attribute and in a middle of the at least one word to be adjacent to a character without the font attribute in the dense image document.

6. The method of claim 1 , wherein determining the characteristics of the characters in the dense image document further comprises:

inputting the dense image document to a semantic segmentation model; and

receiving the characteristics of the characters as an output from the semantic segmentation model.

7. The method of claim 1 , wherein the characteristics of the characters comprise a ratio of pixels divided by a text area.

8. The method of claim 1 , wherein aggregating the characteristics for the at least one word comprises averaging the characteristics of a set of characters corresponding to the at least one word.

9. The method of claim 1 , wherein annotating the at least one word with the font attribute includes generating an annotated document, and wherein the method further comprises:

performing natural language processing (NLP) on the annotated document.

10. The method of claim 1 , wherein the method is performed by one or more computers according to software that is downloaded to the one or more computers from a remote data processing system, and wherein the method further comprises:

metering a usage of the software; and

generating an invoice based on metering the usage.

11. A system comprising:

one or more computer readable storage media storing program instructions; and

one or more processors which, in response to executing the program instructions, are configured to perform a method comprising:

receiving a document having different font attributes amongst a plurality of words respectively comprised of at least one character;

generating a dense image document from the document by:

setting the plurality of words to a predefined size;

removing blank spaces from the document; and

altering an order of characters relative to the document by placing the characters in a random order relative to the document;

determining characteristics of the characters in the dense image document;

aggregating the characteristics for at least one word; and

annotating the at least one word with a font attribute based on the aggregated characteristics.

12. The system of claim 11 , wherein the different font attributes are selected from a group consisting of: bold, italic, and underline.

13. The system of claim 11 , wherein placing the characters in the random order causes at least one character with the font attribute and in a middle of the at least one word to be adjacent to a character without the font attribute in the dense image document.

14. The system of claim 11 , wherein determining the characteristics of the characters in the dense image document further comprises:

inputting the dense image document to a semantic segmentation model; and

receiving the characteristics of the characters as an output from the semantic segmentation model.

15. The system of claim 11 , wherein the characteristics of the characters comprise a ratio of pixels divided by a text area.

16. A computer program product embodied on one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the program instructions comprising instructions configured to cause one or more processors to perform a method comprising:

receiving a document having different font attributes amongst a plurality of words respectively comprised of at least one character;

generating a dense image document from the document by:

setting the plurality of words to a predefined size;

removing blank spaces from the document; and

altering an order of characters relative to the document by placing the characters in a random order relative to the document;

determining characteristics of the characters in the dense image document;

aggregating the characteristics for at least one word; and

annotating the at least one word with a font attribute based on the aggregated characteristics.

17. The computer program product of claim 16 , wherein the different font attributes are selected from a group consisting of: bold, italic, and underline.

18. The computer program product of claim 16 , wherein placing the characters in the random order causes at least one character with the font attribute and in a middle of the at least one word to be adjacent to a character without the font attribute in the dense image document.

19. The computer program product of claim 16 , wherein determining the characteristics of the characters in the dense image document further comprises:

inputting the dense image document to a semantic segmentation model; and

receiving the characteristics of the characters as an output from the semantic segmentation model.

20. The computer program product of claim 16 , wherein the characteristics of the characters comprise a ratio of pixels divided by a text area.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2022
From: AZULAI, OPHIR; ROTMAN, DANIEL NECHEMIA; BARZELAY, UDI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 059732/0743 →
Continuity (1)
Related Publication 20230343124A1 · Oct 26, 2023
References Cited (20)
US 5668891A · Fan · 1997 [cited by applicant]
US 10984295B2 · Wang · 2021 [cited by applicant]
US 11481679B2 · Venkataraman Ganesh · 2022 [cited by examiner]
US 20200160050A1 · Bhotika · 2020 [cited by examiner]
TW 480457B · 2002 [cited by applicant]
GitHub; “Vasile-Peste/Typefont”, Downloaded Feb. 17, 22, 8 PGS. <https://github.com/Vasile-Peste/Typefont>. [cited by applicant]
GitHub_Hwalsuklee;, “Awesome-Deep-Text-Detection-Recognition”, Downloaded Feb. 17, 2022, 29 Pgs. <https://github.com/hwalsuklee/awesome-deep-text-detection-recognition>. [cited by applicant]
GitHub_Kovart; “Font-classificator”, Downloaded Feb. 17, 2022, 4 PGS, <https://github.com/kovart/font-classificator>. [cited by applicant]
GitHub_MGN0015095; “Font Recognition-Project”, Downloaded Feb. 17, 2022, 3 PGS, <https://github.com/MGN00150905/Font-Recognition-Project>. [cited by applicant]
GitHub_Sandeshhegda9; “Font-Style-recognition-Using-Neural_Networks”, Downloaded Feb. 17, 2022, 2 PGS, <https://github.com/sandeshhegde9/Font-Style-Recognition-using-Neural-Networks. [cited by applicant]
Mell et al., “The NIST Definition of Cloud Computing”, National Institute of Standards and Technology, Special Publication 800-145, Sep. 2011, 7 pages. [cited by applicant]
Saikrishna et al., “Script Independent Detection of Bold Words in Multi Font-size Documents”, 4 Pgs., Downloaded Feb. 10, 2022. [cited by applicant]
Tesseract: FontInfo Struct Reference, “3 API for font style”, Downloaded Feb. 10, 2022, 5 PGS, <https://tesseract-ocr.github.io/tessapi/3.x/a00399 . . . >. [cited by applicant]
Zramdini et al., “ApOFIS: an A priori Optical Font Identification System”, Institute of Informatics, University of Fribourg, 6 Pgs, Published 1995, Downloaded Feb. 10, 2022, <https://dblp.org/rec/conf/iciap/Zramdinil95.… [cited by applicant]
Ami Mehta et al. “Multifont Multisize Gujarati OCR with Style Identification” 2017 IEEE 7 pages. [cited by applicant]
Goncalves et al. “Real-time automatice License Plate Recognition Through Deep Multitask Networks” 2018 IEEE, pp. 110-117. [cited by applicant]
J. Dholakia et al “Zone Identification in the Printed Gujarati Text” IEEE 2005 5 pages. [cited by applicant]
Radwan et al. “Neural Networks Pipeline for Offline Machine Printed Arabic OCR” Springer Science & Business Media LLC Oct. 27, 2017 19 pages. [cited by applicant]
Sandu et al. “Context Sensitive Transformer for Bold Worlds Classification” arxiv 2205.07683V1 May 16, 2022, 5 pages. [cited by applicant]
International Search Report and Written Opinion for Application PCT/IB2023/052851, Jun. 12, 2023, 14 pages. [cited by applicant]