IP Library › Granted Patent US 12,567,276
Granted Patent B2
US 12,567,276 · App. 18/459,460 · Granted Mar 3, 2026

Method and apparatus for document analysis through layer separation by machine learning with cross-layer reasoning

Inventors: Wai Kai Arvin Tang (Hong Kong, HK); Ping Yin Koon (Hong Kong, HK)
Assignee: Hong Kong Applied Science and Technology Research Institute Company Limited
G06V30/416G06V10/774G06V30/22G06V30/413G06V2201/09
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,567,276
App. No.
18/459,460
Filed
Sep 1, 2023
Granted
Mar 3, 2026
Kind
B2
Art Unit
2682
USPC
382/176
Abstract

A method for processing electronic documents, comprising: receiving an electronic document; recognizing one or more content components in the electronic document; identifying a content type for each of the recognized content components; creating, by a layer separator, one or more logical layers from the recognized content components such that each of the logical layer contains only the content components of the same content type; and invoking a content-type specific content handler for each of the logical layers created. The layer separator comprises a machine learning (ML) model based on a modified U-Net convolutional neural network and trained to classify the content types of the content components. The modified U-Net CNN is improved over traditional U-Net CNN with transformers at each layer to achieve high recovery rate.

Claims (79)

1 . A method for processing electronic documents, comprising:

receiving an electronic document;

recognizing one or more content components in the electronic document;

identifying a content type for each of the recognized content components; and

creating, by a layer separator, one or more logical layers from the recognized content components such that each of the logical layer contains only the content components of the same content type;

wherein the layer separator comprises a machine learning (ML) model based on a modified U-Net convolutional neural network and trained to classify the content types of the content components;

wherein the modified U-Net convolutional neural network comprises four layers of encoders-decoders;

wherein each encoder is configured to down-sample a feature map through convolution, activation, and pooling operations such that, during contraction, spatial information is reduced while feature information is increased,

wherein each decoder is configured to up-sample a feature map through up-convolution and activation operations such that, during expansion, spatial information is increased while feature information is reduced; and

wherein the ML model is trained using training data comprising a plurality of pairs of a generated document image and correspondingly labelled logical layers of content components of various content types that compose the generated document image, such that encoder and decoder operations using the four layers of encoders-decoders are learned by the ML model through the training, and that the feature map transformations performed during contraction and expansion correspond to labeled logical layers used in the training.

2 . The method of claim 1 , therein the content types comprise printed text content type, handwritten text content type, chop stamp content type, structured content type, barcode content type, and complex content type.

3 . The method of claim 1 , further comprising:

extracting, by a printed text content handler, a region of interest (ROI) containing printed text for each of the content components of printed content type for further processing and disregard empty background space in the logical layer of printed text content type;

depending on a language model chosen for printed text content handler, segmenting the ROI into one or more of sentences and characters;

feeding the ROI as-is or the one or more of sentences and characters to an Optical Character Recognition (OCR) engine for performing a text recognition for the content component of printed text content type;

extracting one or more attributes and a location on the electronic document page of the content component of printed text content type, wherein the attributes comprise at least typeface, font size, and color of the text, and author identification.

4 . The method of claim 1 , further comprising:

extracting, by a handwritten text content handler, a ROI containing handwritten text or signature for each of the content components of handwritten text content type for further recognition as handwritten text or signature, and disregarding empty background space in the logical layer of handwritten text content type;

if the ROI contains handwritten text, feeding the ROI as-is to an OCR engine for performing a text recognition for the content component of handwritten text content type; and

if the ROI contains signature, feeding the ROI as-is to a signature verification engine for performing a signature verification comprising comparing the signature to records of authentic signatures stored in a signature database.

5 . The method of claim 1 , further comprising:

localizing, by a chop stamp content handler, an outline shape of a chop stamp content component in the logical layer of chop stamp content type;

cropping an image region of the outline shape of the chop stamp content component to obtain an chop stamp image;

performing a text recognition of the chop stamp image to extract a text from the chop stamp image; and

comparing and verifying the chop stamp image and the extracted text with records of chop stamp images stored in a chop stamp database.

6 . The method of claim 1 , further comprising:

detecting, recognizing, and extracting, by a structured content handler, one or more structured content components in the logical layer of structured content type using structure and shape analysis; and

detecting a sub-type of each of the extracted structured content components, wherein the sub-types comprising a table, a list, an underlining, a highlighting, a box, and an artifact of a non-arbitrary shape.

7 . The method of claim 1 , further comprising:

detecting, recognizing, and extracting, by a barcode content handler, one or more barcode content components in the logical layer of barcode content type; and

decoding each of the extracted barcode content components into machine readable data.

8 . The method of claim 1 , further comprising:

detecting, recognizing, and extracting, by a complex content handler, one or more complex content components in the logical layer of complex content type;

detecting a sub-type of each of the extracted complex content components; and

invoking a context-sensitive content handling sub-module for each of the extracted complex content component according to its complex content sub-type.

9 . The method of claim 1 , further comprising:

performing, by a multi-layer cross-referencing handler, context-sensitive cross referencing of two or more content components extracted from logical layers of different content types, comprising:

analyzing locations, content types, sub-types, and attributes of the content components to be cross-referenced to determine the relationship between the content components; and

determining a context significance from the determined relationship.

10 . An apparatus for processing electronic documents, comprising:

a layer separator configured to:

receive an electronic document;

recognize one or more content components in the electronic document;

identify a content type for each of the recognized content components; and

create one or more logical layers from the recognized content components such that each of the logical layer contains only the content components of the same content type;

wherein the layer separator comprises a machine learning (ML) model based on a modified U-Net convolutional neural network and trained to classify the content types of the content components;

wherein the modified U-Net convolutional neural network comprises four layers of encoders-decoders;

wherein each encoder is configured to down-sample a feature map through convolution, activation, and pooling operations such that, during contraction, spatial information is reduced while feature information is increased,

wherein each decoder is configured to up-sample a feature map through up-convolution and activation operations such that, during expansion, spatial information is increased while feature information is reduced; and

wherein the ML model is trained using training data comprising a plurality of pairs of a generated document image and correspondingly labelled logical layers of content components of various content types that compose the generated document image, such that encoder and decoder operations using the four layers of encoders-decoders are learned by the ML model through the training, and that the feature map transformations performed during contraction and expansion correspond to labeled logical layers used in the training.

11 . The apparatus of claim 10 , therein the content types comprise printed text content type, handwritten text content type, chop stamp content type, structured content type, barcode content type, and complex content type.

12 . The apparatus of claim 10 , further comprising a printed text content handler configured to:

extract a region of interest (ROI) containing printed text for each of the content components of printed content type for further processing and disregard empty background space in the logical layer of printed text content type;

depending on a language model chosen for printed text content handler, segment the ROI into one or more of sentences and characters;

feed the ROI as-is or the one or more of sentences and characters to an Optical Character Recognition (OCR) engine for performing a text recognition for the content component of printed text content type;

extract one or more attributes and a location on the electronic document page of the content component of printed text content type, wherein the attributes comprise at least typeface, font size, and color of the text, and author identification.

13 . The apparatus of claim 10 , further comprising a handwritten text content handler configured to:

extract a ROI containing handwritten text or signature for each of the content components of handwritten text content type for further recognition as handwritten text or signature, and disregard empty background space in the logical layer of handwritten text content type;

if the ROI contains handwritten text, feed the ROI as-is to an OCR engine for performing a text recognition for the content component of handwritten text content type; and

if the ROI contains signature, feed the ROI as-is to a signature verification engine for performing a signature verification comprising comparing the signature to records of authentic signatures stored in a signature database.

14 . The apparatus of claim 10 , further comprising a chop stamp content handler configured to:

localize an outline shape of a chop stamp content component in the logical layer of chop stamp content type;

crop an image region of the outline shape of the chop stamp content component to obtain an chop stamp image;

perform a text recognition of the chop stamp image to extract a text from the chop stamp image; and

compare and verify the chop stamp image and the extracted text with records of chop stamp images stored in a chop stamp database.

15 . The apparatus of claim 10 , further comprising a structured content handler configured to:

detect, recognize, and extract one or more structured content components in the logical layer of structured content type using structure and shape analysis; and

detect, by the structure and shape analysis, a sub-type of each of the extracted structured content components, wherein the sub-types comprising a table, a list, an underlining, a highlighting, a box, and an artifact of a non-arbitrary shape.

16 . The apparatus of claim 10 , further comprising a barcode content handler configured to:

detect, recognize, and extract one or more barcode content components in the logical layer of barcode content type; and

decode each of the extracted barcode content components into machine readable data.

17 . The apparatus of claim 10 , further comprising a complex content handler configured to:

detect, recognize, and extract one or more complex content components in the logical layer of complex content type;

detect a sub-type of each of the extracted complex content components; and

invoke a context-sensitive content handling sub-module for each of the extracted complex content component according to its complex content sub-type.

18 . The apparatus of claim 10 , further comprising a multi-layer cross-referencing handler configured to:

perform context-sensitive cross referencing of two or more content components extracted from logical layers of different content types, comprising:

analyzing locations, content types, sub-types, and attributes of the content components to be cross-referenced to determine the relationship between the content components; and

determining a context significance from the determined relationship.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 3, 2023
From: TANG, WAI KAI ARVIN; KOON, PING YIN
To: HONG KONG APPLIED SCIENCE AND TECHNOLOGY RESEARCH INSTITUTE COMPANY LIMITED
Reel/Frame 064784/0899 →
Continuity (1)
Related Publication 20250078557A1 · Mar 6, 2025
References Cited (13)
US 6411733B1 · Saund · 2002 [cited by applicant]
US 7616813B2 · Nishida · 2009 [cited by applicant]
US 9384423B2 · Rodriguez-Serrano et al. · 2016 [cited by applicant]
US 10540579B2 · Reisswig et al. · 2020 [cited by applicant]
US 20060204095A1 · Nishida · 2006 [cited by applicant]
US 20090148039A1 · Chen et al. · 2009 [cited by applicant]
US 20110255789A1 · Neogi et al. · 2011 [cited by applicant]
US 20190147103A1 · Bhowan · 2019 [cited by examiner]
US 20210034855A1 · Shorter et al. · 2021 [cited by applicant]
US 20210117666A1 · Kaynig-Fittkau · 2021 [cited by examiner]
CN 114998905A · 2022 [cited by applicant]
WO 2023034125A1 · 2023 [cited by applicant]
International Search Report and Written Opinion of corresponding PCT application No. PCT/CN2023/124523 mailed on May 8, 2024. [cited by applicant]