IP Library › Granted Patent US 12,579,834
Granted Patent B2
US 12,579,834 · App. 18/319,896 · Granted Mar 17, 2026

Information extraction method and apparatus for text with layout

Inventors: Minqin Chen (Shenzhen, CN); Peng Wu (Shenzhen, CN); Rongzhong Yue (Dongguan, CN); Xuan Jiang (Shenzhen, CN); Licui Hao (Shenzhen, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G06V30/414G06V30/18G06V30/19093G06V30/19173
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,834
App. No.
18/319,896
Granted
Mar 17, 2026
Kind
B2
Abstract

An information extraction method includes: determining that a text block that belongs to a target category and that is in text with layout is to be extracted; recognizing, based on feature information at a text block granularity, the text block that belongs to the target category and that is in the text with layout; and outputting an identifier of the text block that belongs to the target category and that is in the text with layout.

Claims (59)

1 . A method comprising:

training a binary classification model to obtain a trained binary classification model for classifying whether a text block belongs to a target category based on spatial relationships with adjacent text blocks in preset orientations, wherein the adjacent text blocks are text blocks that have a preset location relationship with a first to-be-recognized text block, and wherein the preset orientations include at least one of directly above, directly below, directly left, or directly right;

receiving a text with layout comprising a first text block, wherein the first text block belongs to a target category;

receiving a request message to extract the first text block from the text with layout;

classifying, using the trained binary classification model and based on feature information, the first text block to obtain a first classified text block, wherein the feature information comprises at least one of data information of the first text block, metadata information of the first text block, or spatial location information of the first text block; and

outputting a first identifier of the first classified text block.

2 . The method of claim 1 , further comprising recognizing, based on the feature information of a to-be-recognized text block in the text with layout, whether the to-be-recognized text block belongs to the target category.

3 . The method of claim 1 , further comprising:

recognizing, based on a first feature information of a target text block in the text with layout, whether a to-be-recognized text block in the text with layout belongs to the target category; or

recognizing, based on a second feature information of a to-be-recognized text block in the text with layout and the first feature information, whether the to-be-recognized text block belongs to the target category, wherein the target text block is a second text block that has a preset location relationship with the to-be-recognized text block.

4 . The method of claim 3 , wherein the second text block is in a preset range of the to-be-recognized text block or the second text block is adjacent to the to-be-recognized text block in a preset orientation of the to-be-recognized text block.

5 . The method of claim 1 , wherein the data information comprises at least one of the following:

a total length of a character string in the first text block;

whether the first text block comprises a preset character or a preset character string;

a total quantity of preset characters or preset character strings comprised in the first text block;

a proportion of preset characters or preset character strings comprised in the first text block to characters in the first text block;

whether the first text block comprises a preset keyword;

whether the first text block comprises a preset named entity; or

whether the first text block comprises preset layout information.

6 . The method of claim 1 , wherein the metadata information comprises at least one of the following: a font of the first text block, a font size of the first text block, a color of the first text block, whether the first text block is bold, whether the first text block is italicized, or whether the first text block is underlined.

7 . The method of claim 1 , wherein the spatial location information comprises at least one of the following:

a distance of the first text block relative to a page edge of the text with layout; or

a distance of the first text block relative to a reference text block in the text with layout.

8 . The method of claim 1 , further comprising displaying a first user interface, wherein the first user interface comprises first indication information and second indication information, the first indication information indicates to a user to enter the text with layout, and the second indication information indicates to the user to enter a second identifier of the target category.

9 . The method of claim 8 , further comprising displaying a second user interface, wherein the second user interface comprises third indication information indicating that an information extraction process is being performed.

10 . The method of claim 9 , wherein outputting the first identifier comprises displaying a third user interface comprising the first identifier.

11 . The method of claim 1 , further comprising inputting the feature information into the binary classification model to obtain an output result, wherein the binary classification model represents whether a text block belongs to the target category.

12 . The method of claim 11 , further comprising:

obtaining N features of the target category, wherein the N features are features represented by the feature information, and wherein N is an integer greater than or equal to 1;

obtaining a training set comprising a plurality of text blocks belonging to the target category;

performing feature extraction based on the N features on each of the plurality of text blocks to obtain a feature combination corresponding to the target category; and

performing the training based on a plurality of feature combinations obtained for the plurality of text blocks to obtain the binary classification model.

13 . The method of claim 12 , further comprising displaying a fourth user interface, wherein the fourth user interface comprises fourth indication information and fifth indication information, wherein the fourth indication information indicates to a user to enter a second identifier of the target category and the N features, and wherein the fifth indication information indicates the user to enter the training set.

14 . The method of claim 12 , wherein performing the training comprises displaying a fifth user interface comprising sixth indication information indicating that the binary classification model is training.

15 . The method of claim 11 , further comprising receiving the binary classification model from a network device.

16 . An information extraction apparatus comprising:

a memory configured to store a computer program; and

one or more processors coupled to the memory and configured to execute the computer program to:

train a binary classification model to obtain a trained binary classification model for classifying whether a first text block belongs to a target category based on spatial relationships with adjacent text blocks in preset orientations, wherein the adjacent text blocks are text blocks that have a preset location relationship with a first to-be-recognized text block, and wherein the preset orientations include at least one of directly above, directly below, directly left, or directly right;

receive a text with layout comprising a text block, wherein the text block belongs to a target category;

receive an instruction to extract the text block from the text with layout;

classify, using the trained binary classification model and based on feature information, the text block to obtain a classified text block, wherein the feature information comprises at least one of: data information of the text block, metadata information of the text block, or spatial location information of the text block; and

output an identifier of the classified text block.

17 . The information extraction apparatus of claim 16 , wherein the one or more processors are configured to execute the computer program to recognize, based on the feature information of a to-be-recognized text block in the text with layout, whether the to-be-recognized text block belongs to the target category.

18 . The information extraction apparatus of claim 16 , wherein the data information comprises at least one of the following:

a total length of a character string in the first text block;

whether the first text block comprises a preset character or a preset character string;

a total quantity of preset characters or preset character strings comprised in the first text block;

a proportion of preset characters or preset character strings comprised in the first text block to characters in the first text block;

whether the first text block comprises a preset keyword;

whether the first text block comprises a preset named entity; or

whether the first text block comprises preset layout information.

19 . The information extraction apparatus of claim 16 , wherein the metadata information comprises at least one a font of the first text block, a font size of the first text block, a color of the first text block, whether the first text block is bold, whether the first text block is italicized, or whether the first text block is underlined, and wherein the spatial location information comprises at least one of a distance of the first text block relative to a page edge of the text with layout, or a distance of the first text block relative to a reference text block in the text with layout.

20 . A non-volatile computer-readable storage medium storing a computer program, wherein when the computer program is executed on a computer, the computer is enabled to:

train a binary classification model to obtain a trained binary classification model for classifying whether a first text block belongs to a target category based on spatial relationships with adjacent text blocks in preset orientations, wherein the adjacent text blocks are text blocks that have a preset location relationship with a first to-be-recognized text block, and wherein the preset orientations include at least one of directly above, directly below, directly left, or directly right;

receive a text with layout comprising a text block, wherein the text block belongs to a target category;

receive an instruction to extract the text block from the text with layout;

classify, using the trained binary classification model and based on feature information at, the text block to obtain a classified text block, wherein the feature information comprises at least one of: data information of the text block, metadata information of the text block, or spatial location information of the text block; and

output an identifier of the classified text block.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2025
From: YUE, RONGZHONG; JIANG, XUAN; HAO, LICUI
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 073271/0622 →
EMPLOYMENT AGREEMENT Recorded Dec 19, 2025
From: CHEN, MINQIN
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 073818/0292 →
EMPLOYMENT AGREEMENT Recorded Dec 19, 2025
From: WU, PENG
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 073818/0300 →
Priority Claims (1)
CN 202011308474.4 · Nov 19, 2020 · national
Continuity (2)
Continuation PCTCN2021103501 · Jun 30, 2021
Related Publication 20230290169A1 · Sep 14, 2023
References Cited (12)
US 10997195B1 · Sekar · 2021 [cited by examiner]
US 20170293951A1 · Nolan · 2017 [cited by examiner]
US 20200134388A1 · Rohde · 2020 [cited by examiner]
US 20200250459A1 · Sarshogh · 2020 [cited by examiner]
US 20210110527A1 · Wheaton · 2021 [cited by examiner]
US 20230290169A1 · Chen et al. · 2023 [cited by applicant]
CN 110414395A · 2019 [cited by examiner]
CN 11259631A · 2020 [cited by applicant]
CN 111753538A · 2020 [cited by applicant]
CN 112487138A · 2021 [cited by applicant]
EP 3267332A1 · 2018 [cited by applicant]
EP 3349124A1 · 2018 [cited by applicant]