IP Library › Granted Patent US 11,966,455
Granted Patent B2
US 11,966,455 · App. 17/332,478 · Granted Apr 23, 2024

Text partitioning method, text classifying method, apparatus, device and storage medium

Inventor: Bingqian Wang (Beijing, CN)
Assignee: BOE TECHNOLOGY GROUP CO., LTD.
G06F18/2163G06F18/214G06F18/217G06F18/24G06F18/254G06N3/04G06N3/08G06V10/44G06V30/158G06V30/274
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,966,455
App. No.
17/332,478
Granted
Apr 23, 2024
Kind
B2
Abstract

A text partitioning method, a text classifying method, an apparatus, a device and a storage medium, wherein the method includes: parsing a content image, to obtain a target text in a text format; according to a line break in the target text, partitioning the target text into a plurality of text sections; and according to a first data-volume threshold, partitioning sequentially the plurality of text sections into a plurality of text-to-be-predicted sets, wherein a data volume of a last one text section in each of the text-to-be-predicted sets is greater than a second data-volume threshold.

Claims (83)

1. A text partitioning method, wherein the method comprises:

parsing a content image, to obtain a target text in a text format;

according to a line break in the target text, partitioning the target text into a plurality of text sections; and

according to a first data-volume threshold, partitioning sequentially the plurality of text sections into a plurality of text-to-be-predicted sets, wherein a data volume of a last one text section in each of the text-to-be-predicted sets is greater than a second data-volume threshold wherein the plurality of text-to-be-predicted sets are used to be inputted into a semantic identification model to be used for semantic identification;

wherein the step of, according to the first data-volume threshold, partitioning sequentially the plurality of text sections into the plurality of text-to-be-predicted sets comprises:

creating an original text set;

reading through the plurality of text sections, adding the text sections that have been read through currently into the original text set, till a data volume of the original text set obtained after the addition is greater than the first data-volume threshold, and using the original text set obtained after the addition as a candidate text set;

on the condition that a data volume of a last one text section in the candidate text set is greater than the second data-volume threshold, using the candidate text set as a text-to-be-predicted set;

on the condition that a data volume of a last one text section in the candidate text set is less than or equal to the second data-volume threshold, taking out the last one text section from the candidate text set, to use the candidate text set obtained after the taking-out as a text-to-be-predicted set; and

on the condition that a remaining text section exists, partitioning sequentially the remaining text section into the text-to-be-predicted sets.

2. The method according to claim 1 , wherein the step of, according to the line break in the target text, partitioning the target text into the plurality of text sections comprises:

according to a blank character in the target text, partitioning the target text into a plurality of text lines; and

according to the line break in the target text, partitioning the plurality of text lines into the plurality of text sections.

3. The method according to claim 1 , wherein the step of parsing the content image, to obtain the target text in the text format comprises:

determining a text box in the content image;

determining a segment line in the text box;

according to the segment line, partitioning the text box; and

extracting the target text in the text format from the text box obtained after the partitioning.

4. The method according to claim 3 , wherein the step of determining the segment line in the text box comprises:

acquiring a coordinate value of the text box; and

using a vertical line where a modal number of a horizontal coordinate in the coordinate value is located as the segment line.

5. The method according to claim 3 , wherein the step of extracting the target text in the text format from the text box obtained after the partitioning comprises:

according to coordinate values of the text boxes obtained after the partitioning, acquiring weights of the text boxes obtained after the partitioning;

according to the weights, acquiring an extraction sequence of the text boxes obtained after the partitioning; and

according to the extraction sequence, extracting the target text in the text format from the text boxes obtained after the partitioning.

6. A text classifying method, wherein the method comprises:

by using the text partitioning method according to claim 1 , acquiring the text-to-be-predicted sets;

inputting the text-to-be-predicted sets into a target-text classifying model that has been pre-trained, wherein the target-text classifying model comprises at least: a multilayer label pointer network and a multi-label classifying network;

by using the multilayer label pointer network, acquiring a first classification result of the text-to-be-predicted sets, and by using the multi-label classifying network, acquiring a second classification result of the text-to-be-predicted sets; and

according to the first classification result and the second classification result, acquiring a target classification result of the text-to-be-predicted sets.

7. The method according to claim 6 , wherein the target-text classifying model is obtained by training by using the following steps:

marking the text-to-be-predicted sets with classification labels, to obtain sample text sets;

inputting the sample text sets into an original-text classifying model to be trained, and training, wherein the original-text classifying model comprises at least: a multilayer label pointer network and a multi-label classifying network;

by using the multilayer label pointer network, acquiring a third classification result of the sample text sets, and by using the multi-label classifying network, acquiring a fourth classification result of the sample text sets;

according to the third classification result, the fourth classification result and the classification labels, acquiring a loss value of the original-text classifying model obtained after the training; and

on the condition that the loss value is less than a loss-value threshold, using the original-text classifying model obtained after the training as the target-text classifying model.

8. The method according to claim 7 , wherein the step of inputting the sample text sets into the original-text classifying model to be trained, and training comprises:

inputting the sample text sets into a pre-training language model, to obtain a word embedded matrix and a position embedded matrix that correspond to the sample text sets;

combining the word embedded matrix and the position embedded matrix, to obtain an input embedded vector;

according to the input embedded vector, acquiring semantic vectors corresponding to the sample text sets; and

inputting the semantic vectors into the original-text classifying model to be trained, and training.

9. The method according to claim 8 , wherein the step of, by using the multilayer label pointer network, acquiring the third classification result of the sample text sets, and by using the multi-label classifying network, acquiring the fourth classification result of the sample text sets comprises:

by using the multilayer label pointer network, acquiring the semantic vectors of the sample text sets, and a starting-position probability value and an ending-position probability value that are related to each of the classification labels, as the third classification result; and

by using the multi-label classifying network, acquiring the semantic vectors of the sample text sets, and a classification probability value that is related to each of the classification labels, as the fourth classification result.

10. The method according to claim 9 , wherein the step of, according to the third classification result, the fourth classification result and the classification labels, acquiring the loss value of the original-text classifying model obtained after the training comprises:

combining individually the starting-position probability value and the ending-position probability value with the classification probability value, to obtain a target starting-position probability value and a target ending-position probability value of the semantic vectors that are related to each of the classification labels; and

according to the semantic vectors of the sample text sets, and the target starting-position probability value, a standard starting-position probability value, the target ending-position probability value and a standard ending-position probability value that are related to each of the classification labels, acquiring the loss value of the original-text classifying model obtained after the training.

11. An electronic device, wherein the electronic device comprises a memory, a processor and a computer program that is stored in the memory and is executable on the processor, and the processor, when executing the computer program, performs the following operations:

parsing a content image, to obtain a target text in a text format;

according to a line break in the target text, partitioning the target text into a plurality of text sections; and

according to a first data-volume threshold, partitioning sequentially the plurality of text sections into a plurality of text-to-be-predicted sets, wherein a data volume of a last one text section in each of the text-to-be-predicted sets is greater than a second data-volume threshold wherein the plurality of text-to-be-predicted sets are used to be inputted into a semantic identification model to be used for semantic identification;

wherein the operation performed by the processor of according to a first data-volume threshold, partitioning sequentially the plurality of text sections into a plurality of text-to-be-predicted sets comprises:

creating an original text set

reading through the plurality of text sections, adding the text sections that have been read through currently into the original text set, till a data volume of the original text set obtained after the addition is greater than the first data-volume threshold, and using the original text set obtained after the addition as a candidate text set;

on the condition that a data volume of a last one text section in the candidate text set is greater than the second data-volume threshold, using the candidate text set as a text-to-be-predicted set;

on the condition that a data volume of a last one text section in the candidate text set is less than or equal to the second data-volume threshold, taking out the last one text section from the candidate text set, to use the candidate text set obtained after the taking-out as a text-to-be-predicted set; and

on the condition that a remaining text section exists, partitioning sequentially the remaining text section into the text-to-be-predicted sets.

12. The electronic device according to claim 11 , wherein the operation performed by the processor of according to a line break in the target text, partitioning the target text into a plurality of text sections comprises:

according to a blank character in the target text, partitioning the target text into a plurality of text lines; and

according to the line break in the target text, partitioning the plurality of text lines into the plurality of text sections.

13. The electronic device according to claim 11 , wherein the processor further performs the following operations:

by using the operations of claim 11 , acquiring the text-to-be-predicted sets;

inputting the text-to-be-predicted sets into a target-text classifying model that has been pre-trained, wherein the target-text classifying model comprises at least: a multilayer label pointer network and a multi-label classifying network;

by using the multilayer label pointer network, acquiring a first classification result of the text-to-be-predicted sets, and by using the multi-label classifying network, acquiring a second classification result of the text-to-be-predicted sets; and

according to the first classification result and the second classification result, acquiring a target classification result of the text-to-be-predicted sets.

14. A computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program, when executed by the processor, performs the following operations:

parsing a content image, to obtain a target text in a text format;

according to a line break in the target text, partitioning the target text into a plurality of text sections; and

according to a first data-volume threshold, partitioning sequentially the plurality of text sections into a plurality of text-to-be-predicted sets, wherein a data volume of a last one text section in each of the text-to-be-predicted sets is greater than a second data-volume threshold wherein the plurality of text-to-be-predicted sets are used to be inputted into a semantic identification model to be used for semantic identification;

wherein the operation performed by the computer program of according to a first data-volume threshold, partitioning sequentially the plurality of text sections into a plurality of text-to-be-predicted sets comprises:

creating an original text set;

reading through the plurality of text sections, adding the text sections that have been read through currently into the original text set, till a data volume of the original text set obtained after the addition is greater than the first data-volume threshold, and using the original text set obtained after the addition as a candidate text set;

on the condition that a data volume of a last one text section in the candidate text set is greater than the second data-volume threshold, using the candidate text set as a text-to-be-predicted set;

on the condition that a data volume of a last one text section in the candidate text set is less than or equal to the second data-volume threshold, taking out the last one text section from the candidate text set, to use the candidate text set obtained after the taking-out as a text-to-be-predicted set; and

on the condition that a remaining text section exists, partitioning sequentially the remaining text section into the text-to-be-predicted sets.

15. The computer-readable storage medium according to claim 14 , wherein the operation performed by the computer program of according to a line break in the target text, partitioning the target text into a plurality of text sections comprises:

according to a blank character in the target text, partitioning the target text into a plurality of text lines; and

according to the line break in the target text, partitioning the plurality of text lines into the plurality of text sections.

16. The computer-readable storage medium according to claim 14 , wherein the computer program further performs the following operations:

by using the operations of claim 14 , acquiring the text-to-be-predicted sets;

inputting the text-to-be-predicted sets into a target-text classifying model that has been pre-trained, wherein the target-text classifying model comprises at least: a multilayer label pointer network and a multi-label classifying network;

by using the multilayer label pointer network, acquiring a first classification result of the text-to-be-predicted sets, and by using the multi-label classifying network, acquiring a second classification result of the text-to-be-predicted sets; and

according to the first classification result and the second classification result, acquiring a target classification result of the text-to-be-predicted sets.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 27, 2021
From: WANG, BINGQIAN
To: BOE TECHNOLOGY GROUP CO., LTD.
Reel/Frame 056375/0949 →
Priority Claims (1)
CN 202011053820.9 · Sep 29, 2020 · national
Continuity (1)
Related Publication 20220101060A1 · Mar 31, 2022