IP Library Granted Patent US 12,254,274
Granted Patent B2
US 12,254,274 · App. 17/802,349 · Granted Mar 18, 2025

Text classification method and text classification device

Inventors: Feng Liu (Beijing, CN); Mengmeng Liu (Beijing, CN); Zhiming Zhang (Beijing, CN); Nan Liang (Beijing, CN); Jiawei Xu (Beijing, CN); Zhentao Liu (Beijing, CN)
Assignee: Telefonaktiebolaget LM Ericsson (publ)
G06F40/30G06V30/153G06V30/18G06V30/19093G06V30/19107G06V30/19173G06V30/274
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,254,274
App. No.
17/802,349
Granted
Mar 18, 2025
Kind
B2
Abstract

Disclosed is a text classification method and a text classification device. The text classification method includes: receiving text data (S 1 ), the text data comprising one or more text semantic units; replacing the text semantic unit with a corresponding text keyword (S 2 ), based on a correspondence between text semantic elements and text keywords; extracting, with a semantic model, a semantic feature of the text keyword (S 3 ); and classifying, with a classification model, the text keyword at least based on the semantic feature, as a classification result of the text data (S 4 ).

Claims (46)

1. A computer-implemented text classification method comprising:

receiving text data comprising a test log from automatic testing of hardware and/or software, wherein the text data includes one or more text semantic units;

replacing the one or more text semantic units with corresponding one or more text keywords, based on a correspondence between text semantic units and text keywords; and

for each of the one or more text keywords:

extracting a semantic feature of the text keyword using a semantic model; and

classifying the text keyword based on the extracted semantic feature and a classification model, thereby obtaining a classification result of the text data, wherein the classification result indicates or identifies one or more of the following:

whether the automatic testing passed or failed, and

when the automatic testing failed, one or more errors in the test log.

2. The text classification method of claim 1 , wherein classifying the text keyword is further based on a temporal feature associated with the text semantic unit.

3. The text classification method of claim 1 , wherein each semantic feature, on which classifying the text keyword is based, has one or more of the following characteristics specified by a feature standardizer: value, and format.

4. The text classification method of claim 1 , further comprising training the classification model by performing at least the following operations:

receiving text data for training, the text data for training comprising one or more test logs from automatic testing of hardware and/or software, wherein the text data for training includes a plurality of text semantic units and associated category labels, wherein the category labels indicate or identify one or more of the following:

whether the automatic testing passed or failed, and

when the automatic testing failed, one or more errors in the test log;

replacing the plurality of text semantic units with a corresponding plurality of text keywords, based on the correspondence between text semantic units and text keywords;

extracting semantic features of the plurality of text keywords using the semantic model; and

training the classification model based on the extracted semantic features and the category labels associated with the text semantic units.

5. The text classification method of claim 4 , wherein training the classification model is further based on a temporal feature associated with the text data for training.

6. The text classification method of claim 4 , wherein each semantic feature, on which training the classification model is based, has one or more of the following characteristics specified by a feature standardizer: value, and format.

7. The text classification method of claim 1 , further comprising obtaining the correspondence between text semantic units and text keywords by performing at least the following operations:

receiving text data for training, the text data for training comprising one or more test logs from automatic testing of hardware and/or software, wherein the text data for training includes a plurality of text semantic units;

clustering the plurality of text semantic units to obtain a plurality of clusters, wherein each cluster includes text semantic units with a same or similar semantic; and

for each cluster, based on identifying a common word in the text semantic units included in the cluster, assigning the identified common word as a text keyword corresponding to the text semantic units included in the cluster.

8. The text classification method of claim 7 , wherein obtaining the correspondence between text semantic units and text keywords further comprises, for each cluster, based on identifying no common word in the text semantic units included in the cluster, performing the following operations:

dividing the cluster into a plurality of sub-clusters, until a common word is identified in the text semantic units included in a sub-cluster; and

assigning the identified common word as a text keyword corresponding to the text semantic units included in the sub-cluster.

9. The text classification method of claim 7 , wherein clustering the plurality of text semantic units comprises:

converting each text semantic unit into a corresponding eigenvector;

for each eigenvector corresponding to a text semantic unit:

performing a principal component analysis on the eigenvector to obtain a predetermined number of eigenvectors having largest eigenvalues, and

subtracting projection components of the eigenvector on the predetermined number of eigenvectors to obtain semantic eigenvectors; and

clustering the semantic eigenvectors, obtained for the plurality of text semantic units, in a feature space to obtain the plurality of clusters.

10. The text classification method of claim 9 , wherein converting each text semantic unit into an eigenvector comprises:

for each word comprising the text semantic unit,

converting the word into a corresponding word eigenvector, and

calculating an Inverse Document Frequency (IDF) for the word as a word weight of the word;

normalizing the word weights of the words comprising the text semantic unit; and

obtaining the eigenvector corresponding to the text semantic unit based on weighted summing the word eigenvectors for the text semantic unit based on the normalized word weights.

11. The text classification method of claim 1 , further comprising performing data cleaning on the text data to remove a word irrelevant to a semantic of the text data.

12. The text classification method of claim 1 , wherein:

the one or more text semantic units comprise respective lines of the test log, and

each line of the test log includes one or more complete sentences.

13. The text classification method of claim 11 , further comprising performing word segmentation on Chinese text data to obtain the text data on which the data cleaning is performed.

14. A text classification device comprising:

a memory having instructions stored thereon; and

a processor operably coupled to the memory, wherein execution of the instructions by the processor cause the text classification device to perform operations corresponding to the method of claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 25, 2022
From: LIU, FENG; LIU, MENGMENG; ZHANG, ZHIMING; LIANG, NAN; XU, JIAWEI; LIU, ZHENTAO
To: TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
Reel/Frame 060900/0928 →
Priority Claims (1)
CN 202010221049.5 · Mar 25, 2020 · national
Continuity (1)
Related Publication 20230139663A1 · May 4, 2023
References Cited (17)
US 10540610B1 · Yang et al. · 2020 [cited by applicant]
US 11281860B2 · Yue · 2022 [cited by examiner]
US 11580144B2 · Galitsky · 2023 [cited by examiner]
US 12159251B2 · Ekmekci · 2024 [cited by examiner]
US 20170220678A1 · Wren · 2017 [cited by examiner]
US 20200125574A1 · Ghoshal · 2020 [cited by examiner]
US 20200327008A1 · Singh · 2020 [cited by examiner]
US 20200371754A1 · P K · 2020 [cited by examiner]
CN 108170773A · 2018 [cited by applicant]
CN 108427720A · 2018 [cited by applicant]
CN 108536870A · 2018 [cited by applicant]
CN 109840157A · 2019 [cited by applicant]
CN 110232128A · 2019 [cited by applicant]
CN 110532354A · 2019 [cited by applicant]
CN 109947947B · 2021 [cited by applicant]
JP 2008234519A · 2008 [cited by applicant]
WO 2019194343A1 · 2019 [cited by applicant]