IP Library › Granted Patent US 12,437,748
Granted Patent B2
US 12,437,748 · App. 17/965,869 · Granted Oct 7, 2025

Spoken language processing method and apparatus, and storage medium

Inventor: Yinhui Zhang (Beijing, CN)
Assignee: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
G10L15/02G06F17/16G06F18/2415G06F18/253G06F40/166G06F40/20G06F40/279G06F40/30G06N3/045G06N3/08G10L15/183
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,748
App. No.
17/965,869
Granted
Oct 7, 2025
Kind
B2
Abstract

Spoken language processing method and apparatus, a device, and a storage medium, which relate to the field of artificial intelligence and, in particular, to the fields of deep learning, natural-language understanding, intelligent customer service, and the like. The specific implementation solution includes: determining a word feature of a word in spoken text information; determining a correlation feature of the word in the spoken text information; and determining an effect of the word on fluency of the spoken text information according to the word feature of the word and the correlation feature of the word.

Claims (62)

1. A spoken language processing method, comprising:

obtaining spoken text information by collecting speech information and recognizing the speech information through automatic speech recognition;

determining a word feature of a word in the spoken text information;

determining a correlation feature of the word in the spoken text information;

determining an effect of the word on fluency of the spoken text information according to the word feature of the word and the correlation feature of the word; and

correcting the spoken text information according to the effect of the word on the fluency of the spoken text information to obtain target text information so as to improve quality of the spoken text information;

wherein determining the correlation feature of the word in the spoken text information comprises:

determining a word vector matrix of the word in the spoken text information; and

determining the correlation feature of the word according to the word vector matrix of the word; and

wherein determining the correlation feature of the word according to the word vector matrix of the word comprises:

determining a Hadamard product of a word vector matrix of an i-th word and a word vector matrix of a i-th word in the spoken text information;

determining a correlation matrix according to the Hadamard product; and

performing three-dimensional (3D) feature extraction on the correlation matrix to obtain correlation features of the i-th word and the j-th word, and

wherein i and j are natural numbers.

2. The method of claim 1 , wherein determining the correlation matrix according to the Hadamard product comprises:

determining relative position encoding of the i-th word and the j-th word in the spoken text information; and

adding the relative position encoding to the Hadamard product to obtain the correlation matrix of the i-th word and the j-th word.

3. The method of claim 1 , wherein determining the word vector matrix of the word in the spoken text information comprises:

processing the spoken text information by a language representation model to obtain the word vector matrix of the word in the spoken text information;

determining the word feature of the word in the spoken text information comprises:

performing two-dimensional (2D) feature extraction on the word vector matrix of the word to obtain the word feature of the word.

4. A spoken language processing apparatus, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores an instruction executable by the at least one processor to enable the at least one processor to perform the following steps:

obtaining spoken text information by collecting speech information and recognizing the speech information through automatic speech recognition;

determining a word feature of a word in the spoken text information;

determining a correlation feature of the word in the spoken text information;

determining an effect of the word on fluency of the spoken text information according to the word feature of the word and the correlation feature of the word; and

correcting the spoken text information according to the effect of the word on the fluency of the spoken text information to obtain target text information so as to improve quality of the spoken text information;

wherein determining the correlation feature of the word in the spoken text information comprises:

determining a word vector matrix of the word in the spoken text information; and

determining the correlation feature of the word according to the word vector matrix of the word; and

wherein determining the correlation feature of the word according to the word vector matrix of the word comprises:

determining a Hadamard product of a word vector matrix of an i-th word and a word vector matrix of a j-th word in the spoken text information;

determining a correlation matrix according to the Hadamard product; and

performing three-dimensional (3D) feature extraction on the correlation matrix to obtain correlation features of the i-th word and the j-th word, and

wherein i and i are natural numbers.

5. The apparatus of claim 4 , wherein determining the correlation matrix according to the Hadamard product comprises:

determining relative position encoding of the i-th word and the j-th word in the spoken text information; and

adding the relative position encoding to the Hadamard product to obtain the correlation matrix of the word.

6. The apparatus of claim 4 , wherein determining the word vector matrix of the word in the spoken text information comprises:

processing the spoken text information by a language representation model to obtain the word vector matrix of the word in the spoken text information; and

determining the word feature of the word in the spoken text information comprises performing two-dimensional (2D) feature extraction on the word vector matrix of the word to obtain the word feature of the word.

7. A non-transitory computer-readable storage medium, which is configured to store a computer instruction for causing a computer to perform the following steps:

obtaining spoken text information by collecting speech information and recognizing the speech information through automatic speech recognition;

determining a word feature of a word in the spoken text information;

determining a correlation feature of the word in the spoken text information;

determining an effect of the word on fluency of the spoken text information according to the word feature of the word and the correlation feature of the word; and

correcting the spoken text information according to the effect of the word on the fluency of the spoken text information to obtain target text information so as to improve quality of the spoken text information;

wherein determining the correlation feature of the word in the spoken text information comprises:

determining a word vector matrix of the word in the spoken text information; and

determining the correlation feature of the word according to the word vector matrix of the word; and

wherein determining the correlation feature of the word according to the word vector matrix of the word comprises:

determining a Hadamard product of a word vector matrix of an i-th word and a word vector matrix of a j-th word in the spoken text information;

determining a correlation matrix according to the Hadamard product; and

performing three-dimensional (3D) feature extraction on the correlation matrix to obtain correlation features of the i-th word and the j-th word;

wherein i and i are natural numbers.

8. The storage medium of claim 7 , wherein determining the correlation matrix according to the Hadamard product comprises:

determining relative position encoding of the i-th word and the j-th word in the spoken text information; and

adding the relative position encoding to the Hadamard product to obtain the correlation matrix of the i-th word and the j-th word.

9. The storage medium of claim 7 , wherein determining the word vector matrix of the word in the spoken text information comprises:

processing the spoken text information by a language representation model to obtain the word vector matrix of the word in the spoken text information;

determining the word feature of the word in the spoken text information comprises:

performing two-dimensional (2D) feature extraction on the word vector matrix of the word to obtain the word feature of the word.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 14, 2022
From: ZHANG, YINHUI
To: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
Reel/Frame 061422/0738 →
Priority Claims (1)
CN 202210323762.X · Mar 29, 2022 · national
Continuity (1)
Related Publication 20230317058A1 · Oct 5, 2023
References Cited (48)
US 11024194B1 · Beigman Klebanov · 2021 [cited by examiner]
US 20200401938A1 · Etkin · 2020 [cited by examiner]
US 20210216722A1 · Dai · 2021 [cited by examiner]
US 20210374338A1 · Shrivastava · 2021 [cited by examiner]
US 20210383069A1 · Liu · 2021 [cited by examiner]
US 20230317058A1 · Zhang · 2023 [cited by examiner]
CA 2945632C · 2021 [cited by examiner]
CN 102262632A · 2011 [cited by examiner]
CN 104347071A · 2015 [cited by examiner]
CN 106502979A · 2017 [cited by examiner]
CN 109493968A · 2019 [cited by examiner]
CN 110427625A · 2019 [cited by examiner]
CN 104347071B · 2020 [cited by applicant]
CN 110959159A · 2020 [cited by examiner]
CN 111291549A · 2020 [cited by examiner]
CN 111581968A · 2020 [cited by examiner]
CN 111667816A · 2020 [cited by examiner]
CN 111737995A · 2020 [cited by examiner]
CN 112016313A · 2020 [cited by examiner]
CN 112115721A · 2020 [cited by examiner]
CN 112149418A · 2020 [cited by examiner]
CN 112380845A · 2021 [cited by examiner]
CN 112417117A · 2021 [cited by examiner]
CN 112559688A · 2021 [cited by examiner]
CN 112735396A · 2021 [cited by examiner]
CN 112818694A · 2021 [cited by examiner]
CN 113160820A · 2021 [cited by examiner]
CN 113204619A · 2021 [cited by examiner]
CN 113704430A · 2021 [cited by examiner]
CN 113903048A · 2022 [cited by examiner]
CN 114005452A · 2022 [cited by examiner]
CN 114021582A · 2022 [cited by examiner]
CN 114141236A · 2022 [cited by examiner]
CN 114566147A · 2022 [cited by examiner]
CN 115623279A · 2023 [cited by examiner]
CN 114970666B · 2023 [cited by examiner]
TW 202009890A · 2020 [cited by examiner]
WO WO2022022421A1 · 2022 [cited by examiner]
WO WO2022156115A1 · 2022 [cited by examiner]
Supplemental Search Report issued on Jul. 19, 2023 by the CIPO in the corresponding Patent Application No. 202210323762X, with English translation. [cited by applicant]
European Search Report issued on Aug. 4, 2023 in the corresponding Patent Application No. 22201744.4-1203. [cited by applicant]
Bamdev, et al.: “Automated Speech Scoring System Under The Lens Evaluating and interpreting the linguistic cues for language proficiency,” Int'l J. of Artifical Intelligence in Education, arXiv:2111.15156v1 [cs.CL], (20… [cited by applicant]
Office Action issued on Jun. 15, 2023 by the CIPO in the corresponding Patent Application No. 202210323762.X, with English translation. [cited by applicant]
Search Report issued on Jun. 13, 2023 by the CIPO in the corresponding Patent Application No. 202210323762.X, with English translation. [cited by applicant]
Aida, et al.: “A Comprehensive Analysis of PMI-based Models for Measuring Semantic Differences,” Proceedings of 35th Pacific Asia Conf. on Language, Information, and Computation, (2021), pp. 1-11. [cited by applicant]
Jiawei Huang: “Research on Intent Classification in Dialogue Systems,” Thesis, School of Computer Central China Normal Univ., (2018), pp. 1-70, with English Abstract. [cited by applicant]
Linxiu Zhang: “Research on Names Entity Recognition Method for Weibo Text,” Thesis, Database Information Technology Series, (2019), pp. 1-68, with Engish Abstract. [cited by applicant]
Xu, et al.: “Intention Detection in Spoken Language Based in Context Information,” Computer Science, (2019), pp. 1-7; http://www.jsjkx.com, with English Abstract. [cited by applicant]