IP Library Granted Patent US 11,152,007
Granted Patent B2
US 11,152,007 · App. 16/543,155 · Granted Oct 19, 2021

Method, and device for matching speech with text, and computer-readable storage medium

Inventor: Yongshuai Lu (Beijing, CN)
Assignee: Baidu Online Network Technology Co., Ltd.
G10L17/14G10L15/187G10L17/02G06F16/3343G06F40/10G06F40/194G06K9/6215
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,152,007
App. No.
16/543,155
Granted
Oct 19, 2021
Kind
B2
Abstract

Embodiments of a method and device for matching a speech with a text, and a computer-readable storage medium are provided. The method can include: acquiring a speech identification text by identifying a received speech signal; comparing the speech identification text with multiple candidate texts in a first matching mode to determine a first matching text; and comparing phonetic symbols of the speech identification text with phonetic symbols of the multiple candidate texts in a second matching mode to determine a second matching text, in a case that no first matching text is determined.

Claims (74)

1. A method for matching a speech with a text, comprising:

acquiring a speech identification text by identifying a received speech signal;

comparing the speech identification text with multiple candidate texts in a first matching mode to determine a first matching text; and

comparing phonetic symbols of the speech identification text with phonetic symbols of the multiple candidate texts in a second matching mode to determine a second matching text, in response to not determining the first matching text,

wherein comparing phonetic symbols of the speech identification text with phonetic symbols of the multiple candidate texts in the second matching mode to determine the second matching text comprises:

converting the speech identification text into the phonetic symbols of the speech identification text and converting the multiple candidate texts into the phonetic symbols of the multiple candidate texts;

calculating a similarity between the phonetic symbols of the speech identification text and the phonetic symbols of each of the multiple candidate texts; and

determining a candidate text with a largest similarity as a matched candidate text in response to determining that the largest similarity is larger than a set threshold; and

outputting the matched candidate text,

wherein calculating the similarity between the phonetic symbols of the speech identification text and the phonetic symbols of each of the multiple candidate texts is by the following formula:

similarity

=

LCS

(

s

,

q

)

len

(

s

)

wherein s represents phonetic symbols of one of the multiple candidate texts, q represents the phonetic symbols of the speech identification text, LCS(s, q) represents a length of a longest common sequence between the phonetic symbols of the one of the multiple candidate texts and the phonetic symbols of the speech identification text, len(s) represents a length of the phonetic symbols of the one of the multiple candidate texts.

2. The method according to claim 1 , further comprising:

outputting the first matching text as a matched candidate text, in response to determining the first matching text; and

outputting the second matching text as the matched candidate text, in response to determining the second matching text.

3. The method according to claim 1 , further comprising:

calculating a similarity between a sentence vector of the speech identification text and a sentence vector of each of the multiple candidate texts, in response to not determining the second matching text; and outputting a candidate text with a largest similarity as a matched candidate text.

4. The method according to claim 3 , wherein the calculating a similarity between a sentence vector of the speech identification text and a sentence vector of each of the multiple candidate texts comprises:

segmenting the speech identification text and the multiple candidate texts into words;

acquiring a word vector of each word;

adding word vectors of words of the speech identification text to obtain the sentence vector of the speech identification text, and adding word vectors of words of one of the multiple candidate texts to acquire a sentence vector of the one of the multiple candidate texts; and

calculating a cosine similarity between the sentence vector of the speech identification text and the sentence vector of the one of the multiple candidate texts, as the similarity between the sentence vector of the speech identification text and the sentence vector of the one of the multiple candidate texts.

5. A device for matching a speech with a text, comprising:

one or more processors; and

a storage device configured to store one or more programs, that, when executed by the one or more processors, cause the one or more processors to:

acquire a speech identification text by identifying a received speech signal;

compare the speech identification text with multiple candidate texts in a first matching mode to determine a first matching text; and

compare phonetic symbols of the speech identification text with phonetic symbols of the multiple candidate texts in a second matching mode to determine a second matching text, in response to not determining the first matching text,

wherein the one or more programs, when executed by the one or more processors, cause the one or more processors further to:

convert the speech identification text into the phonetic symbols of the speech identification text and convert the multiple candidate texts into the phonetic symbols of the multiple candidate texts;

calculate a similarity between the phonetic symbols of the speech identification text and the phonetic symbols of each of the multiple candidate texts;

determine a candidate text with a largest similarity as a matched candidate text in response to determining that the largest similarity is larger than a set threshold; and

output the matched candidate text,

wherein the similarity between the phonetic symbols of the speech identification text and the phonetic symbols of each of the multiple candidate texts is calculated by the following formula:

similarity

=

LCS

(

s

,

q

)

len

(

s

)

wherein s represents phonetic symbols of one of the multiple candidate texts, q represents the phonetic symbols of the speech identification text, LCS(s, q) represents a length of a longest common sequence between the phonetic symbols of one of the multiple candidate texts and the phonetic symbols of the speech identification text, len(s) represents a length of the phonetic symbols of the one of the multiple candidate texts.

6. The device according to claim 5 , wherein the one or more programs, when executed by the one or more processors, cause the one or more processors further to:

output the first matching text as a matched candidate text, in response to determining the first matching text; and

output the second matching text as the matched candidate text, in response to determining the second matching text.

7. The device according to claim 5 , wherein the one or more programs, when executed by the one or more processors, cause the one or more processors further to:

calculate a similarity between a sentence vector of the speech identification text and a sentence vector of each of the multiple candidate texts, in response to not determining the second matching text; and

output a candidate text with a largest similarity as a matched candidate text.

8. The device according to claim 7 , wherein the one or more programs, when executed by the one or more processors, cause the one or more processors further to:

segment the speech identification text and the multiple candidate texts into words;

acquire a word vector of each word;

add word vectors of words of the speech identification text to obtain the sentence vector of the speech identification text, and add the word vectors of words of one of the multiple candidate texts to acquire a sentence vector of the one of the multiple candidate texts; and

calculate a cosine similarity between the sentence vector of the speech identification text and the sentence vector of the one of the multiple candidate texts, as the similarity between the sentence vector of the speech identification text and the sentence vector of the one of the multiple candidate texts.

9. A non-transitory computer-readable storage medium, in which a computer program is stored, wherein the computer program, when executed by a processor, causes the processor to implement the method of claim 1 .

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2021
From: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.; SHANGHAI XIAODU TECHNOLOGY CO. LTD.
Reel/Frame 056811/0772 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 19, 2019
From: LU, YONGSHUAI
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 050088/0318 →
Priority Claims (1)
CN 201811495921.4 · Dec 7, 2018 · national
Continuity (1)
Related Publication 20200184978A1 · Jun 11, 2020