IP Library › Granted Patent US 10,847,146
Granted Patent B2
US 10,847,146 · App. 16/201,722 · Granted Nov 24, 2020

Multiple voice recognition model switching method and apparatus, and storage medium

Inventors: Bing Jiang (Beijing, CN); Xiangang Li (Beijing, CN); Ke Ding (Beijing, CN)
Assignee: Baidu Online Network Technology (Beijing) Co., Ltd.
G10L15/197G06F40/20G10L15/005G10L15/02G10L15/22G10L25/78
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,847,146
App. No.
16/201,722
Granted
Nov 24, 2020
Kind
B2
Abstract

Embodiments of the present disclosure disclose a method and apparatus for switching multiple speech recognition models. The method includes: acquiring at least one piece of speech information in user input speech; recognizing the speech information and matching a linguistic category for the speech information to determine a corresponding target linguistic category based on a matching degree; and switching a currently used speech recognition model to a speech recognition model corresponding to the target linguistic category. The embodiments of the present disclosure determine the corresponding target linguistic category based on the matching degree by recognizing the speech information and matching the linguistic category for the speech information, and switch the currently used speech recognition model to the speech recognition model corresponding to the target linguistic category.

Claims (45)

1. A method for switching multiple speech recognition models, the method comprising:

acquiring at least one piece of speech information in user input speech;

recognizing the speech information and matching a linguistic category for the speech information to determine a corresponding target linguistic category based on a matching degree; and

switching a currently used speech recognition model to a speech recognition model corresponding to the target linguistic category,

wherein the recognizing the speech information and matching a linguistic category for the speech information to determine a corresponding target linguistic category based on a matching degree comprises:

recognizing at least two pieces of speech sentences included in the speech information to obtain a matching degree between each speech sentence and a linguistic category; and

determining an initial linguistic category based on the matching degree, and calculating a product of probabilities of the speech sentences not belonging to the initial linguistic category, and determining the corresponding target linguistic category based on the product.

2. The method according to claim 1 , wherein the recognizing the speech information and matching a linguistic category for the speech information to determine a corresponding target linguistic category based on a matching degree comprises:

recognizing the speech information based on features of at least two linguistic categories to obtain a similarity between the speech information and each of the linguistic categories, and defining the similarity as the matching degree of the linguistic category.

3. The method according to claim 1 , wherein before the recognizing the speech information and matching a linguistic category for the speech information to determine a corresponding target linguistic category based on a matching degree, the method further comprises:

performing any one of following preprocessing on the speech information: a speech feature extraction, an effective speech detection, a speech vector representation, and a model scoring test.

4. The method according to claim 1 , further comprising:

recognizing the speech information, and displaying a prompt message to prompt a user to perform manual switching if a recognition result does not meet a preset condition.

5. The method according to claim 1 , wherein the recognizing the speech information and matching a linguistic category for the speech information to determine a corresponding target linguistic category based on a matching degree comprises:

recognizing the speech information and matching a linguistic category for the speech information;

determining at least two candidate linguistic categories having matching degrees meeting a preset condition;

querying a user historical speech recognition record to determine a linguistic category used by the user historically; and

selecting a linguistic category consistent with the linguistic category used by the user historically from the at least two candidate linguistic categories as the target linguistic category.

6. An apparatus for switching multiple speech recognition models, the apparatus comprising:

at least one processor; and

a memory storing instructions, the instructions when executed by the at least one processor, cause the at least one processor to perform operations, the operations comprising:

acquiring at least one piece of speech information in user input speech;

recognizing the speech information and matching a linguistic category for the speech information to determine a corresponding target linguistic category based on a matching degree; and

switching, a currently used speech recognition model to a speech recognition model corresponding to the target linguistic category,

wherein the recognizing the speech information and matching a linguistic category for the speech information to determine a corresponding target linguistic category based on a matching degree comprises:

recognizing at least two pieces of speech sentences included in the speech information to obtain a matching degree between each speech sentence and a linguistic category; and

determining an initial linguistic category based on the matching degree, and calculating a product of probabilities of the speech sentences not belonging to the initial linguistic category, and determining the corresponding target linguistic category based on the product.

7. The apparatus according to claim 6 , wherein the recognizing the speech information and matching a linguistic category for the speech information to determine a corresponding target linguistic category based on a matching degree comprises:

recognizing the speech information based on features of at least two linguistic categories to obtain a similarity between the speech information and each of the linguistic categories, and defining the similarity as the matching degree of the linguistic category.

8. The apparatus according to claim 6 , the operations further comprising:

performing any one of following preprocessing on the speech information: a speech feature extraction, an effective speech detection, a speech vector representation, and a model scoring test.

9. The apparatus according to claim 6 , the operations further comprising:

recognizing the speech information, and displaying a prompt message to prompt a user to perform manual switching if a recognition result does not meet a preset condition.

10. The apparatus according to claim 6 , wherein the recognizing the speech information and matching a linguistic category for the speech information to determine a corresponding target linguistic category based on a matching degree comprises:

recognizing the speech information and match a linguistic category for the speech information;

determining at least two candidate linguistic categories having matching degrees meeting a preset condition;

querying a user historical speech recognition record to determine a linguistic category used by the user historically; and

selecting a linguistic category consistent with the linguistic category used by the user historically from the at least two candidate linguistic categories as the target linguistic category.

11. A non-transitory computer storage medium storing a computer program, the computer program when executed by one or more processors, causes the one or more processors to perform operations, the operations comprising:

acquiring at least one piece of speech information in user input speech;

recognizing the speech information and matching a linguistic category for the speech information to determine a corresponding target linguistic category based on a matching degree; and

switching a currently used speech recognition model to a speech recognition model corresponding to the target linguistic category,

wherein the recognizing the speech information and matching a linguistic category for the speech information to determine a corresponding target linguistic category based on a matching degree comprises:

recognizing at least two pieces of speech sentences included in the speech information to obtain a matching degree between each speech sentence and a linguistic category; and

determining an initial linguistic category based on the matching degree, and calculating a product of probabilities of the speech sentences not belonging to the initial linguistic category, and determining the corresponding target linguistic category based on the product.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 14, 2020
From: JIANG, BING
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 054055/0050 →
Priority Claims (1)
CN 2016 1 0429948 · Jun 16, 2016 · national
Continuity (2)
Continuation PCTCN2016097417 · Aug 30, 2016
Related Publication 20190096396A1 · Mar 28, 2019