IP Library Patent Application 16730534
Patent Application
App. No. 16/730,534

SPEECH RECOGNITION CONTROL METHOD AND APPARATUS, ELECTRONIC DEVICE AND READABLE STORAGE MEDIUM

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
16/730,534
Abstract

The present disclosure discloses a speech recognition control method, a speech recognition control apparatus, an electronic device and a readable storage medium. The method includes: querying configuration information of a first operation state to determine whether the first operation state is applicable to the target scene; switching a second operation state to the first operation state in response to determining that the first operation state is applicable to the target scene; in the second operation state, acquiring an audio clip based on a wake-up word to perform speech recognition on the audio clip; and in the first operation state, continuously acquiring audio to obtain an audio stream to perform the speech recognition on the audio stream.

Claims (60)

1 . A speech recognition control method, comprising:

querying configuration information of a first operation state to determine whether the first operation state is applicable to a target scene;

switching a second operation state to the first operation state in response to determining that the first operation state is applicable to the target scene; wherein, in the second operation state, an audio clip is acquired based on a wake-up word to perform speech recognition on the audio clip; and

in the first operation state, continuously acquiring audio to obtain an audio stream to perform the speech recognition on the audio stream.

2 . The speech recognition control method according to claim 1 , further comprising:

detecting whether an application programmer interface relevant to the target scene is called, and in response to detecting that the application programmer interface is called, querying the configuration information of the first operation state; or

determining whether the target scene is detected, and querying the configuration information of the first operation state, in response to detecting the target scene.

3 . The speech recognition control method according to claim 1 , further comprising:

in the second operation state, acquiring a first control intention by performing the speech recognition on the audio clip; and

determining whether the first control intention matches the target scene.

4 . The speech recognition control method according to claim 1 , further comprising:

acquiring an information stream; wherein, the information stream is obtained by performing the speech recognition on the audio stream;

acquiring one or more candidate intentions from the information stream;

obtaining a second control intention matched with the control intention of the target scene from the one or more candidate intentions; and

in response to obtaining the second control intention, executing a control instruction corresponding to the second control intention.

5 . The speech recognition control method according to claim 4 , further comprising:

in response to not obtaining the second control intention within a preset duration, quitting the first operation state; wherein, the preset duration ranges from 20 seconds to 40 seconds.

6 . The speech recognition control method according to claim 4 , further comprising:

refusing to respond to a candidate intention that does not match the control intention of the target scene.

7 . The speech recognition control method according to claim 1 , wherein the configuration information comprises a list of scenes that the first operation state is applicable to, and the list of scenes is generated by selecting one or more of a music scene, an audiobook scene and a video scene, in response to user selection.

8 . An electronic device, comprising:

at least one processor; and

a memory connected in communication with the at least one processor;

wherein the memory is configured to store instructions executable by the at least one processor, and the instructions are executed by the at least one processor such that the at least one processor is configured to:

query configuration information of a first operation state to determine whether the first operation state is applicable to a target scene;

switch a second operation state to the first operation state in response to determining that the first operation state is applicable to the target scene; wherein, in the second operation state, an audio clip is acquired based on a wake-up word to perform speech recognition on the audio clip; and

in the first operation state, continuously acquire audio to obtain an audio stream to perform the speech recognition on the audio stream.

9 . The electronic device according to claim 8 , wherein the at least one processor is further configured to:

detect whether an application programmer interface relevant to the target scene is called, and in response to detecting that the application programmer interface is called, query the configuration information of the first operation state; or

determine whether the target scene is detected, and query the configuration information of the first operation state, in response to detecting the target scene.

10 . The electronic device according to claim 8 , wherein the at least one processor is further configured to:

in the second operation state, acquire a first control intention by performing the speech recognition on the audio clip; and

determine whether the first control intention matches the target scene.

11 . The electronic device according to claim 8 , wherein the at least one processor is further configured to:

acquire an information stream; wherein, the information stream is obtained by performing the speech recognition on the audio stream;

acquire one or more candidate intentions from the information stream;

obtain a second control intention matched with the control intention of the target scene from the one or more candidate intentions; and

in response to obtaining the second control intention, execute a control instruction corresponding to the second control intention.

12 . The electronic device according to claim 11 , wherein the at least one processor is further configured to:

in response to not obtaining the second control intention within a preset duration, quit the first operation state; wherein, the preset duration ranges from 20 seconds to 40 seconds.

13 . The electronic device according to claim 11 , wherein the at least one processor is further configured to:

refuse to respond to a candidate intention that does not match the control intention of the target scene.

14 . The electronic device according to claim 8 , wherein the configuration information comprises a list of scenes that the first operation state is applicable to, and the list of scenes is generated by selecting one or more of a music scene, an audiobook scene and a video scene, in response to user selection.

15 . A non-transitory computer-readable storage medium, having computer instructions stored thereon, wherein the computer instructions are executed by a computer such that the computer is configured to execute a speech recognition control method, the method comprises:

querying configuration information of a first operation state to determine whether the first operation state is applicable to a target scene;

switching a second operation state to the first operation state in response to determining that the first operation state is applicable to the target scene; wherein, in the second operation state, an audio clip is acquired based on a wake-up word to perform speech recognition on the audio clip; and

in the first operation state, continuously acquiring audio to obtain an audio stream to perform the speech recognition on the audio stream.

16 . The non-transitory computer-readable storage medium according to claim 15 , wherein the method further comprises:

in the second operation state, acquiring a first control intention by performing the speech recognition on the audio clip; and

determining whether the first control intention matches the target scene.

17 . The non-transitory computer-readable storage medium according to claim 15 , wherein the method further comprises:

acquiring an information stream; wherein, the information stream is obtained by performing the speech recognition on the audio stream;

acquiring one or more candidate intentions from the information stream;

obtaining a second control intention matched with the control intention of the target scene from the one or more candidate intentions; and

in response to obtaining the second control intention, executing a control instruction corresponding to the second control intention.

18 . The non-transitory computer-readable storage medium according to claim 17 , wherein the method further comprises:

in response to not obtaining the second control intention within a preset duration, quitting the first operation state; wherein, the preset duration ranges from 20 seconds to 40 seconds.

19 . The non-transitory computer-readable storage medium according to claim 17 , wherein the method further comprises:

refusing to respond to a candidate intention that does not match the control intention of the target scene.

20 . The non-transitory computer-readable storage medium according to claim 15 , wherein the configuration information comprises a list of scenes that the first operation state is applicable to, and the list of scenes is generated by selecting one or more of a music scene, an audiobook scene and a video scene, in response to user selection.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2021
From: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.; SHANGHAI XIAODU TECHNOLOGY CO. LTD.
Reel/Frame 056811/0772 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 30, 2019
From: LUO, YONGXI; WANG, SHASHA
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 051388/0678 →