IP Library Patent Application 16728696
Patent Application
App. No. 16/728,696

SPEECH CONTROL METHOD AND APPARATUS, ELECTRONIC DEVICE, AND READABLE STORAGE MEDIUM

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
16/728,696
Abstract

The present disclosure discloses a speech control method, a speech control apparatus, an electronic device, and a readable storage medium. The method may be applied to an electronic device, and includes: in a target scenario, controlling the electronic device to operate in a first operation state, and collecting an audio clip based on a wake word in the first operation state; performing speech recognition on the audio clip to obtain a first control intent; performing a first control instruction corresponding to the first control intent, and controlling the electronic device to switch from the first operation state to a second operation state; in the second operation state, continuously collecting audio to obtain an audio stream, and performing speech recognition on the audio stream to obtain a second control intent; and performing a second control instruction corresponding to the second control intent when the second control intent matches the target scenario.

Claims (59)

1 . A speech control method, applied to an electronic device, and comprising:

in a target scenario, controlling the electronic device to operate in a first operation state, and collecting an audio clip based on a wake word in the first operation state;

performing speech recognition on the audio clip to obtain a first control intent;

performing a first control instruction corresponding to the first control intent, and controlling the electronic device to switch from the first operation state to a second operation state;

in the second operation state, continuously collecting audio to obtain an audio stream, and performing speech recognition on the audio stream to obtain a second control intent; and

performing a second control instruction corresponding to the second control intent when the second control intent matches the target scenario.

2 . The speech control method of claim 1 , after continuously collecting audio to obtain the audio stream, and performing speech recognition on the audio stream to obtain the second control intent, further comprising:

performing speech recognition on the audio stream to obtain information stream;

obtaining at least one candidate intent based on the information stream;

selecting the second control intent matching the target scenario from the at least one candidate intent; and

controlling the electronic device to quit the second operation state when the second control intent matching the target scenario is not obtained within a preset period, the preset period ranging from 20 seconds to 40 seconds.

3 . The speech control method of claim 2 , after obtaining the at least one candidate intent based on the information stream, further comprising:

controlling the electronic device to reject responding to the candidate intent that does not match the target scenario.

4 . The speech control method of claim 1 , wherein controlling the electronic device to switch from the first operation state to the second operation state comprises:

replacing a first element with a second element, and displaying a third element, wherein the first element is configured to indicate that the electronic device is in the first operation state, the second element is configured to indicate that the electronic device is in the second operation state, and the third element is configured to prompt inputting the wake word and/or broadcasting an audio or video.

5 . The speech control method of claim 1 , before controlling the electronic device to switch from the first operation state to the second operation state, further comprising:

determining that the first control intent matches the target scenario.

6 . The speech control method of claim 1 , wherein the target scenario comprises a game scenario.

7 . A speech control apparatus, comprising:

at least one processor; and

a memory, configured to store instructions, and coupled to the at least one processor;

wherein when the instructions are executed by the at least one processor, the at least one processor is caused to:

in a target scenario, control an electronic device to operate in a first operation state, and collect an audio clip based on a wake word in the first operation state;

perform speech recognition on the audio clip to obtain a first control intent;

perform a first control instruction corresponding to the first control intent, and control the electronic device to switch from the first operation state to a second operation state;

in the second operation state, continuously collect audio to obtain an audio stream, and perform speech recognition on the audio stream to obtain a second control intent; and

perform a second control instruction corresponding to the second control intent when the second control intent matches the target scenario.

8 . The speech control apparatus of claim 7 , wherein the at least one processor is further configured to:

perform speech recognition on the audio stream to obtain information stream;

obtain at least one candidate intents based on the information stream;

select the second control intent matching the control intent of the target scenario from the at least one candidate intents; and

control the electronic device to quit the second operation state when the second control intent matching the target scenario is not obtained within a preset period, the preset period ranging from 20 seconds to 40 seconds.

9 . The speech control apparatus of claim 8 , wherein the at least one processor is further configured to: control the electronic device to reject responding to the candidate intent that does not match the target scenario.

10 . The speech control apparatus of claim 7 , wherein the at least one processor is further configured to:

replace a first element with a second element, and display a third element, wherein the first element is configured to indicate that the electronic device is in the first operation state, the second element is configured to indicate that the electronic device is in the second operation state, and the third element is configured to indicate inputting the wake word and/or broadcasting an audio or video.

11 . The speech control apparatus of claim 7 , wherein the at least one processor is further configured to: determine that the first control intent matches the target scenario.

12 . The speech control apparatus of claim 7 , wherein the target scenario comprises a game scenario.

13 . A non-transitory computer readable storage medium having computer instructions stored thereon, wherein when the computer instructions are executed by a processor, the processor is caused execute a speech control method, wherein the speech control method is applied to an electronic device, and comprises:

in a target scenario, controlling the electronic device to operate in a first operation state, and collecting an audio clip based on a wake word in the first operation state;

performing speech recognition on the audio clip to obtain a first control intent;

performing a first control instruction corresponding to the first control intent, and controlling the electronic device to switch from the first operation state to a second operation state;

in the second operation state, continuously collecting audio to obtain an audio stream, and performing speech recognition on the audio stream to obtain a second control intent; and

performing a second control instruction corresponding to the second control intent when the second control intent matches the target scenario.

14 . The non-transitory computer readable storage medium of claim 13 , wherein after continuously collecting audio to obtain the audio stream, and performing speech recognition on the audio stream to obtain the second control intent, the method further comprises:

performing speech recognition on the audio stream to obtain information stream;

obtaining at least one candidate intent based on the information stream;

selecting the second control intent matching the target scenario from the at least one candidate intent; and

controlling the electronic device to quit the second operation state when the second control intent matching the target scenario is not obtained within a preset period, the preset period ranging from 20 seconds to 40 seconds.

15 . The non-transitory computer readable storage medium of claim 14 , wherein after obtaining the at least one candidate intent based on the information stream, the method further comprises:

controlling the electronic device to reject responding to the candidate intent that does not match the target scenario.

16 . The non-transitory computer readable storage medium of claim 13 , wherein controlling the electronic device to switch from the first operation state to the second operation state comprises:

replacing a first element with a second element, and displaying a third element, wherein the first element is configured to indicate that the electronic device is in the first operation state, the second element is configured to indicate that the electronic device is in the second operation state, and the third element is configured to prompt inputting the wake word and/or broadcasting an audio or video.

17 . The non-transitory computer readable storage medium of claim 13 , wherein before controlling the electronic device to switch from the first operation state to the second operation state, the method further comprises:

determining that the first control intent matches the target scenario.

18 . The non-transitory computer readable storage medium of claim 13 , wherein the target scenario comprises a game scenario.

19 . The speech control apparatus of claim 8 , wherein the at least one processor is further configured to:

replace a first element with a second element, and display a third element, wherein the first element is configured to indicate that the electronic device is in the first operation state, the second element is configured to indicate that the electronic device is in the second operation state, and the third element is configured to indicate inputting the wake word and/or broadcasting an audio or video.

20 . The speech control apparatus of claim 9 , wherein the at least one processor is further configured to:

replace a first element with a second element, and display a third element, wherein the first element is configured to indicate that the electronic device is in the first operation state, the second element is configured to indicate that the electronic device is in the second operation state, and the third element is configured to indicate inputting the wake word and/or broadcasting an audio or video.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2021
From: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.; SHANGHAI XIAODU TECHNOLOGY CO. LTD.
Reel/Frame 056811/0772 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 27, 2019
From: LUO, YONGXI; WANG, SHASHA
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 051378/0111 →