IP Library Granted Patent US 10,600,415
Granted Patent B2
US 10,600,415 · App. 16/133,336 · Granted Mar 24, 2020

Method, apparatus, device, and storage medium for voice interaction

Inventors: Jianan Xu (Beijing, CN); Guoguo Chen (Beijing, CN); Qinggeng Qian (Beijing, CN)
Assignee: Baidu Online Network Technology (Beijing) Co., Ltd.
G10L15/22G10L15/1815G10L17/04G10L17/22G10L2015/223G10L2015/225G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,600,415
App. No.
16/133,336
Granted
Mar 24, 2020
Kind
B2
Abstract

This disclosure provides a method, apparatus, device, and storage medium for voice interaction, where the method is applied to an AI device to determine whether a current scenario of the AI device is a preset scenario and waken a voice interaction function of the AI device to facilitate voice interaction with a user in response to the current scenario of the AI device being the preset scenario. A scenario directly triggers the voice interaction process, thereby avoiding the process of wakening by physical wakening or a wakening word, simplifying the process of using voice interaction, reducing the costs of learning voice interaction, and improving user experience.

Claims (113)

1. A method for voice interaction, applied to an artificial intelligence (AI) device, and comprising:

determining whether a current scenario of the AI device is a preset scenario;

wakening a voice interaction function of the AI device to facilitate voice interaction with a user, in response to the current scenario of the AI device being the preset scenario,

wherein the wakening a voice interaction function of the AI device to facilitate voice interaction with a user comprises:

acquiring voice data of the user;

identifying and understanding the voice data using an acoustic model and a semantic understanding model to obtain a semantic understanding result; and

executing an operation indicated by the semantic understanding result when a confidence level of the semantic understanding result is greater than a preset threshold, and

wherein the method is performed by at least one processor.

2. The method according to claim 1 , wherein the determining whether a current scenario of the AI device is a preset scenario comprises:

detecting whether an operation state of the AI device is changed; and

determining, in response to the operation state of the AI device being changed, whether a scenario of the AI device is the preset scenario after the operation state is changed;

or,

receiving a scenario setting instruction entered by a user on the AI device; and

determining whether the current scenario of the AI device is the preset scenario based on the scenario setting instruction;

or,

periodically detecting and determining whether the current scenario of the AI device is the preset scenario based on a preset period;

or,

detecting whether a microphone of the AI device is in an on-state; and

determining whether the current scenario of the AI device is the preset scenario in response to the microphone being in the on-state.

3. The method according to claim 1 , wherein

the preset scenario comprises a calling scenario, and the determining whether a current scenario of the AI device is a preset scenario comprises:

detecting whether the AI device is in a calling process or receives a request for calling; and

determining the current scenario of the AI device is the preset scenario in response to the AI device being in the calling process or receiving the request for calling;

or,

the preset scenario comprises a media file playing scenario, and the determining whether a current scenario of the AI device is a preset scenario comprises:

detecting whether the AI device is playing a media file, the media file comprising at least one of an image file, an audio file, or a video file; and

determining the current scenario of the AI device is the preset scenario in response to the AI device being playing the media file;

or,

the preset scenario comprises a mobile scenario, and the determining whether a current scenario of the AI device is a preset scenario comprises:

detecting a moving speed of the AI device, and determining whether the moving speed is greater than a preset value; and

determining the current scenario of the AI device is the preset scenario in response to the moving speed being greater than the preset value;

or,

the preset scenario comprises a messaging scenario, and the determining whether a current scenario of the AI device is a preset scenario comprises:

detecting whether the AI device receives a short message or a notification message; and

determining the current scenario of the AI device is the preset scenario in response to the AI device receiving the short message or the notification message.

4. The method according to claim 1 , wherein the wakening a voice interaction function of the AI device to facilitate voice interaction with a user comprises:

performing voice interaction based on the voice data and a preset instruction set corresponding to the current scenario of the AI device.

5. The method according to claim 1 , wherein the acquiring voice data of the user comprises:

controlling the microphone of the AI device to collect the voice data of the user;

or,

controlling a Bluetooth or a headset microphone connected to the AI device to collect a voice of the user and acquire the voice data of the user;

or,

receiving the voice data of the user sent by an other device.

6. The method according to claim 1 , wherein before the identifying and understanding the voice data using an acoustic model and a semantic understanding model, the method further comprises:

processing the voice data by noise cancellation and echo cancellation.

7. The method according to claim 1 , wherein the identifying and understanding the voice data using an acoustic model and a semantic understanding model to obtain a semantic understanding result comprises:

matching the voice data using the acoustic model to identify semantic data; and

understanding and analyzing the semantic data based on the semantic understanding model to obtain the semantic understanding result.

8. The method according to claim 1 , further comprising:

evaluating the confidence level of the semantic understanding result based on the current scenario of the AI device, the instruction set corresponding to the current scenario of the AI device, and a state of the AI device;

determining whether the confidence level of the semantic understanding result is greater than the preset threshold; and

discarding the executing an operation indicated by the semantic understanding result in response to determining that the confidence level of the semantic understanding result is smaller than the preset threshold.

9. The method according to claim 1 , wherein the executing the operation indicated by the semantic understanding result comprises:

outputting the semantic understanding result to a software interface for execution through a specified instruction.

10. An apparatus for voice interaction, comprising:

at least one processor; and

a memory storing instructions, the instructions when executed by the at least one processor, cause the at least one processor to perform operations, the operations comprising:

determining whether a current scenario of an apparatus for voice interaction is a preset scenario; and

wakening a voice interaction function of the apparatus for voice interaction to facilitate voice interaction with a user, in response to the current scenario of the apparatus for voice interaction being the preset scenario,

wherein the wakening a voice interaction function of the AI device to facilitate voice interaction with a user comprises:

acquiring voice data of the user;

identifying and understanding the voice data using an acoustic model and a semantic understanding model to obtain a semantic understanding result; and

executing an operation indicated by the semantic understanding result when a confidence level of the semantic understanding result is greater than a preset threshold.

11. The apparatus according to claim 10 , wherein the determining whether a current scenario of the AI device is a preset scenario comprises:

detecting whether an operation state of the apparatus for voice interaction is changed; and

determining, in response to the operation state of the apparatus for voice interaction being changed, whether a scenario of the apparatus for voice interaction is the preset scenario after the operation state is changed;

or,

receiving a scenario setting instruction entered by a user on the apparatus for voice interaction; and

determining whether the current scenario of the apparatus for voice interaction is the preset scenario based on the scenario setting instruction;

or,

periodically detecting and determining whether the current scenario of the apparatus for voice interaction is the preset scenario based on a preset period;

or,

detecting whether a microphone of the apparatus for voice interaction is in an on-state; and

determining whether the current scenario of the apparatus for voice interaction is the preset scenario in response to the microphone being in the on-state.

12. The apparatus according to claim 10 , wherein

the preset scenario comprises a calling scenario, and the first processing module is further used for:

detecting whether the apparatus for voice interaction is in a calling process or receives a request for calling; and

determining the current scenario of the apparatus for voice interaction is the preset scenario in response to the apparatus for voice interaction being in the calling process or receiving the request for calling;

or,

the preset scenario comprises a media file playing scenario, and the first processing module is further used for:

detecting whether the apparatus for voice interaction is playing a media file, the media file comprising at least one of an image file, an audio file, or a video file; and

determining the current scenario of the apparatus for voice interaction is the preset scenario in response to the apparatus for voice interaction being playing the media file;

or,

the preset scenario comprises a mobile scenario, and the first processing module is further used for:

detecting a moving speed of the apparatus for voice interaction, and determining whether the moving speed is greater than a preset value; and

determining the current scenario of the apparatus for voice interaction is the preset scenario in response to the moving speed being greater than the preset value;

or,

the preset scenario comprises a messaging scenario, and the first processing module is further used for:

detecting whether the apparatus for voice interaction receives a short message or a notification message; and

determining the current scenario of the apparatus for voice interaction is the preset scenario in response to the apparatus for voice interaction receiving the short message or the notification message.

13. The apparatus according to claim 10 , wherein the wakening a voice interaction function of the AI device to facilitate voice interaction with a user comprises:

performing voice interaction based on the voice data and a preset instruction set corresponding to the current scenario of the apparatus for voice interaction.

14. The apparatus according to claim 10 , wherein the acquiring voice data of the user comprises:

controlling the microphone of the apparatus for voice interaction to collect the voice data of the user;

or,

controlling a Bluetooth or a headset microphone connected to the apparatus for voice interaction to collect a voice of the user and acquire the voice data of the user;

or,

receiving the voice data of the user sent by an other device.

15. The apparatus according to 10 , wherein before the identifying and understanding the voice data using an acoustic model and a semantic understanding model, the operations further comprises:

processing the voice data by noise cancellation and echo cancellation.

16. The apparatus according to claim 10 , wherein the identifying and understanding the voice data using an acoustic model and a semantic understanding model to obtain a semantic understanding result comprises:

matching the voice data using the acoustic model to identify semantic data; and

understanding and analyzing the semantic data based on the semantic understanding model to obtain the semantic understanding result.

17. The apparatus according to claim 10 , wherein the operation further comprises:

evaluating the confidence level of the semantic understanding result based on the current scenario of the apparatus for voice interaction, the instruction set corresponding to the current scenario of the apparatus for voice interaction, and a state of the apparatus for voice interaction;

determining whether the confidence level of the semantic understanding result is greater than the preset threshold; and

discarding the executing an operation indicated by the semantic understanding result in response to determining that the confidence level of the semantic understanding result is smaller than the preset threshold.

18. A non-transitory computer readable storage medium storing a computer program, wherein the computer program, when executed by a processor, cause the processor to perform operations, the operation comprising:

determining whether a current scenario of the AI device is a preset scenario; and wakening a voice interaction function of the AI device to facilitate voice interaction with a user, in response to the current scenario of the AI device being the preset scenario,

wherein the wakening a voice interaction function of the AI device to facilitate voice interaction with a user comprises:

acquiring voice data of the user;

identifying and understanding the voice data using an acoustic model and a semantic understanding model to obtain a semantic understanding result; and

executing an operation indicated by the semantic understanding result when a confidence level of the semantic understanding result is greater than a preset threshold.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2021
From: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.; SHANGHAI XIAODU TECHNOLOGY CO. LTD.
Reel/Frame 056811/0772 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 12, 2018
From: CHEN, GUOGUO; QIAN, QINGGENG
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 047759/0191 →
Priority Claims (1)
CN 2017 1 1427997 · Dec 26, 2017 · national
Continuity (1)
Related Publication 20190198019A1 · Jun 27, 2019