IP Library Granted Patent US 11,322,138
Granted Patent B2
US 11,322,138 · App. 16/268,865 · Granted May 3, 2022

Voice awakening method and device

Inventors: Jun Li (Beijing, CN); Rui Yang (Beijing, CN); Lifeng Zhao (Beijing, CN); Xiaojian Chen (Beijing, CN); Yushu Cao (Beijing, CN)
Assignees: Baidu Online Network Technology (Beijing) Co., Ltd.; Shanghai Xiaodu Technology Co., Ltd.
G10L15/22G06N3/08G10L15/063G10L15/08G10L15/16G10L15/32G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,322,138
App. No.
16/268,865
Granted
May 3, 2022
Kind
B2
Abstract

A voice awakening method and device are provided. According to an embodiment, the method includes: receiving voice information of a user; obtaining an awakening confidence level corresponding to the voice information based on the voice information; determining, on the basis of the awakening confidence level, whether the voice information is suspected wake-up voice information; and performing, in response to determining the voice information being the suspected wake-up voice information, a secondary determination on the voice information to obtain a secondary determination result, and determining whether to perform a wake-up operation on the basis of the secondary determination result. The embodiment implements a secondary verification on the voice information, thereby reducing the probability that the smart device is mistakenly awakened.

Claims (60)

1. A voice awakening method, comprising:

receiving voice information of a user;

inputting the voice information into a pre-established recognition model to obtain an awakening confidence level for the voice information by: extracting characteristic information in the voice information to obtain a characteristic vector, and obtaining the awakening confidence level for the voice information according to a pre-established corresponding relationship table, the pre-established corresponding relationship table storing a plurality of corresponding relationships between characteristic vectors and awakening confidence levels;

determining, on the basis of the awakening confidence level, whether the voice information is suspected wake-up voice information; and

performing, in response to determining the voice information being the suspected wake-up voice information, a secondary determination on the voice information to obtain a secondary determination result and determining whether to perform a wake-up operation on the basis of the secondary determination result;

wherein determining the voice information being the suspected wake-up voice information comprises:

in response to the awakening confidence level being greater than a preset first threshold and less than a preset second threshold, determining that the voice information is the suspected wake-up voice information.

2. The method according to claim 1 , wherein the recognition model is used to represent a corresponding relationship between the voice information and the awakening confidence level.

3. The method according to claim 2 , wherein the recognition model is a neural network model, and the neural network model is trained by:

acquiring a sample set, wherein a sample includes sample voice information and an annotation indicating whether the sample voice information is wake-up voice information; and

performing following training: inputting respectively at least one piece of sample voice information in the sample set into an initial neural network model, to obtain a prediction corresponding to each of the at least one piece of sample voice information, wherein the prediction represents a probability of the sample voice information being the wake-up voice information; comparing the prediction corresponding to the each of the at least one piece of sample voice information with the annotation; determining, on the basis of the comparison result, whether the initial neural network model achieves a preset optimization objective; and using, in response to determining the initial neural network model achieving the preset optimization objective, the initial neural network model as a trained neural network model.

4. The method according to claim 3 , wherein the training the neural network model further comprises:

adjusting, in response to determining the initial neural network model not achieving the preset optimization objective, a network parameter of the initial neural network model, and continuing to perform the training.

5. The method according to claim 3 , wherein the annotation includes a first identifier and a second identifier, the first identifier represents that the sample voice information is the wake-up voice information, and the second identifier represents that the sample voice information is not the wake-up voice information.

6. The method according to claim 1 , wherein the secondary determination result includes a wake-up confirmation and a non-wake-up confirmation, and

the performing a secondary determination on the voice information to obtain a secondary determination result and determining whether to perform a wake-up operation on the basis of the secondary determination result comprises:

sending the voice information to a server end, and generating, by the server end, the secondary determination result based on the voice information;

receiving the secondary determination result; and

performing, in response to determining the secondary determination result being the wake-up confirmation, the wake-up operation.

7. A voice awakening method, comprising:

receiving suspected wake-up voice information sent by a terminal, wherein an awakening confidence level corresponding to the suspected wake-up voice information is greater than a preset first threshold and less than a preset second threshold, and the wakening confidence level is obtained by inputting voice information of a user into a pre-established recognition model, extracting characteristic information in the voice information to obtain a characteristic vector, and obtaining the awakening confidence level for the voice information according to a pre-established corresponding relationship table, the pre-established corresponding relationship table storing a plurality of corresponding relationships between characteristic vectors and awakening confidence levels;

performing speech recognition on the suspected wake-up voice information to obtain a speech recognition result;

comparing the speech recognition result and a target wake-up word of the terminal; and

sending, based on a comparison result, a secondary determination result to the terminal, and determining, by the terminal, whether to perform a wake-up operation on the basis of the secondary determination result, the secondary determination result including a wake-up confirmation and a non-wake-up confirmation.

8. The method according to claim 7 , wherein the sending, based on a comparison result, a secondary determination result to the terminal comprises:

sending, in response to determining the speech recognition result matching the target wake-up word, the wake-up confirmation to the terminal; and

sending, in response to determining the speech recognition result not matching the target wake-up word, the non-wake-up confirmation to the terminal.

9. A voice awakening device, comprising:

at least one processor; and

a memory storing instructions, the instructions when executed by the at least one processor, cause the at least one processor to perform operations, the operations comprising:

receiving voice information of a user;

inputting the voice information into a pre-established recognition model to obtain an awakening confidence level for the voice information by: extracting characteristic information in the voice information to obtain a characteristic vector, and obtaining the awakening confidence level for the voice information according to a pre-established corresponding relationship table, the pre-established corresponding relationship table storing a plurality of corresponding relationships between characteristic vectors and awakening confidence levels;

determining, on the basis of the awakening confidence level, whether the voice information is suspected wake-up voice information; and

performing, in response to determining the voice information being the suspected wake-up voice information, a secondary determination on the voice information to obtain a secondary determination result, and determining whether to perform a wake-up operation on the basis of the secondary determination result;

wherein determining the voice information being the suspected wake-up voice information comprises:

in response to the awakening confidence level being greater than a preset first threshold and less than a preset second threshold, determining that the voice information is the suspected wake-up voice information.

10. The device according to claim 9 , wherein the recognition model is used to represent a corresponding relationship between the voice information and the awakening confidence level.

11. The device according to claim 10 , wherein the recognition model is a neural network model, and the neural network model is trained by:

acquiring a sample set, wherein a sample includes sample voice information and an annotation indicating whether the sample voice information is wake-up voice information; and

performing following training: inputting respectively at least one piece of sample voice information in the sample set into an initial neural network model, to obtain a prediction corresponding to each of the at least one piece of sample voice information, wherein the prediction represents a probability of the sample voice information being the wake-up voice information; comparing the prediction corresponding to the each of the at least one piece of sample voice information with the annotation; determining, on the basis of a comparison result, whether the initial neural network model achieves a preset optimization objective; and using, in response to determining the initial neural network model achieving the preset optimization objective, the initial neural network model as a trained neural network model.

12. The device according to claim 11 , wherein the training the neural network model further comprises:

adjusting, in response to determining the initial neural network model not achieving the preset optimization objective, a network parameter of the initial neural network model, and continuing to perform the training.

13. The device according to claim 11 , wherein the annotation includes a first identifier and a second identifier, the first identifier represents that the sample voice information is the wake-up voice information, and the second identifier represents that the sample voice information is not the wake-up voice information.

14. The device according to claim 9 , wherein the secondary determination result includes a wake-up confirmation and a non-wake-up confirmation, and

the performing a secondary determination on the voice information to obtain a secondary determination result and determining whether to perform a wake-up operation on the basis of the secondary determination result comprises:

sending the voice information to a server end, and generating, by the server end, the secondary determination result based on the voice information;

receiving the secondary determination result; and

performing, in response to determining the secondary determination result being the wake-up confirmation, the wake-up operation.

15. A voice awakening device, comprising:

at least one processor; and

a memory storing instructions, the instructions when executed by the at least one processor, cause the at least one processor to perform operations, the operations comprising:

receiving suspected wake-up voice information sent by a terminal, wherein an awakening confidence level corresponding to the suspected wake-up voice information is greater than a preset first threshold and less than a preset second threshold, and the wakening confidence level is obtained by inputting voice information of a user into a pre-established recognition model, extracting characteristic information in the voice information to obtain a characteristic vector, and obtaining the awakening confidence level for the voice information according to a pre-established corresponding relationship table, the pre-established corresponding relationship table storing a plurality of corresponding relationships between characteristic vectors and awakening confidence levels;

performing speech recognition on the suspected wake-up voice information to obtain a speech recognition result;

comparing the speech recognition result and a target wake-up word of the terminal; and

sending, based on a comparison result, a secondary determination result to the terminal, and determining, by the terminal, whether to perform a wake-up operation on the basis of the secondary determination result, the secondary determination result including a wake-up confirmation and a non-wake-up confirmation.

16. The device according to claim 15 , wherein the sending, based on a comparison result, a secondary determination result to the terminal comprises:

sending, in response to determining the speech recognition result matching the target wake-up word, the wake-up confirmation to the terminal; and

sending, in response to determining the speech recognition result not matching the target wake-up word, the non-wake-up confirmation to the terminal.

17. A non-transitory computer-readable storage medium storing a computer program, the computer program when executed by one or more processors, causes the one or more processors to perform operations of the method of claim 1 .

18. A non-transitory computer-readable storage medium storing a computer program, the computer program when executed by one or more processors, causes the one or more processors to perform operations of the method of claim 7 .

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2021
From: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.; SHANGHAI XIAODU TECHNOLOGY CO. LTD.
Reel/Frame 056811/0772 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2020
From: LI, JUN; YANG, RUI; ZHAO, LIFENG; CHEN, XIAOJIAN; CAO, YUSHU
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 053624/0850 →
Priority Claims (1)
CN 201810133876.1 · Feb 9, 2018 · national
Continuity (1)
Related Publication 20190251963A1 · Aug 15, 2019
Cited By (1)
US 12,393,617