IP Library Granted Patent US 9,685,161
Granted Patent B2
US 9,685,161 · App. 14/585,486 · Granted Jun 20, 2017

Method for updating voiceprint feature model and terminal

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,685,161
App. No.
14/585,486
Granted
Jun 20, 2017
Kind
B2
Abstract

A method for updating a voiceprint feature model and a terminal are provided that are applicable to the field of voice recognition technologies. The method includes: obtaining an original audio stream including at least one speaker; obtaining a respective audio stream of each speaker of the at least one speaker in the original audio stream according to a preset speaker segmentation and clustering algorithm; separately matching the respective audio stream of each speaker of the at least one speaker with an original voiceprint feature model, to obtain a successfully matched audio stream; and using the successfully matched audio stream as an additional audio stream training sample for generating the original voiceprint feature model, and updating the original voiceprint feature model.

Claims (62)

1. A method for updating a voiceprint feature model, comprising:

obtaining an original audio stream comprising at least one speaker;

obtaining a respective audio stream of each speaker of the at least one speaker in the original audio stream according to a preset speaker segmentation and clustering algorithm;

separately matching the respective audio stream of each speaker of the at least one speaker with an original voiceprint feature model to obtain a successfully matched audio stream;

using the successfully matched audio stream as an additional audio stream training sample for generating the original voiceprint feature model; and

updating the original voiceprint feature model to improve a voice recognition capability of a computing device that uses the original voiceprint feature model to identify the at least one speaker.

2. The method according to claim 1 , wherein before obtaining the original audio stream comprising the at least one speaker, the method further comprises establishing the original voiceprint feature model according to a preset audio stream training sample.

3. The method according to claim 2 , wherein obtaining the respective audio stream of each speaker of the at least one speaker in the original audio stream according to the preset speaker segmentation and clustering algorithm comprises:

segmenting the original audio stream into a plurality of audio clips according to a preset speaker segmentation algorithm, wherein each audio clip of the plurality of audio clips comprises only audio information of a same speaker of the at least one speaker; and

clustering, according to a preset speaker clustering algorithm, the audio clips that comprise only the same speaker of the at least one speaker, to generate an audio stream that comprises only the audio information of the same speaker of the at least one speaker.

4. The method according to claim 3 , wherein separately matching the respective audio stream of each speaker of the at least one speaker with the original voiceprint feature model to obtain the successfully matched audio stream comprises:

obtaining a matching degree between the audio stream of each speaker of the at least one speaker and the original voiceprint feature model according to the audio stream of each speaker of the at least one speaker and the original voiceprint feature model; and

selecting an audio stream corresponding to a matching degree that is the highest and is greater than a preset matching threshold as the successfully matched audio stream.

5. The method according to claim 2 , wherein separately matching the respective audio stream of each speaker of the at least one speaker with the original voiceprint feature model to obtain the successfully matched audio stream comprises:

obtaining a matching degree between the audio stream of each speaker of the at least one speaker and the original voiceprint feature model according to the audio stream of each speaker of the at least one speaker and the original voiceprint feature model; and

selecting an audio stream corresponding to a matching degree that is the highest and is greater than a preset matching threshold as the successfully matched audio stream.

6. The method according to claim 1 , wherein obtaining the respective audio stream of each speaker of the at least one speaker in the original audio stream according to the preset speaker segmentation and clustering algorithm comprises:

segmenting the original audio stream into a plurality of audio clips according to a preset speaker segmentation algorithm, wherein each audio clip of the plurality of audio clips comprises only audio information of a same speaker of the at least one speaker; and

clustering, according to a preset speaker clustering algorithm, the audio clips that comprise only the same speaker of the at least one speaker to generate an audio stream that comprises only the audio information of the same speaker of the at least one speaker.

7. The method according to claim 6 , wherein separately matching the respective audio stream of each speaker of the at least one speaker with the original voiceprint feature model to obtain the successfully matched audio stream comprises:

obtaining a matching degree between the audio stream of each speaker of the at least one speaker and the original voiceprint feature model according to the audio stream of each speaker of the at least one speaker and the original voiceprint feature model; and

selecting an audio stream corresponding to a matching degree that is the highest and is greater than a preset matching threshold as the successfully matched audio stream.

8. The method according to claim 1 , wherein separately matching the respective audio stream of each speaker of the at least one speaker with the original voiceprint feature model to obtain a successfully matched audio stream comprises:

obtaining a matching degree between the audio stream of each speaker of the at least one speaker and the original voiceprint feature model according to the audio stream of each speaker of the at least one speaker and the original voiceprint feature model; and

selecting an audio stream corresponding to a matching degree that is the highest and is greater than a preset matching threshold as the successfully matched audio stream.

9. The method according to claim 1 , wherein using the successfully matched audio stream as the additional audio stream training sample for generating the original voiceprint feature model and updating the original voiceprint feature model comprises:

generating a corrected voiceprint feature model according to the successfully matched audio stream and the preset audio stream training sample, wherein the preset audio stream training sample is an audio stream for generating the original voiceprint feature model; and

updating the original voiceprint feature model to the corrected voiceprint feature model.

10. The method according to claim 1 , further comprising unlocking a screen of a mobile phone based upon matching the original voiceprint feature model.

11. A terminal, comprising:

a non-transitory computer readable medium having instructions stored thereon; and

a computer processor coupled to the non-transitory computer readable medium and configured to execute the instructions to:

obtain an original audio stream comprising at least one speaker;

obtain a respective audio stream of each speaker of the at least one speaker in the original audio stream according to a preset speaker segmentation and clustering algorithm;

separately match the respective audio stream of each speaker of the at least one speaker with an original voiceprint feature model, to obtain a successfully matched audio stream;

use the successfully matched audio stream as an additional audio stream training sample for generating the original voiceprint feature model; and

update the original voiceprint feature model to improve a voice recognition capability of a computing device that uses the original voiceprint feature model to identify the at least one speaker.

12. The terminal according to claim 11 , wherein the computer processor is further configured to execute the instructions to:

obtain a preset audio stream training sample; and

establish the original voiceprint feature model according to the preset audio stream training sample.

13. The terminal according to claim 12 , wherein the computer processor is further configured to execute the instructions to:

segment the original audio stream into a plurality of audio clips according to a preset speaker segmentation algorithm, wherein each audio clip of the plurality of audio clips comprises only audio information of a same speaker of the at least one speaker; and

cluster, according to a preset speaker clustering algorithm, the audio clips that comprise only the same speaker of the at least one speaker, to generate an audio stream that comprises only the audio information of the same speaker of the at least one speaker.

14. The terminal according to claim 13 , wherein the computer processor is further configured to execute the instructions to:

obtain a matching degree between the audio stream of each speaker of the at least one speaker and the original voiceprint feature model according to the audio stream of each speaker of the at least one speaker and the original voiceprint feature model; and

select an audio stream corresponding to a matching degree that is the highest and is greater than a preset matching threshold as the successfully matched audio stream.

15. The terminal according to claim 11 , wherein the computer processor is further configured to execute the instructions to:

segment the original audio stream into a plurality of audio clips according to a preset speaker segmentation algorithm, wherein each audio clip of the plurality of audio clips comprises only audio information of a same speaker of the at least one speaker; and

cluster, according to a preset speaker clustering algorithm, the audio clips that comprise only the same speaker of the at least one speaker to generate an audio stream that comprises only the audio information of the same speaker of the at least one speaker.

16. The terminal according to claim 15 , wherein the computer processor is further configured to execute the instructions to:

obtain a matching degree between the audio stream of each speaker of the at feast one speaker and the original voiceprint feature model according to the audio stream of each speaker of the at least one speaker and the original voiceprint feature model; and

select an audio stream corresponding to a matching degree that is the highest and is greater than a preset matching threshold as the successfully matched audio stream.

17. The terminal according to claim 11 , wherein the computer processor is further configured to execute the instructions to:

obtain a matching degree between the audio stream of each speaker of the at least one speaker and the original voiceprint feature model according to the audio stream of each speaker of the at least one speaker and the original voiceprint feature model; and

select an audio stream corresponding to a matching degree that is the highest and is greater than a preset matching threshold as the successfully matched audio stream.

18. The terminal according to claim 12 , wherein the computer processor is further configured to execute the instructions to:

obtain a matching degree between the audio stream of each speaker of the at least one speaker and the original voiceprint feature model according to the audio stream of each speaker of the at least one speaker and the original voiceprint feature model; and

select an audio stream corresponding to a matching degree that is the highest and is greater than a preset matching threshold as the successfully matched audio stream.

19. The terminal according to claim 11 , wherein the computer processor is further configured to execute the instructions to:

generate a corrected voiceprint feature model according to the successfully matched audio stream and the preset audio stream training sample; and

update the original voiceprint feature model to the corrected voiceprint feature model.

20. The terminal according to claim 11 , wherein the computer processor is further configured to execute the instructions to unlock a screen of a mobile phone based upon matching the original voiceprint feature model.

Assignments (3)
CHANGE OF NAME Recorded Mar 11, 2019
From: HUAWEI DEVICE (DONGGUAN) CO.,LTD.
To: HUAWEI DEVICE CO.,LTD.
Reel/Frame 048555/0951 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 5, 2017
From: HUAWEI DEVICE CO., LTD.
To: HUAWEI DEVICE (DONGGUAN) CO., LTD.
Reel/Frame 043750/0393 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2015
From: LU, TING
To: HUAWEI DEVICE CO., LTD.
Reel/Frame 034802/0149 →