IP Library Granted Patent US 10,593,334
Granted Patent B2
US 10,593,334 · App. 15/484,082 · Granted Mar 17, 2020

Method and apparatus for generating voiceprint information comprised of reference pieces each used for authentication

Inventor: Jian Xiong (Zhejiang, CN)
Assignee: ALIBABA GROUP HOLDING LIMITED
G10L17/22G10L15/26G10L17/00G10L17/005G10L17/02G10L17/04G10L17/06G06F21/32G10L17/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,593,334
App. No.
15/484,082
Filed
Apr 10, 2017
Granted
Mar 17, 2020
Kind
B2
Art Unit
2657
USPC
704/235
Abstract

A method for generating voiceprint information is provided. The method includes acquiring a historical voice file generated by a call between a first user and a second user; executing text recognition processing on the voice information to obtain text information corresponding to the voice information; and storing the voice information and the corresponding text information as reference voiceprint information of the first user, and storing an identifier of the first user. Furthermore each voiceprint information comprises a plurality of pieces of reference voiceprint information, each of which is sufficient to authenticate a user.

Claims (105)

1. A method for generating voiceprint information, comprising:

acquiring a plurality of historical voice files generated by a plurality of calls between a first user and one or more second users;

executing filtering processing on the plurality of historical voice files to obtain voice information of the first user, wherein the voice information includes a plurality of pieces of reference voiceprint information of the first user;

executing text recognition processing on the plurality of pieces of reference voiceprint information to obtain text information corresponding to the plurality of pieces of reference voiceprint information; and

storing an identifier of the first user and the plurality of pieces of reference voiceprint information and the corresponding text information;

wherein a randomly selected piece of the reference voiceprint information is sufficient for authenticating the first user.

2. The method according to claim 1 , further comprising:

segmenting the text information into a plurality of pieces of sub-text information;

marking a start time and an end time of each piece of the sub-text information; and

acquiring, according to the start time and the end time of the sub-text information, sub-voice information corresponding to each piece of the sub-text information from the voice information.

3. The method according to claim 2 , wherein storing the plurality of pieces of reference voiceprint information and the corresponding text information comprises:

storing each pair of sub-voice information and sub-text information as a piece of reference voiceprint information of the first user.

4. The method according to claim 1 , wherein storing the reference voiceprint information and the identifier of the first user comprises:

determining whether second reference voiceprint information exists in a voiceprint library, wherein the second reference voiceprint information includes second text information and second identifier, the second text info nation is the same as the text information in the reference voiceprint information, and the second identifier is the same as the identifier of the first user;

in response to the second reference voiceprint information existing in the voiceprint library, comparing a quality of the voice information in the reference voiceprint information with a quality of second voice information in the second reference voiceprint information;

in response to the quality of the voice information being lower than the quality of the second voice information, deleting the reference voiceprint information; and

in response to the quality of the voice information is higher than the quality of the second voice information, deleting the second reference voiceprint information.

5. A system for generating voiceprint information, comprising:

a voice filter configured to acquire a plurality of historical voice files generated by a plurality of calls between a first user and one or more second users and execute filtering processing on the plurality of historical voice files to obtain voice information of the first user, wherein the voice information includes a plurality of pieces of reference voiceprint information of the first user;

a text recognizer configured to execute text recognition processing on the plurality of pieces of reference voiceprint information to obtain text information corresponding to the plurality of pieces of reference voiceprint information; and

a voiceprint generator configured to store the plurality of pieces of reference voiceprint information and the corresponding text information;

wherein a randomly selected piece of the reference voiceprint information is sufficient for authenticating the first user.

6. The system according to claim 5 , further comprising:

a text segmenter configured to segment the text information into a plurality of pieces of sub-text information, and mark a start time and an end time of each piece of the sub-text information; and

a voiceprint segmenter configured to acquire, according to the start time and the end time of the sub-text information, sub-voice information corresponding to each piece of the sub-text information from the voice information.

7. The system according to claim 6 , wherein the voiceprint generator is further configured to store each pair of sub-voice information and sub-text information as a piece of reference voiceprint information of the first user.

8. The system according to claim 5 , wherein the voiceprint generator is further configured to:

determine whether second reference voiceprint information exists in a voiceprint library, wherein the second reference voiceprint information includes second text information and second identifier, the second text information is the same as the text information in the reference voiceprint information, and the second identifier is the same as the identifier of the first user;

in response to the second reference voiceprint information existing in the voiceprint library, comparing a quality of the voice information in the reference voiceprint information with a quality of second voice information in the second reference voiceprint information;

in response to the quality of the voice information being lower than the quality of the second voice information, delete the reference voiceprint information; and

in response to the quality of the voice information being higher than the quality of the second voice information, delete the second reference voiceprint information.

9. An identity authentication method, comprising:

acquiring a plurality of historical voice files generated by a call between a first user and one or more second users;

filtering processing on the plurality of historical voice files to obtain voice information of the first user wherein the voice information includes a plurality of pieces of reference voiceprint information of the first user;

text recognition processing on the plurality of pieces of reference voiceprint information of the first user to obtain text information corresponding to the plurality of pieces of reference voiceprint information of the first user;

storing an identifier of the first user and the plurality of pieces of reference voiceprint information and the corresponding text information;

acquiring one of the plurality of pieces of reference voiceprint information corresponding to an identifier of a user to be authenticated;

outputting text information in the acquired reference voiceprint information, and receiving voice information to be authenticated;

comparing voice information in the acquired reference voiceprint information with the voice information to be authenticated;

in response to the voice information in the acquired reference voiceprint information matching with the voice information to be authenticated, determining that the authentication of the user succeeds; and

in response to the voice information in the acquired reference voiceprint information not matching with the voice information to be authenticated, determining that the authentication of the user fails;

wherein a randomly selected piece of the reference voiceprint information is sufficient for authenticating the first user.

10. The identity authentication method according to claim 9 , further comprising: segmenting the text information into a plurality of pieces of sub-text information; marking a start time and an end time of each piece of the sub-text information; and acquiring, according to the start time and the end time of the sub-text information, sub-voice information corresponding to each piece of the sub-text information from the voice information.

11. The identity authentication method according to claim 10 , wherein storing the plurality of pieces of reference voiceprint information and the corresponding text information comprises:

storing each pair of sub-voice information and sub-text information as a piece of reference voiceprint information of the first user.

12. The identity authentication method according to claim 9 , wherein storing the reference voiceprint information and the identifier of the first user comprises:

determining whether second reference voiceprint information exists in a voiceprint library, wherein the second reference voiceprint information includes second text information and second identifier, the second text information is the same as the text information in the reference voiceprint information, and the second identifier is the same as the identifier of the first user;

if the second reference voiceprint information exists in the voiceprint library, comparing a quality of the voice information in the reference voiceprint information with a quality of second voice information in the second reference voiceprint information;

if the quality of the voice information is lower than the quality of the second voice information, deleting the reference voiceprint information; and

if the quality of the voice information is higher than the quality of the second voice information, deleting the second reference voiceprint information.

13. An identity authentication system, comprising:

a voice filter configured to acquire a plurality of historical voice files generated by a plurality of calls between a first user and one or more second users, and execute filtering processing on the plurality of historical voice files to obtain voice information of the first user, wherein the voice information includes a plurality of pieces of reference voiceprint information of the first user;

a text recognizer configured to execute text recognition processing on the plurality of pieces of reference voiceprint information to obtain text information corresponding to the plurality of pieces of reference voiceprint information;

a voiceprint generator configured to store the plurality of pieces of reference voiceprint information and the corresponding text information, and store an identifier of the first user;

a voiceprint extractor configured to acquire one of the plurality of pieces of reference voiceprint information corresponding to an identifier of a user to be authenticated;

a user interface configured to output text information in the acquired reference voiceprint information, and receive voice information to be authenticated; and

a voiceprint matcher configured to compare voice information in the acquired reference voiceprint information with the voice information to be authenticated, the voiceprint matcher further configured to determine that the authentication of the user succeeds if the voice information in the acquired reference voiceprint information matches with the voice information to be authenticated, and determine that the authentication of the user fails if the voice information in the acquired reference voiceprint information does not match with the voice information to be authenticated;

wherein a randomly selected piece of the reference voiceprint information is sufficient for authenticating the first user.

14. The identity authentication system according to claim 13 , further comprising:

a text segmenter configured to segment the text information into a plurality of pieces of sub-text information, and mark a start time and an end time of each piece of the sub-text information; and

a voiceprint segmenter configured to acquire, according to the start time and the end time of the sub-text information, sub-voice information corresponding to each piece of the sub-text information from the voice information.

15. The identity authentication system according to claim 14 , wherein the voiceprint generator is further configured to store each pair of sub-voice information and sub-text information as a piece of reference voiceprint information of the first user.

16. The identity authentication system according to claim 13 , wherein the voiceprint generator is further configured to:

determine whether second reference voiceprint information exists in a voiceprint library, wherein the second reference voiceprint information includes second text information and second identifier, the second text information is the same as the text information in the reference voiceprint information, and the second identifier is the same as the identifier of the first user;

if the second reference voiceprint information exists in the voiceprint library, compare a quality of the voice information in the reference voiceprint information with a quality of second voice information in the second reference voiceprint information;

if the quality of the voice information is lower than the quality of the second voice information, delete the reference voiceprint information; and

if the quality of the voice information is higher than the quality of the second voice information, delete the second reference voiceprint information.

17. A non-transitory computer readable medium that stores a set of instructions that is executable by at least one processor of a system to cause the system to perform a method for generating voiceprint information, the method comprising:

acquiring a plurality of historical voice files generated by a plurality of calls between a first user and one or more second users;

executing filtering processing on the plurality of historical voice files to obtain voice information of the first user, wherein the voice information includes a plurality of pieces of reference voiceprint information of the first user;

executing text recognition processing on the plurality of pieces of reference voiceprint information to obtain text information corresponding to the plurality of pieces of reference voiceprint information; and

storing an identifier of the first user and the plurality of pieces of reference voiceprint information and the corresponding text information;

wherein a randomly selected piece of the reference voiceprint information is sufficient for authenticating the first user.

18. The non-transitory computer readable medium of claim 17 , wherein the set of instructions that is executable by the at least one processor of the system to cause the system to further perform:

segmenting the text information into a plurality of pieces of sub-text information;

marking a start time and an end time of each piece of the sub-text information; and

acquiring, according to the start time and the end time of the sub-text information, sub-voice information corresponding to each piece of the sub-text information from the voice information.

19. The non-transitory computer readable medium of claim 18 , wherein the set of instructions that is executable by the at least one processor of the system to cause the system to further perform:

storing each pair of sub-voice information and sub-text information as a piece of reference voiceprint information of the first user.

20. The non-transitory computer readable medium of claim 17 , wherein the set of instructions that is executable by the at least one processor of the system to cause the system to further perform:

determining whether second reference voiceprint information exists in a voiceprint library, wherein the second reference voiceprint information includes second text information and second identifier, the second text information is the same as the text information in the reference voiceprint information, and the second identifier is the same as the identifier of the first user;

in response to the second reference voiceprint information existing in the voiceprint library, comparing a quality of the voice information in the reference voiceprint information with a quality of the second voice information in the second reference voiceprint information;

in response to the quality of the voice information being lower than the quality of second voice information, deleting the reference voiceprint information; and

in response to the quality of the voice information is higher than the quality of the second voice information, deleting the second reference voiceprint information.

21. A non-transitory computer readable medium that stores a set of instructions that is executable by at least one processor of a system to cause the system to perform a method for identity authentication, the method comprising:

acquiring a plurality of historical voice files generated by a plurality of calls between a first user and one or more second users;

filtering processing on the plurality of historical voice files to obtain voice information of the first user, wherein the voice information includes a plurality of pieces of reference voiceprint information of the first user;

text recognition processing on the plurality of pieces of reference voiceprint information of the first user to obtain text information corresponding to the plurality of pieces of reference voiceprint information of the first user;

storing an identifier of the first user and the plurality of pieces of reference voiceprint information and the corresponding text information;

acquiring one of the plurality of pieces of reference voiceprint information corresponding to an identifier of a user to be authenticated;

outputting text information in the acquired reference voiceprint information, and receiving voice information to be authenticated;

comparing voice information in the acquired reference voiceprint information with the voice information to be authenticated;

in response to the voice information in the acquired reference voiceprint information matching with the voice information to be authenticated, determining that the authentication of the user succeeds; and

in response to the voice information in the acquired reference voiceprint information not matching with the voice information to be authenticated, determining that the authentication of the user fails;

wherein a randomly selected piece of the reference voiceprint information is sufficient for authenticating the first user.

22. The non-transitory computer readable medium of claim 21 , wherein the set of instructions that is executable by the at least one processor of the system to cause the system to further perform:

segmenting the text information into a plurality of pieces of sub-text information; marking a start time and an end time of each piece of the sub-text information; and acquiring, according to the start time and the end time of the sub-text information,

sub-voice information corresponding to each piece of the sub-text information from the voice information.

23. The non-transitory computer readable medium of claim 22 , wherein the set of instructions that is executable by the at least one processor of the system to cause the system to further perform:

storing each pair of sub-voice information and sub-text information as a piece of reference voiceprint information of the first user.

24. The non-transitory computer readable medium of claim 21 , wherein the set of instructions that is executable by the at least one processor of the system to cause the system to further perform:

determining whether second reference voiceprint information exists in a voiceprint library, wherein the second reference voiceprint information includes second text information and second identifier, the second text information is the same as the text information in the reference voiceprint information, and the second identifier is the same as the identifier of the first user;

if the second reference voiceprint information exists in the voiceprint library, comparing a quality of the voice information in the reference voiceprint information with a quality of second voice information in the second reference voiceprint information;

if the quality of the voice information is lower than the quality of the second voice information, deleting the reference voiceprint information; and

if the quality of the voice information is higher than the quality of the second voice information, deleting the second reference voiceprint information.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 13, 2020
From: ADVANTAGEOUS NEW TECHNOLOGIES CO., LTD.
To: ADVANCED NEW TECHNOLOGIES CO., LTD.
Reel/Frame 053761/0338 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 5, 2020
From: ALIBABA GROUP HOLDING LIMITED
To: ADVANTAGEOUS NEW TECHNOLOGIES CO., LTD.
Reel/Frame 053713/0665 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 17, 2019
From: XIONG, JIAN
To: ALIBABA GROUP HOLDING LIMITED
Reel/Frame 051310/0021 →
Priority Claims (1)
CN 2014 1 0532530 · Oct 10, 2014 · national
Continuity (2)
Continuation PCTCN2015091260 · Sep 30, 2015
Related Publication 20170221488A1 · Aug 3, 2017