IP Library Granted Patent US 10,789,343
Granted Patent B2
US 10,789,343 · App. 16/192,401 · Granted Sep 29, 2020

Identity authentication method and apparatus

Inventors: Peng Li (Hangzhou, CN); Yipeng Sun (Hangzhou, CN); Yongxiang Xie (Hangzhou, CN); Liang Li (Hangzhou, CN)
Assignee: Alibaba Group Holding Limited
G06F21/32G06K9/00G06K9/00288G10L17/005G10L17/06H04L29/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,789,343
App. No.
16/192,401
Granted
Sep 29, 2020
Kind
B2
Abstract

An audio/video stream generated by a target object to be authenticated is obtained. The target object is associated with a user. A determination is made whether a lip reading component and voice component in the audio/video stream are consistent. In response to determining that the lip reading component and voice component are consistent, voice recognition is performed on an audio stream in the audio/video stream to obtain voice content. The voice content is used as an object identifier of the target object. A model physiological feature corresponding to the object identifier is obtained from object registration information. Physiological recognition is performed on the audio/video stream to obtain a physiological feature of the target object. The physiological feature of the target object is compared with the model physiological feature to obtain a comparison result. If the comparison result satisfies an authentication condition, the target object is authenticated.

Claims (52)

1. A computer-implemented method, comprising:

obtaining an audio/video stream of a user that is to be authenticated;

determining that the user's voice in the audio/video stream matches the user's lips in the audio/video stream;

in response to determining that the user's voice in the audio/video stream matches the user's lips in the audio/video stream, determining, based on performing automated speech recognition on the audio/video stream, a user identifier for the user;

determining, based on performing automated physiological feature extraction on the audio/video stream, a physiological feature of the user;

obtaining, from stored, object registration information, a stored, model physiological feature corresponding to the determined, user identifier;

generating a comparison result based on comparing the physiological feature of the target object that was determined based on performing automated physiological feature extraction on the audio/video stream with the stored, model physiological feature; and

in response to determining that the comparison result satisfies an authentication condition, determining that the user has been authenticated.

2. The method of claim 1 , wherein the physiological feature of the user comprises a facial feature of the user.

3. The method of claim 1 , wherein the comparison result comprises a similarity score.

4. The method of claim 1 , wherein determining that the comparison result satisfies an authentication condition comprises determining that a similarity score exceeds a threshold score.

5. The method of claim 1 , wherein determining that the user's voice in the audio/video stream matches the user's lips comprises:

determining a lip reading syllable in a video image of the audio/video stream at a particular point in time;

determining a voice syllable in audio of the audio/video stream at the particular point in time; and

determining that the lip reading syllable and the voice syllable match.

6. The method of claim 1 , comprising storing the model physiological feature in the object registration information.

7. The method of claim 1 , comprising receiving a request from the user to authenticate, wherein the audio/video stream of the user is obtained in response to the request.

8. A non-transitory, computer-readable medium storing one or more instructions executable by a computer system to perform operations comprising:

obtaining an audio/video stream of a user that is to be authenticated;

determining that the user's voice in the audio/video stream matches the user's lips in the audio/video stream;

in response to determining that the user's voice in the audio/video stream matchesthe user's lips in the audio/video stream, determining, based on performing automated speech recognition on the audio/video stream, a user identifier for the user;

determining, based on performing automated physiological feature extraction on the audio/video stream, a physiological feature of the user;

obtaining, from stored, object registration information, a stored, model physiological feature corresponding to the determined, user identifier;

generating a comparison result based on comparing the physiological feature of the target object that was determined based on performing automated physiological feature extraction on the audio/video stream with the stored, model physiological feature; and

in response to determining that the comparison result satisfies an authentication condition, determining that the user has been authenticated.

9. The medium of claim 8 , wherein the physiological feature of the user comprises a facial feature of the user.

10. The medium of claim 8 , wherein the comparison result comprises a similarity score.

11. The medium of claim 8 , wherein determining that the comparison result satisfies an authentication condition comprises determining that a similarity score exceeds a threshold score.

12. The medium of claim 8 , wherein determining that the user's voice in the audio/video stream matches the user's lips comprises:

determining a lip reading syllable in a video image of the audio/video stream at a particular point in time;

determining a voice syllable in audio of the audio/video stream at the particular point in time; and

determining that the lip reading syllable and the voice syllable match.

13. The medium of claim 8 , wherein the operations comprise storing the model physiological feature in the object registration information.

14. The medium of claim 8 , wherein the operations comprise receiving a request from the user to authenticate, wherein the audio/video stream of the user is obtained in response to the request.

15. A computer-implemented system, comprising:

one or more computers; and

one or more computer memory devices interoperably coupled with the one or more computers and having tangible, non-transitory, machine-readable media storing one or more instructions that, when executed by the one or more computers, perform one or more operations comprising:

obtaining an audio/video stream of a user that is to be authenticated;

determining that the user's voice in the audio/video stream matches the user's lips in the audio/video stream;

in response to determining that the user's voice in the audio/video stream matches the user's lips in the audio/video stream, determining, based on performing automated speech recognition on the audio/video stream, a user identifier for the user;

determining, based on performing automated physiological feature extraction on the audio/video stream, a physiological feature of the user;

obtaining, from stored, object registration information, a stored, model physiological feature corresponding to the determined, user identifier;

generating a comparison result based on comparing the physiological feature of the target object that was determined based on performing automated physiological feature extraction on the audio/video stream with the stored, model physiological feature; and

in response to determining that the comparison result satisfies an authentication condition, determining that the user has been authenticated.

16. The system of claim 15 , wherein the physiological feature of the user comprises a facial feature of the user.

17. The system of claim 15 , wherein the comparison result comprises a similarity score.

18. The system of claim 15 , wherein determining that the comparison result satisfies an authentication condition comprises determining that a similarity score exceeds a threshold score.

19. The system of claim 15 , wherein determining that the user's voice in the audio/video stream matches the user's lips comprises:

determining a lip reading syllable in a video image of the audio/video stream at a particular point in time;

determining a voice syllable in audio of the audio/video stream at the particular point in time; and

determining that the lip reading syllable and the voice syllable match.

20. The system of claim 15 , wherein the operations comprise storing the model physiological feature in the object registration information.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 10, 2020
From: ADVANTAGEOUS NEW TECHNOLOGIES CO., LTD.
To: ADVANCED NEW TECHNOLOGIES CO., LTD.
Reel/Frame 053754/0625 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 31, 2020
From: ALIBABA GROUP HOLDING LIMITED
To: ADVANTAGEOUS NEW TECHNOLOGIES CO., LTD.
Reel/Frame 053743/0464 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2019
From: LI, PENG; SUN, YIPENG; XIE, YONGXIANG; LI, LIANG
To: ALIBABA GROUP HOLDING LIMITED
Reel/Frame 048213/0624 →