IP Library Granted Patent US 8,897,500
Granted Patent B2
US 8,897,500 · App. 13/101,704 · Granted Nov 25, 2014

System and method for dynamic facial features for speaker recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,897,500
App. No.
13/101,704
Granted
Nov 25, 2014
Kind
B2
Abstract

Disclosed herein are systems, methods, and non-transitory computer-readable storage media for performing speaker verification. A system configured to practice the method receives a request to verify a speaker, generates a text challenge that is unique to the request, and, in response to the request, prompts the speaker to utter the text challenge. Then the system records a dynamic image feature of the speaker as the speaker utters the text challenge, and performs speaker verification based on the dynamic image feature and the text challenge. Recording the dynamic image feature of the speaker can include recording video of the speaker while speaking the text challenge. The dynamic feature can include a movement pattern of head, lips, mouth, eyes, and/or eyebrows of the speaker. The dynamic image feature can relate to phonetic content of the speaker speaking the challenge, speech prosody, and the speaker's facial expression responding to content of the challenge.

Claims (42)

1. A method comprising:

receiving a request from a speaker to confirm a user identity;

retrieving a user profile associated with the user identity;

generating, based on the user profile and a level of security desired, a text challenge, wherein generating the text challenge is based on eliciting distinctive behavior of the speaker when compared to behavior of other users, the behavior of other users and the highly distinctive behavior of the speaker being stored in a database of dynamic image features;

prompting the speaker to utter the text challenge;

recording an audio recording and a video recording of the speaker as the speaker utters the text challenge; and

performing speaker verification using the audio recording, the video recording, and the user profile.

2. The method of claim 1 , wherein the recording of the audio recording and the video recording yields a dynamic image feature.

3. The method of claim 2 , wherein the dynamic image feature comprises a pattern of movement.

4. The method of claim 3 , wherein the pattern of movement is based on one of head, lips, mouth, eyes, and eyebrows.

5. The method of claim 2 , wherein the dynamic image feature relates to one of phonetic content of the speaker speaking the text challenge, speech prosody, a facial expression of the speaker in response to content of the text challenge, and a non-facial physically manifested response.

6. The method of claim 1 , wherein performing speaker verification is based on a database of speaker behaviors.

7. The method of claim 1 , wherein performing speaker verification is further based on a location of the speaker.

8. A system comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

receiving a request from a speaker to confirm a user identity;

retrieving a user profile associated with the user identity;

generating, based on the user profile and a level of security desired, a text challenge, wherein generating the text challenge is based on eliciting distinctive behavior of the speaker when compared to behavior of other users, the behavior of other users and the highly distinctive behavior of the speaker being stored in a database of dynamic image features;

prompting the speaker to utter the text challenge;

recording an audio recording and a video recording of the speaker as the speaker utters the text challenge; and

performing speaker verification using the audio recording, the video recording, and the user profile.

9. The system of claim 8 , wherein the speaker verification further comprises ensuring that the audio and the video match.

10. The system of claim 8 , wherein the speaker verification further comprises:

identifying features of the speaker in the video recording;

analyzing the features; and

temporally aligning the features to the audio recording based on the text challenge.

11. The system of claim 10 , wherein the features comprise one of a degree of a mouth opening, a symmetry of the mouth opening, a lip rounding, a lip spreading, a visible tongue position, a head movement, an eyebrow movement, and an eye shape.

12. The system of claim 10 , wherein the features comprise a facial expression of the speaker in response to the text challenge.

13. A computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

receiving a request to confirm a user identity;

retrieving a user profile associated with the user identity;

generating, based on the user profile and a level of security desired, a text challenge, wherein generating the text challenge is based on eliciting distinctive behavior of the speaker when compared to behavior of other users, the behavior of other users and the highly distinctive behavior of the speaker being stored in a database of dynamic image features;

prompting the speaker to utter the text challenge;

recording an audio recording and a video recording of the speaker as the speaker utters the text challenge; and

performing an analysis of the audio recording and the video recording based on the user profile.

14. The computer-readable storage device of claim 13 , wherein the user profile is generated as part of a user enrollment process.

15. The computer-readable storage device of claim 13 , wherein the user verification device uses the confirmation as part of a multi-factor authentication of the user.

16. The computer-readable storage device of claim 13 , further comprising:

receiving from the user verification device an indication of desired user verification certainty; and

setting the verification threshold based on the desired user verification certainty.

17. The computer-readable storage device of claim 13 , wherein performing the analysis further comprises temporally aligning the audio recording and the video recording, and determining whether the audio recording and the video recording match.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065530/0871 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 13, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065238/0239 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041504/0952 →