IP Library Granted Patent US 10,950,237
Granted Patent B2
US 10,950,237 · App. 14/953,984 · Granted Mar 16, 2021

System and method for dynamic facial features for speaker recognition

Inventors: Ann K. Syrdal (San Jose, CA); Sumit Chopra (Jersey City, NJ); Patrick Haffner (Atlantic Highlands, NJ); Taniya Mishra (New York, NY); Ilija Zeljkovic (Scotch Plains, NJ); Eric Zavesky (Austin, TX)
Assignee: Nuance Communications, Inc.
G10L15/25G06F21/32G06K9/00255G06K9/00281G06K9/00288G06K9/00315G06K9/00335G10L17/24G10L21/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,950,237
App. No.
14/953,984
Granted
Mar 16, 2021
Kind
B2
Abstract

Disclosed herein are systems, methods, and non-transitory computer-readable storage media for performing speaker verification. A system configured to practice the method receives a request to verify a speaker, generates a text challenge that is unique to the request, and, in response to the request, prompts the speaker to utter the text challenge. Then the system records a dynamic image feature of the speaker as the speaker utters the text challenge, and performs speaker verification based on the dynamic image feature and the text challenge. Recording the dynamic image feature of the speaker can include recording video of the speaker while speaking the text challenge. The dynamic feature can include a movement pattern of head, lips, mouth, eyes, and/or eyebrows of the speaker. The dynamic image feature can relate to phonetic content of the speaker speaking the challenge, speech prosody, and the speaker's facial expression responding to content of the challenge.

Claims (36)

1. A method comprising:

generating, by a processor associated with a device, a text challenge, wherein the text challenge is designed to elicit distinctive facial motion behavior on an individual basis for a user when speaking the text challenge when compared to facial motion behavior of other users when speaking the text challenge;

prompting, by the processor, the user to utter the text challenge;

recording, via an electronic image recorder associated with the device, video of the user as the user utters the text challenge, to capture the distinctive facial motion behavior that is associated with the text challenge;

performing, by the processor, user verification using the video; and

when the user verification confirms an identity of the user, authorizing access to a local resource or network resource.

2. The method of claim 1 , wherein the user verification is performed by comparing an image to the distinctive facial motion behavior of the user.

3. The method of claim 1 , wherein the generating of the text challenge and the performing of user verification are further based on a user profile.

4. The method of claim 1 , wherein the recording of the video further comprises recording audio and video of the user to yield a dynamic image feature.

5. The method of claim 4 , wherein the dynamic image feature comprises a pattern of movement.

6. The method of claim 5 , wherein the pattern of movement is based on one of head, lips, mouth, eyes, and eyebrows.

7. The method of claim 4 , wherein the dynamic image feature relates to one of phonetic content of the user speaking the text challenge, speech prosody, a facial expression of the user in response to content of the text challenge, and a non-facial physically manifested response.

8. The method of claim 1 , wherein the performing of user verification is further based on a database of user behaviors.

9. The method of claim 1 , wherein the performing of user verification is further based on a location of the user.

10. A system comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

generating a text challenge, wherein the text challenge is designed to elicit distinctive facial motion behavior on an individual basis for a user when speaking the text challenge when compared to facial motion behavior of other users when speaking the text challenge;

prompting the user to utter the text challenge;

recording, using an image recorder associated with the system, video of the user as the user utters the text challenge, to capture the distinctive facial motion behavior that is associated with the text challenge; performing user verification using the video; and

when the user verification confirms an identity of the user, authorizing access to a local resource or network resource.

11. The system of claim 10 , wherein the user verification is performed by comparing an image to the distinctive facial motion behavior of the user.

12. The system of claim 10 , wherein the generating of the text challenge and the performing of user verification are further based on a user profile.

13. The system of claim 10 , wherein the recording of the video further comprises recording audio and video to yield a dynamic image feature.

14. The system of claim 13 , wherein the dynamic image feature comprises a pattern of movement.

15. The system of claim 14 , wherein the pattern of movement is based on one of head, lips, mouth, eyes, and eyebrows.

16. The system of claim 13 , wherein the dynamic image feature relates to one of phonetic content of the user speaking the text challenge, speech prosody, a facial expression of the user in response to content of the text challenge, and a non-facial physically manifested response.

17. The system of claim 10 , wherein the performing of user verification is further based on a database of user behaviors.

18. The system of claim 10 , wherein the performing of user verification is further based on a location of the user.

19. A computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

generating a text challenge, wherein the text challenge is designed to elicit distinctive facial motion behavior on an individual basis for a user when speaking the text challenge when compared to facial motion behavior of other users when speaking the text challenge;

prompting the user to utter the text challenge;

recording, using an image recorder, video of the user as the user utters the text challenge, to capture the distinctive facial motion behavior that is associated with the text challenge;

performing user verification using the video; and

when the user verification confirms an identity of the user, authorizing access to a local resource or network resource.

20. The computer-readable storage device of claim 19 , wherein the user verification is performed by comparing an image to the distinctive facial motion behavior of the user.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065530/0871 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065238/0417 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041504/0952 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 29, 2015
From: SYRDAL, ANN K.; CHOPRA, SUMIT; HAFFNER, PATRICK; MISHRA, TANIYA; ZELJKOVIC, ILIJA; ZAVESKY, ERIC
To: AT&T INTELLECTUAL PROPERTY I, L.P.
Reel/Frame 037396/0698 →
Continuity (3)
Continuation 14551907 · Nov 24, 2014
Continuation 13101704 · May 5, 2011
Related Publication 20160078869A1 · Mar 17, 2016