IP Library › Granted Patent US 12,236,717
Granted Patent B2
US 12,236,717 · App. 18/325,928 · Granted Feb 25, 2025

Spoof detection based on challenge response analysis

Inventors: Spandana Vemulapalli (Kansas City, MO); Reza R. Derakhshani (Shawnee, KS)
Assignee: JUMIO CORPORATION
G06V40/45G06V40/172G06V40/174
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,236,717
App. No.
18/325,928
Granted
Feb 25, 2025
Kind
B2
Abstract

Methods, systems, and computer-readable storage media for determining that a subject is a live person include capturing a set of images of a subject instructed to perform a facial expression. A region of interest for the facial expression is determined in a first image of the set, the first image representing a first facial state that includes the facial expression. A set of facial features is identified in the region of interest, the facial features being indicative of interaction between facial muscles and skin of the subject due to the subject performing the facial expression. A determination is made, based on the facial features, that the first image substantially matches a template image of the facial expression of the subject. Responsive to determining that the first image substantially matches the template image, identifying the subject as a live person.

Claims (58)

1. A computer-implemented method comprising:

causing display of an animated image of an avatar performing a facial expression over a reference time period;

capturing a set of images of a subject as a response of the subject to the display of the animated image of the avatar performing the facial expression during a substantially similar time period to the reference time period, wherein the set of images of the subject includes multiple images depicting a transition in the facial expression over the substantially similar time period;

determining a set of points of interest for the facial expression in a first image of the set of images, the set of points of interest used to determine one or more expression features in the facial expression;

identifying the one or more expression features in the facial expression, the one or more expression features comprising one or more subject-specific features, the one or more expression features identified based on the set of points of interest for the facial expression;

determining, in a reference image of the subject performing a neutral expression, at least one subject-specific feature identified in the first image exists in the reference image using a deep learning model of a machine learning process;

determining that the subject in the first image substantially matches the subject in the reference image based, at least in part, on the at least one subject-specific feature identified in the first image exists in the reference image and a dynamic metric representing a quantification of continuous motion of the facial expression based on the multiple images depicting the transition from the neutral expression; and

in response to determining that the subject in the first image substantially matches the subject in the reference image, identifying the subject as a live person.

2. The computer-implemented method of claim 1 , wherein the set of points of interest for the facial expression in the first image are correlated to a set of points of interest for the neutral expression in the reference image to determine one or more temporal displacement parameters.

3. The computer-implemented method of claim 1 , wherein the reference image was previously captured in an enrollment process and stored in a database.

4. The computer-implemented method of claim 1 , wherein a dissimilarity score measures a dissimilarity of the first image in comparison to the reference image, the dissimilarity score calculated based on the one or more expression features in the set of points of interest, the method further comprising:

determining that the subject in the first image substantially matches the subject in the reference image based, at least in part, on the dissimilarity score exceeding a threshold.

5. The computer-implemented method of claim 1 , wherein the facial expression is one of: a smile, a scowl, a frown, or raising eyebrows.

6. The computer-implemented method of claim 1 , further comprising:

determining, from the captured set of images of the subject performing the facial expression, a continuous transition from a neutral phase to a facial expression phase based on tracking a temporal displacement of the set of points of interest used to determine the one or more expression features identified in the facial expression; and

determining that the subject in the first image substantially matches the subject in the reference image based, at least in part, on the continuous transition determined from the neutral phase to the facial expression phase.

7. The computer-implemented method of claim 1 , further comprising:

determining, from the captured set of images of the subject performing the facial expression, an abrupt transition from a neutral phase to a facial expression phase based on determining a displacement trajectory of at least one point of interest in the set of points of interest being different from a step function; and

determining that the subject in the first image does not substantially match the subject in the reference image based on the abrupt transition determined from the neutral phase to the facial expression phase.

8. A non-transitory, computer-readable medium storing one or more instructions executable by a computer system to perform operations comprising:

causing display of an animated image of an avatar performing a facial expression over a reference time period;

capturing a set of images of a subject as a response of the subject to the display of the animated image of the avatar performing the facial expression during a substantially similar time period to the reference time period, wherein the set of images of the subject includes multiple images depicting a transition in the facial expression over the substantially similar time period;

determining a set of points of interest for the facial expression in a first image of the set of images, the set of points of interest used to determine one or more expression features in the facial expression;

identifying the one or more expression features in the facial expression, the one or more expression features comprising one or more subject-specific features, the one or more expression features identified based on the set of points of interest for the facial expression;

determining, in a reference image of the subject performing a neutral expression, at least one subject-specific feature identified in the first image exists in the reference image using a deep learning model of a machine learning process;

determining that the subject in the first image substantially matches the subject in the reference image based, at least in part, on the at least one subject-specific feature identified in the first image exists in the reference image and a dynamic metric representing a quantification of continuous motion of the facial expression based on the multiple images depicting the transition from the neutral expression; and

in response to determining that the subject in the first image substantially matches the subject in the reference image, identifying the subject as a live person.

9. The non-transitory, computer-readable medium of claim 8 , wherein the set of points of interest for the facial expression in the first image are correlated to a set of points of interest for the neutral expression in the reference image to determine one or more temporal displacement parameters.

10. The non-transitory, computer-readable medium of claim 8 , wherein the reference image was previously captured in an enrollment process and stored in a database.

11. The non-transitory, computer-readable medium of claim 8 , wherein a dissimilarity score measures a dissimilarity of the first image in comparison to the reference image, the dissimilarity score calculated based on the one or more expression features in the set of points of interest, wherein the operations further comprise:

determining that the subject in the first image substantially matches the subject in the reference image based, at least in part, on the dissimilarity score exceeding a threshold.

12. The non-transitory, computer-readable medium of claim 8 , wherein the facial expression is one of: a smile, a scowl, a frown, or raising eyebrows.

13. The non-transitory, computer-readable medium of claim 8 , wherein the operations further comprise:

determining, from the captured set of images of the subject performing the facial expression, a continuous transition from a neutral phase to a facial expression phase based on tracking a temporal displacement of the set of points of interest used to determine the one or more expression features identified in the facial expression; and

determining that the subject in the first image substantially matches the subject in the reference image based, at least in part, on the continuous transition determined from the neutral phase to the facial expression phase.

14. The non-transitory, computer-readable medium of claim 8 , wherein the operations further comprise:

determining, from the captured set of images of the subject performing the facial expression, an abrupt transition from a neutral phase to a facial expression phase based on determining a displacement trajectory of at least one point of interest in the set of points of interest being different from a step function; and

determining that the subject in the first image does not substantially match the subject in the reference image based on the abrupt transition determined from the neutral phase to the facial expression phase.

15. A computer-implemented system, comprising:

one or more computers; and

one or more computer memory devices interoperably coupled with the one or more computers and having tangible, non-transitory, machine-readable media storing one or more instructions that, when executed by the one or more computers, perform operations comprising:

causing display of an animated image of an avatar performing a facial expression over a reference time period;

capturing a set of images of a subject as a response of the subject to the display of the animated image of the avatar performing the facial expression during a substantially similar time period to the reference time period, wherein the set of images of the subject includes multiple images depicting a transition in the facial expression over the substantially similar time period;

determining a set of points of interest for the facial expression in a first image of the set of images, the set of points of interest used to determine one or more expression features in the facial expression;

identifying the one or more expression features in the facial expression, the one or more expression features comprising one or more subject-specific features, the one or more expression features identified based on the set of points of interest for the facial expression;

determining, in a reference image of the subject performing a neutral expression, at least one subject-specific feature identified in the first image exists in the reference image using a deep learning model of a machine learning process;

determining that the subject in the first image substantially matches the subject in the reference image based, at least in part, on the at least one subject-specific feature identified in the first image exists in the reference image and a dynamic metric representing a quantification of continuous motion of the facial expression based on the multiple images depicting the transition from the neutral expression; and

in response to determining that the subject in the first image substantially matches the subject in the reference image, identifying the subject as a live person.

16. The system of claim 15 , wherein the set of points of interest for the facial expression in the first image are correlated to a set of points of interest for the neutral expression in the reference image to determine one or more temporal displacement parameters.

17. The system of claim 15 , wherein the reference image was previously captured in an enrollment process and stored in a database.

18. The system of claim 15 , wherein a dissimilarity score measures a dissimilarity of the first image in comparison to the reference image, the dissimilarity score calculated based on the one or more expression features in the set of points of interest, wherein the operations further comprise:

determining that the subject in the first image substantially matches the subject in the reference image based, at least in part, on the dissimilarity score exceeding a threshold.

19. The system of claim 15 , wherein the operations further comprise:

determining, from the captured set of images of the subject performing the facial expression, a continuous transition from a neutral phase to a facial expression phase based on tracking a temporal displacement of the set of points of interest used to determine the one or more expression features identified in the facial expression; and

determining that the subject in the first image substantially matches the subject in the reference image based, at least in part, on the continuous transition determined from the neutral phase to the facial expression phase.

20. The system of claim 15 , wherein the operations further comprise:

determining, from the captured set of images of the subject performing the facial expression, an abrupt transition from a neutral phase to a facial expression phase based on determining a displacement trajectory of at least one point of interest in the set of points of interest being different from a step function; and

determining that the subject in the first image does not substantially match the subject in the reference image based on the abrupt transition determined from the neutral phase to the facial expression phase.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE ERRONEOUSLY ENTERED RECEIVING PARTY DATA PREVIOUSLY RECORDED AT REEL: 063800 FRAME: 0479. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jun 8, 2023
From: VEMULAPALLI, SPANDANA; DERAKHSHANI, REZA R.
To: EYEVERIFY, INC.
Reel/Frame 063939/0820 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2023
From: VEMULAPALLI, SPANDANA; DERAKHSHANI, REZA R.
To: JUMIO CORPORATION
Reel/Frame 063800/0479 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2023
From: EYEVERIFY INC.
To: JUMIO CORPORATION
Reel/Frame 063800/0490 →
Continuity (2)
Continuation 17463199 · Aug 31, 2021
Related Publication 20230306792A1 · Sep 28, 2023
References Cited (14)
US 8457367B1 · Sipe · 2013 [cited by examiner]
US 10929984B2 · Zhang · 2021 [cited by examiner]
US 11710353B2 · Vemulapalli · 2023 [cited by examiner]
US 20130083052A1 · Dahlkvist · 2013 [cited by examiner]
US 20160019410A1 · Komogortsev · 2016 [cited by applicant]
US 20190318077A1 · Narasimhan · 2019 [cited by applicant]
US 20200057846A1 · Gautam · 2020 [cited by examiner]
US 20200342245A1 · Lubin · 2020 [cited by examiner]
US 20210327431A1 · Stewart · 2021 [cited by examiner]
US 20220207263A1 · Vexler et al. · 2022 [cited by applicant]
US 20230063229A1 · Vemulapalli · 2023 [cited by applicant]
WO 2023034251A1 · 2023 [cited by applicant]
PCT International Search Report and Written Opinion; Application No. PCT/US2022/041967 Jumio Corporation; International filing date of Aug. 30, 2022, date of mailing Dec. 29, 2022, 13 pages. [cited by applicant]
VAAS, sophos.com [online], “Google files patent to let you unlock your phone by grimacing at it,” Jun. 12, 2013, retrieved on Oct. 18, 2021, retrieved from URL https://nakedsecurity.sophos.com/2013/06/12/google-files-pa… [cited by applicant]