IP Library › Granted Patent US 11,443,559
Granted Patent B2
US 11,443,559 · App. 17/006,256 · Granted Sep 13, 2022

Facial liveness detection with a mobile device

Inventors: Mikhail Vorobiev (Zurich, CH); Nevena Shamoska (Zürich, CH); Magdalena Polac (Zürich, CH); Benjamin Fankhauser (Biel-Bienne, CH); Michael Goettlicher (Biel-Bienne, CH); Marcus Hudritsch (Biel-Bienne, CH); Stamatios Georgoulis (Zurich, CH); Suman Saha (Zurich, CH); Luc van Gool (Zurich, CH)
Assignee: PXL Vision AG
G06V40/45B42D25/23B42D25/328G06K9/6256G06N3/08G06V10/25G06V10/751G06V30/413G06V40/168G06V40/172G06V40/67
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,443,559
App. No.
17/006,256
Granted
Sep 13, 2022
Kind
B2
Abstract

A system for remote identification of users. The system uses deep learning techniques for authenticating a user from an identification document, using automated verification of identification documents and detection that a live person identified by the document is present. Liveness of a user indicated by the identification document may be determined with a deep learning model trained for identification of facial spoofing attacks. The deep learning model may be trained using training data extracted from facial feature locations of training images.

Claims (47)

1. A non-transitory computer-readable medium comprising computer-executable instructions which, when executed by a computing device, cause the computing device to carry out a method, the method comprising:

using at least one processor to perform:

accessing a plurality of images comprising a face obtained by a camera;

providing the plurality of images to a trained deep learning model to obtain output indicating one or more likelihoods that the plurality of images comprise images of a live user and one or more likelihoods that the plurality of images comprise images of a spoof attack; and

identifying the plurality of images as comprising at least one of a live user and a spoof attack based on the output obtained from the trained deep learning model;

wherein the trained deep learning model comprises at least one convolutional neural network, the at least one convolutional neural network being trained using at least one feedback network configured to improve performance of the at least one convolutional neural network in classifying images from multiple domains at least in part by generating, during training, one or more domain classifier scores representing one or more probabilities of an input training data example belonging to one or more domains.

2. The non-transitory computer-readable medium of claim 1 , wherein the trained deep learning model is trained based on any one of:

training data comprising facial feature locations extracted from training images; and

feedback from the at least one feedback network.

3. The non-transitory computer-readable medium of claim 1 , wherein the trained deep learning model is configured to identify spoof attacks in images from the multiple domains, the multiple domains including pre-recorded videos comprising a face, still images comprising a face, and live users wearing a mask.

4. The non-transitory computer-readable medium of claim 1 , wherein during training, the at least one convolutional neural network is trained using the one or more domain classifier scores during backpropagation.

5. The non-transitory computer-readable medium of claim 1 , wherein the camera and the trained deep learning model are disposed in a user device.

6. The non-transitory computer-readable medium of claim 1 , wherein the at least one convolutional neural network comprises a residual network.

7. A computing system comprising:

a camera;

a server;

at least one processor; and

at least one non-transitory computer-readable medium comprising instructions which, when executed by the at least one processor, cause the computing system to perform a method of:

using the at least one processor to perform:

accessing, from the server, a plurality of images comprising a face obtained by the camera;

providing the plurality of images to a trained deep learning model to obtain output indicating one or more likelihoods that the plurality of images comprise images of a live user and one or more likelihoods that the plurality of images comprise images of a spoof attack; and

identifying the plurality of images as comprising at least one of a live user and a spoof attack based on the output obtained from the trained deep learning model,

wherein the trained deep learning model comprises at least one convolutional neural network, the at least one convolutional neural network being trained using at least one feedback network configured to improve performance of the at least one convolutional neural network in classifying images from multiple domains at least in part by generating, during training, one or more domain classifier scores representing one or more probabilities of an input training data example belonging to one or more domains.

8. The computing system of claim 7 , wherein the trained deep learning model is trained based on any one of:

training data comprising facial feature locations extracted from training images; and

feedback from the at least one feedback network.

9. The computing system of claim 7 , wherein the trained deep learning model is configured to identify spoof attacks in images from the multiple domains, the multiple domains including pre-recorded videos comprising a face, still images comprising a face, and live users wearing a mask.

10. The computing system of claim 9 , wherein the camera is disposed in a user device and the trained deep learning model is stored on computer-readable medium of the server.

11. The computing system of claim 9 , wherein the at least one convolutional neural network comprises a residual network.

12. A method of training a deep learning model for identifying facial spoofing, the method comprising:

using at least one computer hardware processor to perform:

accessing training data obtained by extracting facial feature locations from training images; and

training the deep learning model using the training data, wherein the deep learning model comprises at least one residual network, and wherein training the deep learning model comprises using at least one feedback network to improve performance of the at least one residual network in classifying images from multiple domains at least in part by generating, during training, one or more domain classifier scores representing one or more probabilities of an input training data example belonging to one or more domains.

13. The method of claim 12 , wherein the at least one residual network is configured to pass a feature vector to a label classifier network and a gradient reversal layer.

14. The method of claim 13 , wherein the label classifier network is configured to process the feature vector into a classification score vector and to classify whether the at least one image is an image of a live user or a spoof attack based on the classification score vector.

15. The method of claim 13 , wherein the gradient reversal layer is configured to optimize parameters of the at least one residual network during training.

16. The method of claim 13 , wherein the gradient reversal layer is configured to pass the feature vector to the at least one feedback network during forward propagation.

17. A method of training a deep learning model for identifying facial spoofing, the method comprising:

using at least one computer hardware processor to perform:

accessing training data obtained by extracting facial feature locations from training images; and

training the deep learning model using the training data, wherein:

the deep learning model comprises at least one residual network,

training the deep learning model comprises using at least one feedback network to improve performance of the at least one residual network in classifying images from multiple domains at least in part by generating, during training, one or more domain classifier scores representing one or more probabilities of an input training data example belonging to one or more domains, and

the at least one feedback network comprises a domain discriminator network configured to produce a domain classification loss vector, wherein during training:

the domain discriminator network minimizes the domain classification loss vector; and

the at least one residual network maximizes the domain classification loss vector until an optimization between the domain discriminator network and the at least one residual network is reached.

18. The method of claim 12 , wherein the deep learning model is trained to identify spoof attacks including pre-recorded videos comprising a face, still images comprising a face, and live users wearing a mask.

Assignments (6)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 19, 2020
From: VOROBIEV, MIKHAIL; SHAMOSKA, NEVENA; POLAC, MAGDALENA
To: PXL VISION AG
Reel/Frame 054418/0363 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 19, 2020
From: SAHA, SUMAN; GEORGOULIS, STAMATIOS; GOOL, LUC VAN
To: SWISS FEDERAL INSTITUTE OF TECHNOLOGY ZURICH (A.K.A. ETH ZURICH)
Reel/Frame 054418/0778 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 19, 2020
From: BERN UNIVERSITY OF APPLIED SCIENCES
To: PXL VISION AG
Reel/Frame 054418/0973 →
CHANGE OF ADDRESS Recorded Nov 19, 2020
From: PXL VISION AG
To: PXL VISION AG
Reel/Frame 054481/0597 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 19, 2020
From: FANKHAUSER, BENJAMIN; GOETTLICHER, MICHAEL; HUDRITSCH, MARCUS
To: BERN UNIVERSITY OF APPLIED SCIENCES
Reel/Frame 054481/0599 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 19, 2020
From: SWISS FEDERAL INSTITUTE OF TECHNOLOGY ZURICH (A.K.A. ETH ZURICH)
To: PXL VISION AG
Reel/Frame 054977/0631 →
Continuity (2)
Provisional Application 62893556 · Aug 29, 2019
Related Publication 20210064901A1 · Mar 4, 2021
Cited By (1)
US 12,511,942