Digital verification of users based on real-time video stream
Techniques are disclosed relating to receiving, by a computer system from a user, a request to authorize the user. The technique may further include receiving, by the computer system, a live video stream of the user, and requesting, by the computer system, the user to submit an identification document from a list of authorized identification documents. While the user submits the identification document, the technique may also include detecting liveliness of the user in the live video stream by using a gesture recognition operation. Additionally, the technique may include determining, using the submitted identification document and the detected liveliness, whether to proceed with the request to authorize the user.
1 . A method comprising:
receiving, by a computer system from a user, a request to authorize the user;
receiving, by the computer system, a live video stream of the user that includes audio from the user;
requesting, by the computer system, the user to display an identification document via the live video stream;
detecting, by the computer system while the user displays the identification document, liveliness of the user in the live video stream by using a gesture recognition operation;
determining, based on the displayed identification document and the detected liveliness, whether to proceed with the request to authorize the user; and
based on a decision to proceed, performing an analysis model that includes:
analyzing images of the user captured from the live video stream to determine a probable emotional state of the user;
comparing the probable emotional state to answers provided by the user via the audio from the user;
identifying, based on the comparison, indications of deceit in the answers provided by the user; and
using the indications of deceit to determine whether to authorize the user.
2 . The method of claim 1 , further comprising, in response to determining to proceed with the request to authorize the user, performing one or more detection models, including face spoofing detection, face morphing detection, or facial marker detection.
3 . The method of claim 2 , further comprising:
prior to performing the analysis model, performing two or more of the detection models.
4 . The method of claim 2 , wherein the computer system is a particular local client server, and the one or more detection models are performed by a central online server that is in communication with a plurality of local client servers; and
wherein parameters learned by the particular local client server, excluding personal data, are passed to the central online server.
5 . The method of claim 1 , further comprising:
receiving, by the computer system from the user, a second image of a hardcopy of the identification document prior to the request to authorize the user; and
comparing a first image of the identification document taken from the live video stream to the second image of the identification document.
6 . The method of claim 5 , wherein the comparing includes:
performing an optical character recognition operation on the first image to generate a first text version of the identification document; and
comparing the first text version of the identification document to a second text version of the second image of the identification document.
7 . The method of claim 5 , further comprising:
requesting, by the computer system, the user to submit an address verification document from a list of authorized address verification documents;
receiving, by the computer system from the user, an image of a hardcopy of the address verification document; and
comparing the received image of the address verification document to a previously stored version of the address verification document.
8 . The method of claim 1 , further comprising storing, by the computer system after determining whether to proceed, information associated with a first image of the displayed identification document and the detected liveliness by adding a new block to a blockchain.
9 . The method of claim 1 , wherein detecting the liveliness includes aligning motions of the user to input received by the computer system via a user interface.
10 . A computer-readable, non-transient memory including instructions that when executed by a computer system, cause the computer system to perform operations including:
determining that a request from a user on a user device requires authentication of the user to proceed;
receiving, from a camera coupled to the user device, a live video stream of the user that includes audio from the user;
requesting the user to display, in front of the camera during the live video stream, an identification document from a list of authorized identification documents;
monitoring, while the user displays the identification document, liveliness of the user in the live video stream, wherein the monitoring includes using a gesture recognition operation;
determining, based on the displayed identification document and recognized gestures, whether to proceed with the authentication of the user; and
based on a decision to proceed, performing an analysis model that includes:
analyzing images of facial features and movements of the user captured from the live video stream to determine a probable emotional state of the user;
correlating the facial features and movements of the user to audio of questions and answers between the user and an interviewer participating in the authentication of the user;
identifying, based on the correlation, indications of deceit in the answers provided by the user; and
using the indications of deceit to determine whether to authorize the user.
11 . The computer-readable, non-transient memory of claim 10 , wherein the operations further include, in response to determining to proceed with the authentication of the user, performing a face spoofing detection model, including:
identifying one or more particular facial features from a plurality of images of the live video stream of the user;
assigning respective activation vectors to the one or more particular facial features; and
determining whether the respective activation vectors are consistent with movement of a real face.
12 . The computer-readable, non-transient memory of claim 10 , wherein the operations further include, in response to determining to proceed with the authentication of the user, performing a face morphing detection model, including:
generating a respective set of outputs from each of twin neural networks using one or more images from the live video stream of the user; and
measuring contrastive loss between the respective sets of outputs to determine whether a face of the user is real.
13 . The computer-readable, non-transient memory of claim 10 , wherein the operations further include, in response to determining to proceed with the authentication of the user, performing a facial marker detection model, including:
identifying one or more particular facial markers from one or more images of the live video stream of the user;
identifying a contour of the user's face in the one or more images and mapping the one or more particular facial markers to the contour; and
determining whether the user's face is real based on the mapping.
14 . The computer-readable, non-transient memory of claim 10 , wherein the operations further include monitoring liveliness using a facial depth detection model that includes a first convolutional neural network (CNN) that is operable to:
extract one or more local face blocks from one or more images of the user's face; and
assign a corresponding score to respective ones of the local face blocks, each score indicating a likelihood that the user's face is real.
15 . The computer-readable, non-transient memory of claim 14 , wherein the facial depth detection model includes a second CNN that is operable to estimate a depth map of the user's face using the one or more images of the user's face.
16 . A system comprising:
a processor circuit;
a camera circuit; and
a memory circuit including instructions that when executed by the processor circuit, cause the system to perform operations including:
receiving, from a user, a request to access restricted information associated with the user;
in response to receiving an indication of user approval, generating, using the camera circuit, a live video stream of the user that includes audio from the user;
requesting, during the live video stream, the user to display an identification document from a list of authorized identification documents;
detecting, while the user displays the identification document, liveliness of the user in the live video stream by using a gesture recognition operation;
determining, based on the displayed identification document and the detected liveliness, whether to proceed with authorization of the user; and
based on a decision to proceed, performing an analysis model to determine an indication of the user's honesty, the analysis model including:
correlating images of facial features and movements of the user to audio of answers provided by the user;
identifying, based on the correlation, indications of deceit in the answers provided by the user; and
using the indications of deceit to determine whether to authorize the user.
17 . The system of claim 16 , wherein the operations further include detecting the liveliness by:
requesting the user to perform a particular action in view of the camera circuit; and
determine, using the live video stream, whether the user performed the particular action.
18 . The system of claim 16 , wherein the operations further include detecting, using the live video stream, the liveliness by detecting micro-motions in the user's face while the user displays the identification document.
19 . The system of claim 16 , wherein the operations further include detecting the liveliness using a facial depth detection model that includes:
a first convolutional neural network (CNN) that is operable to:
extract one or more local face blocks from one or more images of the user's face; and
assign a corresponding score to respective ones of the local face blocks, each score indicating a likelihood that the user's face is real; and
a second CNN that is operable to estimate a depth map of the user's face using the one or more images of the user's face.
20 . The system of claim 19 , wherein the operations further include:
receiving, from a federated server computer, updates for the first and second CNNs; and
sending, to the federated server computer, parameters from usage of the first and second CNNs with the live video stream, wherein the parameters exclude personal information of the user.