Systems and methods for providing mobility-impaired transactional or identity documents processing
Disclosed embodiments may include a method for systems and methods for providing mobility-impaired transactional or identify documents processing. The method may include receiving login credentials associated with an account and a request to deposit a check into the account, and identifying, using an image recognition model, a presence of a check in a visual field of the image capture device. Then the method may include obtaining, via the image capture device, a plurality of digital images of a front side of the check, and generating, based on the plurality of digital images of the front side of the check, a composite image of the front side of the check. The method may further include creating a recording of an endorsement gesture and transmitting the composite image of the front side of the check and the recording of the endorsement gesture to a back-end server.
1 . A system comprising:
an image capture device;
one or more processors; and
a memory in communication with the one or more processors and storing instructions that, when executed by the one or more processors, are configured to cause the system to:
identify, using an image recognition model, a presence of a document in a visual field of the image capture device;
obtain, via the image capture device, a plurality of digital images of a first side of the document;
generate, based on the plurality of digital images of the first side of the document using a machine learning model, a composite image of the first side of the document by juxtaposing portions of the plurality of digital images together; and
transmit the composite image of the first side of the document to a back-end server for storage.
2 . The system of claim 1 , wherein the system comprises an augmented reality device.
3 . The system of claim 1 , wherein generating the composite image of the first side of the document comprises:
analyzing, using optical character recognition (OCR) techniques, the plurality of digital images to identify a first set of document information;
transmitting the first set of document information to a back-end system;
receiving, from the back-end system, a second set of document information, wherein the back-end system identified the second set of document information based on the first set of document information; and
overlaying the second set of document information onto a digital image that comprises the first set of document information.
4 . The system of claim 3 , wherein the first set of document information comprises a name and a first portion of an account number and the second set of document information comprises a second portion of the account number, wherein when juxtaposed, the first portion of the account number and the second portion of the account number form a complete account number.
5 . The system of claim 3 , wherein the instructions are further configured to cause the system to:
determine that the first set of document information is insufficient to determine the second set of document information;
generate a prompt to a user to move such that the image capture device can obtain a new set of digital images of the document from a different vantage point;
analyze, using OCR techniques, the new set of digital images to obtain new document information; and
responsive to adding the new document information to the first set of document information, determine that the first set of document information is sufficient to determine the second set of document information.
6 . The system of claim 1 , wherein the composite image of the first side of the document comprises:
a routing number;
an account number;
a signature of a payor of the document;
an amount;
a date;
a document number; or
combinations thereof.
7 . The system of claim 1 , wherein the machine learning model has been trained using training data comprising a plurality of sets of digital document images and corresponding sets of resultant composite images.
8 . The system of claim 1 , wherein the instructions are further configured to cause the system to:
create a recording of an endorsement gesture by recording a gesture performed by a user that signifies an endorsement of the document; and
transmit the recording of the endorsement gesture to the back-end server for storage in association with a record of the document,
wherein recording the gesture performed by the user that signifies an endorsement of the document comprises one or more of:
recording, via the image capture device, a hand gesture of the user;
recording, via the image capture device, one or more taps performed by the user with a finger;
recording, via an audio recording device, an audio signature performed by the user; and
recording, via the image capture device, an eye movement of the user.
9 . The system of claim 8 , wherein the instructions are further configured to cause the system to:
generate, based on the recording of the endorsement gesture, a composite image of a second side of the document, wherein the composite image of the second side of the document comprises an indication that the document has been endorsed; and
transmit the composite image of the second side of the document to the back-end server for storage.
10 . The system of claim 1 , wherein the instructions are further configured to cause the system to:
receive, via a voice recording, a request to process the document.
11 . A system comprising:
an image capture device;
one or more processors; and
a memory in communication with the one or more processors and storing instructions that, when executed by the one or more processors, are configured to cause the system to:
identify, using an image recognition model, a presence of a document in a visual field of the image capture device;
obtain, via the image capture device, a plurality of digital images of a first side of a document;
generate, based on the plurality of digital images of the first side of the document, a composite image of the first side of the document;
obtain, via the image capture device, a second plurality of digital images of a second side of the document;
generate, based on the second plurality of digital images of the second side of the document, a composite image of the second side of the document; and
transmit the composite image of the first side of the document and the composite image of the second side of the document to a back-end server for storage.
12 . The system of claim 11 , wherein the instructions are further configured to cause the system to:
create a recording of an endorsement gesture by recording a gesture performed by a user that signifies an endorsement of the document;
determine, based on one or more of the plurality of digital images of the first side of the document and using a machine learning model, whether the document has already been endorsed; and
transmit the recording of the endorsement gesture to the back-end server for storage in association with a record of the document,
wherein the machine learning model has been trained with training data that comprises training images of first sides of a plurality of test documents and a corresponding classification for each of the plurality of test documents, wherein the corresponding classification provides an indication of whether the test document has already been endorsed.
13 . The system of claim 12 , wherein determining whether the document has already been endorsed comprises the machine learning model receiving as inputs, a portion of one or more of the plurality of digital images, wherein the portion corresponds to an area of the first side of the document that is designated for providing a signature.
14 . The system of claim 12 , wherein the instructions are further configured to cause the system to:
responsive to the machine learning model determining that the document has already been endorsed:
generate a prompt to the user to move such that the visual field of the image capture device changes; and
identify, using the image recognition model, a presence of another document in the visual field of the image capture device.
15 . The system of claim 11 , wherein the instructions are further configured to cause the system to:
scan an environment to detect the presence of the document prior to receiving a request to process the document; and
transmit a notification to a user to confirm whether to process the document.
16 . A system comprising:
one or more processors; and
a memory in communication with the one or more processors and storing instructions that, when executed by the one or more processors, are configured to cause the system to:
receive a plurality of digital images of a first side of a document, wherein the plurality of digital images of the first side of the document are obtained by an image capture device of an augmented reality (AR) device;
generate, based on the plurality of digital images of the first side of the document, a composite image of the first side of the document;
analyze the composite image to extract a first identification number;
verify the first identification number associated with the document by comparing a second identification number associated with a user to portions of the first identification number from the plurality of digital images; and
store the composite image of the first side of the document such that the user may remotely access the composite image of the first side of the document.
17 . The system of claim 16 , wherein generating the composite image of the first side of the document comprises digitally stitching portions of the plurality of digital images together.
18 . The system of claim 16 , wherein generating the composite image of the first side of the document comprises:
analyzing, using optical character recognition (OCR) techniques, the plurality of digital images to identify a first set of document information;
accessing account information based on the first set of document information to determine a second set of document information; and
overlaying the second set of document information onto a digital image that comprises the first set of document information.
19 . The system of claim 18 , wherein the instructions are further configured to cause the system to:
determine that the first set of document information is insufficient to determine the second set of document information;
transmit a request to the AR device to prompt the user to move such that the image capture device of the AR device can obtain a new set of digital images of the document from a different vantage point;
analyze, using OCR techniques, the new set of digital images to obtain new document information; and
responsive to adding the new document information to the first set of document information, determine that the first set of document information is sufficient to determine the second set of document information.
20 . The system of claim 16 , wherein the instructions are further configured to cause the system to:
create a recording of an endorsement gesture by recording a gesture performed by a user that signifies an endorsement of the document,
wherein the recording of the endorsement gesture comprises one or more of:
a digital video recording of a hand gesture of the user;
a digital video recording of one or more taps performed by the user with a finger;
an audio recording of an audio signature performed by the user; and
a digital video recording of an eye movement of the user.