IP Library › Granted Patent US 12,444,224
Granted Patent B2
US 12,444,224 · App. 17/841,571 · Granted Oct 14, 2025

Machine learning for data extraction

Inventors: Pierre-Éric Michaël Alix Melchy (Montreal, CA); Labhesh Patel (Palo Alto, CA); Philipp Pointner (Vienna, AT); Radu Rogojanu (Vienna, AT); Antoine Bon (Vienna, AT)
Assignee: Jumio Corporation
G06V30/413G06V10/82G06V30/18057
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,444,224
App. No.
17/841,571
Filed
Jun 15, 2022
Granted
Oct 14, 2025
Kind
B2
Art Unit
2668
USPC
382/138
Abstract

Computer systems and methods are provided for extracting information from an image of a document. A computer system receives image data, the image data including an image of a document. The computer system determines a portion of the received image data that corresponds to a predefined document field. The computer system utilizes a neural network system to assign a label to the determined portion of the received image data. The computer system performs text recognition on the portion of the received image data and stores the recognized text in association with the assigned label.

Claims (97)

1. A computer-implemented method, comprising:

at a server system including one or more processors and memory storing one or more programs for execution by the one or more processors:

receiving image data, the image data including an image of a document, wherein the image of the document includes facial image data;

determining a portion of the received image data that corresponds to a predefined document field;

determining a position for the image of the document, by identifying one or more facial features within the facial image data;

utilizing a neural network system to assign a label to the determined portion of the received image data;

performing text recognition on the portion of the received image data; and

storing recognized text in association with the assigned label.

2. The method of claim 1 , further comprising:

after receiving the image data, determining a document type corresponding to the image of the document, the document type including document characteristics for the document type.

3. The method of claim 2 , further comprising:

comparing the portion of the received image data with the document type to determine respective document characteristics for the portion of the received image data.

4. The method of claim 3 , further comprising:

generating, based on the respective document characteristics for the portion of the received image data, sanitized document information, wherein the sanitized document information stores the document characteristics corresponding to the document type with a predetermined format.

5. The method of claim 2 , wherein the document characteristics corresponding to the document type include a predetermined layout for the document type, wherein the predetermined layout includes a landscape layout or a portrait layout.

6. The method of claim 2 , wherein the document characteristics corresponding to the document type include at least one or more selected from the group consisting of: one or more anchors, date format, or text format.

7. The method of claim 1 , wherein the predefined document field corresponds to at least one of a name, a location, a date, a document type, or a document number.

8. A computer-implemented method, further comprising:

at a server system including one or more processors and memory storing one or more programs for execution by the one or more processors:

receiving image data, the image data including an image of a document;

determining a portion of the received image data that corresponds to a predefined document field;

utilizing a neural network system to assign a label to the determined portion of the received image data;

performing text recognition on the portion of the received image data;

storing recognized text in association with the assigned label;

before performing the text recognition on the portion of the received image data, determining a position of the image of the document, wherein the image of the document includes facial image data;

determining whether the position of the image of the document meets orientation criteria;

in accordance with a determination that the position of the document meets the orientation criteria, performing the text recognition on the portion of the received image data; and

determining the position for the image of the document, including:

determining one or more facial features corresponding to the facial image data; and

determining the position of the image of the document of the image data based on the one or more facial features.

9. The method of claim 8 , wherein determining the position of the image of the document includes:

identifying respective corners of the image of the document; and

comparing the respective corners of the image of the document with document characteristics corresponding to a document type to determine the position of the document.

10. The method of claim 8 , wherein:

determining a saliency value for the predefined document field;

determining whether the saliency value for the predefined document field meets a predetermined saliency threshold; and

upon a determination that the saliency value does not meet the predetermined saliency threshold, requesting new image data that includes an image of the document.

11. The method of claim 8 , wherein:

the predefined document field includes text; and

determining the position for the image of the document includes:

determining, based on the predefined document field, a text position; and

determining the position of the image of the document based on the text position.

12. The method of claim 8 , wherein determining the position of the image of the document includes cropping the image of the document in the image data.

13. The method of claim 8 , further comprising:

in accordance with a determination that the position of the image of the document does not meet the orientation criteria, adjusting the image of the document to satisfy the orientation criteria; and

in accordance with a determination that the position corresponding to the adjusted image of the document meets the orientation criteria, performing the text recognition on the adjusted image of the document.

14. A computer-implemented method, further comprising:

at a server system including one or more processors and memory storing one or more programs for execution by the one or more processors:

receiving image data, the image data including an image of a document;

determining a portion of the received image data that corresponds to a predefined document field;

utilizing a neural network system to assign a label to the determined portion of the received image data;

performing text recognition on the portion of the received image data;

storing recognized text in association with the assigned label;

determining a saliency value for the predefined document field;

determining whether the saliency value for the predefined document field meets a predetermined saliency threshold; and

in accordance with a determination that the saliency value does not meet the predetermined saliency threshold, requesting new image data that includes an image of the document.

15. The method of claim 14 , further comprising:

in accordance with a determination that the saliency value meets the predetermined saliency threshold, generating a bounding box for the predefined document field; and

performing the text recognition on the generated bounding box.

16. The method of claim 15 , wherein utilizing the neural network system to assign the label to the determined portion of the received image data includes:

determining a label for the generated bounding box; and

assigning the label to the generated bounding box.

17. The method of claim 16 , wherein the neural network system includes at least one recurrent neural network (RNN), or a convolutional neural network (CNN).

18. The method of claim 17 , wherein the neural network system includes a plurality of neural networks, the plurality of neural networks including both the RNN and the CNN.

19. The method of claim 17 , wherein:

the neural network system includes a plurality of neural networks, the plurality of neural networks including both the RNN and the CNN; and

determining the label for the generated bounding box includes:

determining, using a first neural network of the plurality of neural networks, a first label for the generated bounding box;

determining, using a second neural network of the plurality of neural networks, a second label for the generated bounding box; and

comparing the first label and the second label to determine whether the first label and the second label match; and

in accordance with a determination that the first label and the second label match, assigning the first label or the second label to the generated bounding box.

20. The method of claim 19 , further comprising in accordance with a determination that the first label and the second label do not match, assigning a respective label of the first label or the second label with a highest relevance score.

21. The method of claim 16 , wherein the neural network system includes a registration system, the registration system including a first template, wherein the first template includes a first predetermined label, the first predetermined label associated with a first predetermined label location; and wherein:

determining the label for the generated bounding box includes:

determining whether the first predetermined label corresponds to the generated bounding box by superimposing the first template over the image of the document;

comparing the first predetermined label location with the generated bounding box to determine a template value;

determining whether the template value meets similarity threshold; and

in accordance with a determination that the template value meets the similarity threshold, determining a relevant label based on the first predetermined label; and

in accordance with a determination that the template value for the first template does not meet the similarity threshold, determining the label for the generated bounding box based on a second template.

22. The method of claim 21 , wherein determining the template value includes determining respective distances between the first predetermined label location and one or more edges of the image of the document, the respective distances measured based on one or more pixels between the first predetermined label location and the one or more edges.

23. A non-transitory computer readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed, cause a device to:

receive image data, the image data including an image of a document, wherein the image of the document includes facial image data;

determine a portion of the received image data that corresponds to a predefined document field;

determine a position for the image of the document, by identifying one or more facial features within the facial image data;

utilize a neural network system to assign a label to the determined portion of the received image data;

perform text recognition on the portion of the received image data; and

store recognized text in association with the assigned label.

24. A system, comprising:

one or more processors;

memory; and

one or more programs, wherein the one or more programs are stored in the memory and are configured for execution by the one or more processors, the one or more programs including instructions for:

receiving image data, the image data including an image of a document, wherein the image of the document includes facial image data;

determining a portion of the received image data that corresponds to a predefined document field;

determine a position for the image of the document, by identifying one or more facial features within the facial image data;

utilizing a neural network system to assign a label to the determined portion of the received image data;

performing text recognition on the portion of the received image data; and

storing recognized text in association with the assigned label.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 16, 2022
From: MELCHY, PIERRE-ÉRIC MICHAËL ALIX; PATEL, LABHESH; POINTNER, PHILIPP; ROGOJANU, RADU; BON, ANTOINE
To: JUMIO CORPORATION
Reel/Frame 060230/0426 →
Continuity (2)
Continuation PCTUS2019067747 · Dec 20, 2019
Related Publication 20220309813A1 · Sep 29, 2022
References Cited (16)
US 10467464B2 · Chen · 2019 [cited by examiner]
US 10860785B2 · Miyamoto · 2020 [cited by examiner]
US 11087123B2 · Chitta · 2021 [cited by examiner]
US 11853406B2 · Benkreira · 2023 [cited by examiner]
US 12019675B2 · Tripuraneni · 2024 [cited by examiner]
US 20170351913A1 · Chen · 2017 [cited by examiner]
US 20190370688A1 · Patel · 2019 [cited by examiner]
US 20210200937A1 · Wheaton · 2021 [cited by examiner]
US 20230206619A1 · Lund · 2023 [cited by examiner]
US 20240126855A1 · Benkreira · 2024 [cited by examiner]
US 20240135700A1 · Flament · 2024 [cited by examiner]
US 20240346069A1 · Tripuraneni · 2024 [cited by examiner]
CN 110222695A · 2019 [cited by applicant]
Jumio Corporation, Communication Pursuant to Article 94(3), EP19842953.2, Apr. 25, 2024, 5 pgs. [cited by applicant]
Jumio Corporation, International Search Report and Writen Opinion, PCT/US2019/067747, Sep. 7, 2020, 11 pgs. [cited by applicant]
Jumio Corporation, International Preliminary Report on Patentability, PCT/US2019/067747, May 17, 2022, 8 pgs. [cited by applicant]