IP Library Granted Patent US 12,266,145
Granted Patent B1
US 12,266,145 · App. 18/629,259 · Granted Apr 1, 2025

Machine-learning models for image processing

Inventors: Ashutosh K. Sureka (Irving, TX); Venkata Sesha Kiran Kumar Adimatyam (Irving, TX); Miriam Silver (Tel Aviv, IL); Daniel Funken (Irving, TX)
Assignee: CITIBANK, N.A.
G06V10/25G06T7/13G06V20/70G06T2207/30176G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,266,145
App. No.
18/629,259
Granted
Apr 1, 2025
Kind
B1
Abstract

Presented herein are systems and methods for the employment of machine learning models for image processing. A method may include a capture of a video feed including image data of a document at a client device. The client device can provide the video feed to another computing device. The method can include, by the client device or the other computing device object recognition for recognizing a type of document and capturing an image exceeding a quality threshold of the document amongst the frames within the video feed. The method may further include the execution of other image processing operations on the image data to improve the quality of the image or features extracted therefrom. The method may further include anti-fraud detection or scoring operations to determine an amount of risk associated with the image data.

Claims (33)

1. A method for capturing document imagery using object recognition bounding boxes, the method comprising:

obtaining, by a computer via one or more networks, image data depicting an object from a user device;

executing, by the computer, an object recognition engine of a machine-learning architecture using the image data as an input, the object recognition engine trained for detecting a type of document in the image data;

in response to detecting, by the computer, the object as a document of the type of document in the image data, generating, by the computer, a bounding box for the document according to pixel data corresponding to the document in the image data; and

generating, by the computer, for display at a user interface a dynamic alignment indicator based upon the bounding box for the document, wherein the computer generates a first label annotation of the bounding box and a second label annotation for the image data, the second label annotation corresponding to the bounding box.

2. The method of claim 1 , further comprising, generating an augmented image overlay comprising the bounding box over the image data.

3. The method of claim 1 , further comprising generating a prompt for display via the user interface, the prompt to provide an indication to adjust a camera device.

4. The method of claim 3 , wherein the prompt is generated for display at a periodic interval.

5. The method of claim 1 , further comprising receiving, from the user interface, an indication of a location of the bounding box.

6. The method of claim 1 , wherein the dynamic alignment indicator comprises an overlay element visually corresponding to at least one of:

an edge of the document;

a corner of the document; or

a data field of the document.

7. The method of claim 1 , further comprising receiving, from a user via a control element of the user device, a confirmation of an element of the dynamic alignment indicator.

8. The method of claim 7 , wherein the confirmation of the element of the dynamic alignment indicator is based on a selection of one of a plurality of displayed elements of the dynamic alignment indicator.

9. The method of claim 1 , wherein the dynamic alignment indicator is provided for the object in a first frame, based on the object detected in a second frame.

10. The method of claim 1 , further comprising updating a machine learning model of the object recognition engine responsive to a detection of content of the document identified in the image data.

11. A system comprising:

a server comprising at least one processor, configured to:

obtain, via one or more networks, image data depicting an object from a user device;

execute an object recognition engine of a machine-learning architecture using the image data as an input, the object recognition engine trained for detecting a type of document in the image data;

generate a bounding box for a document according to pixel data corresponding to the document in the image data in response to detecting the object as the document of the type of document in the image data; and

generate a dynamic alignment indicator for display at a user interface based upon the bounding box for the document, wherein the server generates a first label annotation of the bounding box and a second label annotation for the image data, the second label annotation corresponding to the bounding box.

12. The system of claim 11 , wherein the dynamic alignment indicator includes an augmented image overlay comprising the bounding box over the image data.

13. The system of claim 11 , wherein the server is configured to cause a generation of a prompt for display via the user interface, the prompt to provide an indication to adjust a camera device.

14. The system of claim 11 , wherein the dynamic alignment indicator comprises an overlay element visually corresponding to at least one of:

an edge of the document;

a corner of the document; or

a data field of the document.

15. The system of claim 11 , further comprising the user device to receive, from a user via a control element, a confirmation of an element of the dynamic alignment indicator.

16. The system of claim 15 , wherein the confirmation of the element of the dynamic alignment indicator is based on a selection of one of a plurality of displayed elements of the dynamic alignment indicator.

17. The system of claim 11 , wherein the dynamic alignment indicator is provided for the object in a first frame, based on the object detected in a second frame.

18. The system of claim 11 , wherein the server is configured to update a machine learning model of the object recognition engine responsive to a detection of content of the document identified in the image data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 20, 2024
From: SUREKA, ASHUTOSH K.; ADIMATYAM, VENKATA SESHA KIRAN KUMAR; SILVER, MIRIAM; FUNKEN, DANIEL
To: CITIBANK, N.A.
Reel/Frame 068343/0898 →
References Cited (42)
US 8688579B1 · Ethington et al. · 2014 [cited by applicant]
US 9672510B2 · Roach et al. · 2017 [cited by applicant]
US 10115031B1 · Pashintsev et al. · 2018 [cited by applicant]
US 10339374B1 · Pribble et al. · 2019 [cited by applicant]
US 10635898B1 · Pribble · 2020 [cited by applicant]
US 10803431B2 · Hinski · 2020 [cited by applicant]
US 11068976B1 · Voutour et al. · 2021 [cited by applicant]
US 11321679B1 · Bueche, Jr. · 2022 [cited by examiner]
US 11900755B1 · Bueche · 2024 [cited by applicant]
US 12039504B1 · Foster et al. · 2024 [cited by applicant]
US 12106590B1 · Kinsey · 2024 [cited by applicant]
US 20070233615A1 · Tumminaro · 2007 [cited by applicant]
US 20110091092A1 · Nepomniachtchi et al. · 2011 [cited by applicant]
US 20110243459A1 · Deng · 2011 [cited by applicant]
US 20120230577A1 · Calman et al. · 2012 [cited by applicant]
US 20150003666A1 · Wang et al. · 2015 [cited by applicant]
US 20150120572A1 · Slade · 2015 [cited by applicant]
US 20150278819A1 · Song et al. · 2015 [cited by applicant]
US 20150294523A1 · Smith · 2015 [cited by examiner]
US 20150309966A1 · Gupta et al. · 2015 [cited by applicant]
US 20160037071A1 · Emmett et al. · 2016 [cited by applicant]
US 20160125613A1 · Shustorovich et al. · 2016 [cited by applicant]
US 20160253569A1 · Eid · 2016 [cited by examiner]
US 20170116494A1 · Isaev · 2017 [cited by applicant]
US 20170185833A1 · Wang et al. · 2017 [cited by applicant]
US 20180211243A1 · Ekpenyong · 2018 [cited by examiner]
US 20180330342A1 · Prakash et al. · 2018 [cited by applicant]
US 20180376193A1 · Sullivan et al. · 2018 [cited by applicant]
US 20190019020A1 · Flament · 2019 [cited by examiner]
US 20190197693A1 · Zagaynov et al. · 2019 [cited by applicant]
US 20200410291A1 · Kriegman · 2020 [cited by examiner]
US 20210350516A1 · Tang et al. · 2021 [cited by applicant]
US 20210360149A1 · Mukul · 2021 [cited by applicant]
US 20210365677A1 · Anzenberg · 2021 [cited by applicant]
US 20220224816A1 · Pribble et al. · 2022 [cited by applicant]
US 20220414955A1 · Ota · 2022 [cited by applicant]
US 20230281629A1 · Shevyrev et al. · 2023 [cited by applicant]
US 20240176951A1 · Krishnamoorthy · 2024 [cited by applicant]
US 20240256955A1 · Goodsitt et al. · 2024 [cited by applicant]
US 20240303658A1 · Cohen et al. · 2024 [cited by applicant]
CA 3188665A1 · 2024 [cited by examiner]
Chernov, Timofey S., Sergey A llyuhin, and Vladimir V. Arlazarov. “Application of dynamic saliency maps to the video stream recognition systems with image quality assessment.” Eleventh International Conference on Machin… [cited by applicant]
Cited By (1)
US 12,579,832