IP Library › Granted Patent US 12,272,111
Granted Patent B1
US 12,272,111 · App. 18/890,276 · Granted Apr 8, 2025

Machine-learning models for image processing

Inventors: Ashutosh K. Sureka (Irving, TX); Venkata Sesha Kiran Kumar Adimatyam (Irving, TX); Miriam Silver (Tel Aviv, IL); Daniel Funken (Irving, TX)
Assignee: Citibank, N.A.
G06V10/25G06T7/13G06V20/70G06T2207/30176G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,272,111
App. No.
18/890,276
Granted
Apr 8, 2025
Kind
B1
Abstract

Presented herein are systems and methods for the employment of machine learning models for image processing. A method may include a capture of a video feed including image data of a document at a client device. The client device can provide the video feed to another computing device. The method can include, by the client device or the other computing device object recognition for recognizing a type of document and capturing an image exceeding a quality threshold of the document amongst the frames within the video feed. The method may further include the execution of other image processing operations on the image data to improve the quality of the image or features extracted therefrom. The method may further include anti-fraud detection or scoring operations to determine an amount of risk associated with the image data.

Claims (31)

1. A method for capturing document imagery using augmented reality bounding boxes, the method comprising:

obtaining, by a computing device, image data containing an object;

detecting, by the computing device, that the object is a document by executing an object recognition engine of a machine-learning architecture using the image data as an input, the object recognition engine trained for detecting a type of document of the document in the image data; and

generating, by the computing device, an augmented reality (AR) overlay using the image data for display at a graphical user interface, the AR overlay including a bounding box for display in the image data containing the object according to pixel data corresponding to the document, the user interface including a dynamic alignment indicator based upon the bounding box for the document having a first label annotation of the bounding box, and a second label annotation for the image data disposed at a portion of the bounding box.

2. The method according to claim 1 , wherein the computing device obtains the image data from a client device via one or more networks, and wherein the computing device transmits the AR overlay to the client device via the one or more networks.

3. The method according to claim 1 , wherein the generating the AR overlay includes generating, by the computing device, the bounding box for the document according to the pixel data corresponding to the document in the image data.

4. The method according to claim 1 , further comprising generating, by the computing device, a prompt for display via the user interface, the prompt to provide an indication to adjust a camera device.

5. The method according to claim 4 , wherein the prompt is generated for display at a periodic interval.

6. The method according to claim 1 , further comprising receiving, by the computing device from the user interface, an indication of a location of the bounding box.

7. The method according to claim 1 , wherein generating the AR overlay includes generating, by the computing device, for display at the user interface the dynamic alignment indicator based upon the bounding box for the document.

8. The method according to claim 7 , wherein generating the dynamic alignment indicator includes:

generating, by the computing device, the first label annotation of the bounding box; and

generating, by the computing device, the second label annotation for the image data, the second label annotation corresponding to the bounding box.

9. The method according to claim 7 , wherein the dynamic alignment indicator of the AR overlay comprises an overlay element visually corresponding to at least one of: an edge of the document, a corner of the document, or a data field of the document.

10. The method according to claim 1 , wherein the computing device detects the object as being the type of document in a video frame of the image data.

11. A system comprising:

a computing device comprising at least one processor, configured to:

obtain image data containing an object;

detect that the object is a document by executing an object recognition engine of a machine-learning architecture using the image data as an input, the object recognition engine trained for detecting a type of document of the document in the image data; and

generate an augmented reality (AR) overlay using the image data for display at a graphical user interface, the AR overlay including a bounding box for display in the image data containing the object according to pixel data corresponding to the document, the user interface including a dynamic alignment indicator based upon the bounding box for the document having a first label annotation of the bounding box, and a second label annotation for the image data disposed at a portion of the bounding box.

12. The system according to claim 11 , wherein the computing device obtains the image data from a client device via one or more networks, and wherein the computing device transmits the AR overlay to the client device via the one or more networks.

13. The system according to claim 11 , wherein when generating the AR overlay the computing device is configured to generate the bounding box for the document according to the pixel data corresponding to the document in the image data.

14. The system according to claim 11 , wherein the computing device is further configured to generate a prompt for display via the user interface, the prompt to provide an indication to adjust a camera device.

15. The system according to claim 14 , wherein the prompt is generated for display at a periodic interval.

16. The system according to claim 11 , wherein the computing device is further configured to receive, from the user interface, an indication of a location of the bounding box.

17. The system according to claim 11 , wherein when generating the AR overlay the computing device is configured to generate for display at the user interface the dynamic alignment indicator based upon the bounding box for the document.

18. The system according to claim 17 , wherein when generating the dynamic alignment indicator the computing device is configured to:

generate the first label annotation of the bounding box; and

generate the second label annotation for the image data, the second label annotation corresponding to the bounding box.

19. The system according to claim 17 , wherein the dynamic alignment indicator of the AR overlay comprises an overlay element visually corresponding to at least one of: an edge of the document, a corner of the document, or a data field of the document.

20. The system according to claim 11 , wherein the computing device detects the object as being the type of document in a video frame of the image data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 19, 2024
From: SUREKA, ASHUTOSH K.; ADIMATYAM, VENKATA SESHA KIRAN KUMAR; SILVER, MIRIAM; FUNKEN, DANIEL
To: CITIBANK, N.A.
Reel/Frame 068638/0486 →
Continuity (1)
Continuation 18629259 · Apr 8, 2024
References Cited (44)
US 8688579B1 · Ethington et al. · 2014 [cited by applicant]
US 9672510B2 · Roach · 2017 [cited by examiner]
US 10115031B1 · Pashintsev · 2018 [cited by examiner]
US 10339374B1 · Pribble et al. · 2019 [cited by applicant]
US 10635898B1 · Pribble · 2020 [cited by examiner]
US 10803431B2 · Hinski · 2020 [cited by applicant]
US 11068976B1 · Voutour et al. · 2021 [cited by applicant]
US 11321679B1 · Bueche et al. · 2022 [cited by applicant]
US 11900755B1 · Bueche · 2024 [cited by applicant]
US 12039504B1 · Foster et al. · 2024 [cited by applicant]
US 12106590B1 · Kinsey · 2024 [cited by applicant]
US 20070233615A1 · Tumminaro · 2007 [cited by applicant]
US 20110091092A1 · Nepomniachtchi et al. · 2011 [cited by applicant]
US 20110243459A1 · Deng · 2011 [cited by applicant]
US 20120230577A1 · Calman et al. · 2012 [cited by applicant]
US 20150003666A1 · Wang et al. · 2015 [cited by applicant]
US 20150120572A1 · Slade · 2015 [cited by applicant]
US 20150278819A1 · Song et al. · 2015 [cited by applicant]
US 20150294523A1 · Smith · 2015 [cited by applicant]
US 20150309966A1 · Gupta et al. · 2015 [cited by applicant]
US 20160037071A1 · Emmett et al. · 2016 [cited by applicant]
US 20160125613A1 · Shustorovich et al. · 2016 [cited by applicant]
US 20160253569A1 · Eid et al. · 2016 [cited by applicant]
US 20170116494A1 · Isaev · 2017 [cited by applicant]
US 20170185833A1 · Wang et al. · 2017 [cited by applicant]
US 20180211243A1 · Ekpenyong et al. · 2018 [cited by applicant]
US 20180330342A1 · Prakash et al. · 2018 [cited by applicant]
US 20180376193A1 · Sullivan et al. · 2018 [cited by applicant]
US 20190019020A1 · Flament et al. · 2019 [cited by applicant]
US 20190197693A1 · Zagaynov · 2019 [cited by examiner]
US 20200410291A1 · Kriegman et al. · 2020 [cited by applicant]
US 20210350516A1 · Tang et al. · 2021 [cited by applicant]
US 20210360149A1 · Mukul · 2021 [cited by applicant]
US 20210365677A1 · Anzenberg · 2021 [cited by applicant]
US 20220224816A1 · Pribble · 2022 [cited by examiner]
US 20220414955A1 · Ota · 2022 [cited by applicant]
US 20230281629A1 · Shevyrev et al. · 2023 [cited by applicant]
US 20240061992A1 · Bhatia · 2024 [cited by examiner]
US 20240176951A1 · Krishnamoorthy · 2024 [cited by applicant]
US 20240256955A1 · Goodsitt et al. · 2024 [cited by applicant]
US 20240303658A1 · Cohen et al. · 2024 [cited by applicant]
CA 3188665A1 · 2024 [cited by applicant]
Chernov, Timofey S., Sergey Allyuhin, and Vladimir V. Arlazarov. “Application of dynamic saliency maps to the video stream recognition systems with image quality assessment.” Eleventh International Conference on Machine… [cited by applicant]
Non-Final Office Action on U.S. Appl. No. 18/943,493 Dtd Jan. 7, 2025. [cited by applicant]
Cited By (12)
US 12,387,512 US 12,444,213 US 12,456,322 US 12,456,323 US 12,462,366 US 12,482,286 US 12,525,048 US 12,541,989 US 12,586,402 US 12,608,743 US 12,688,667 US 12,718,605