IP Library Granted Patent US 12,444,213
Granted Patent B1
US 12,444,213 · App. 19/084,341 · Granted Oct 14, 2025

Machine-learning models for image processing

Inventors: Ashutosh K. Sureka (Irving, TX); Venkata Sesha Kiran Kumar Adimatyam (Irving, TX); Miriam Silver (Tel Aviv, IL); Daniel Funken (Irving, TX)
Assignee: CITIBANK, N.A.
G06V20/95G06T5/60G06T7/0002G06V30/18G06V30/191G06V30/418G06V30/42G06T2200/24G06T2207/10016G06T2207/20081G06T2207/20084G06T2207/30168G06T2207/30176G06V2201/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,444,213
App. No.
19/084,341
Granted
Oct 14, 2025
Kind
B1
Abstract

Presented herein are systems and methods for the employment of machine learning models for image processing. A method may include a capture of a video feed including image data of a document at a client device. The client device can provide the video feed to another computing device. The method can include, by the client device or the other computing device object recognition for recognizing a type of document and capturing an image exceeding a quality threshold of the document amongst the frames within the video feed. The method may further include the execution of other image processing operations on the image data to improve the quality of the image or features extracted therefrom. The method may further include anti-fraud detection or scoring operations to determine an amount of risk associated with the image data.

Claims (43)

1. A method for remotely processing document imagery, the method comprising:

receiving, by a computer remote from a user device, via one or more networks, a video feed comprising a plurality of frames from the user device, at least one frame including image data depicting an object;

executing, by the computer, an object recognition engine of a machine-learning architecture using the image data of the plurality of frames, the object recognition engine trained for detecting content data of a document in the image data;

generating, by the computer, a risk score based upon a distance of the content data to a predefined template to validate the content data, wherein the predefined template comprises a plurality of fields comprising a plurality of characters; and

generating, by the computer, an output image representing the document having the content data based upon the content data on the document in each frame of the at least one frame, responsive to validating the document using the risk score.

2. The method of claim 1 , further comprising:

comparing, by the computer, one or more features of the output image to a quality threshold to determine a compliance with a predefined standard; and

responsive to the computer determining the compliance, transmitting, by the computer, the output image to a provider server.

3. The method of claim 1 , further comprising providing, by the computer, the content data to a same destination device as the output image responsive to validating the document.

4. The method of claim 3 , wherein the validation further comprises:

generating, by the computer, a prompt comprising the content data for presentation via a user interface of the user device; and

receiving, by the computer, from a control element of the user device a confirmation of the content data as presented in the prompt.

5. The method of claim 1 , wherein the predefined template comprises one or more fields including a format including a quantity of the plurality of characters.

6. The method of claim 1 , wherein the output image is generated from a set of one or more frames of the plurality of frames of the image data.

7. The method of claim 1 , wherein generating the output image comprises generating, by the computer, a reconstructed image based on the image data from at least one frame of the plurality of frames.

8. The method of claim 7 , further comprising executing, by the computer, a decoder of the machine-learning architecture to generate the output image.

9. The method of claim 7 , further comprising determining, by the computer, a metric for the reconstructed image indicative of a deviation of the reconstructed image from the document.

10. The method of claim 9 , wherein the metric is based on a difference between one or more frames of the image data and the reconstructed image.

11. The method of claim 7 , further comprising:

generating, by the computer, a prompt comprising the reconstructed image for presentation via a user interface; and

receiving, by the computer, from a control element of the user device a confirmation of the reconstructed image.

12. The method of claim 11 , further comprising:

presenting, by the computer, the prompt comprising the content data via the user interface of the user device, the content data derived from the reconstructed image; and

providing, by the computer, textual content of the content data as metadata of the reconstructed image to a provider server.

13. A system comprising:

a computing device remote from a user device, the computing device comprising at least one processor configured to:

receive, via one or more networks, a video feed comprising a plurality of frames from the user device, at least one frame including image data depicting an object;

execute an object recognition engine of a machine-learning architecture using the image data of the plurality of frames, the object recognition engine trained for detecting content data of a document in the image data;

generate a risk score based upon a distance between the content data and a predefined template to validate the content data, wherein the predefined template comprises a plurality of fields comprising a plurality of characters; and

generate an output image representing the document having the content data based upon the content data on the document in each frame of the at least one frame, responsive to validating the document using the risk score.

14. The system of claim 13 , wherein the computing device is further configured to:

compare one or more features of the output image to a quality threshold to determine a compliance with a predefined standard; and

responsive to determining the compliance, transmit the output image to a provider server.

15. The system of claim 13 , wherein the computing device is further configured to provide the content data to a same destination device as the output image responsive to validating the document.

16. The system of claim 15 , wherein the predefined template comprises one or more fields including a format including a quantity of the plurality of characters.

17. The system of claim 13 , further comprising the user device configured to:

generate a prompt comprising the content data for display via a user interface of the user device; and

receive, from a control element of the user device a confirmation of the content data.

18. The system of claim 13 , wherein the output image is generated from a set of one or more frames of the plurality of frames of the image data.

19. The system of claim 13 , wherein when generating the output image, the computing device is further configured to generate a reconstructed image based on the image data from at least one frame of the plurality of frames.

20. The system of claim 19 , wherein the computing device is further configured to:

determine a metric for the reconstructed image indicative of a deviation of the reconstructed image from the document,

wherein the metric is based on a difference between one or more frames of the image data and the reconstructed image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2025
From: SUREKA, ASHUTOSH K.; ADIMATYAM, VENKATA SESHA KIRAN KUMAR; SILVER, MIRIAM; FUNKEN, DANIEL
To: CITIBANK, N.A.
Reel/Frame 070562/0209 →
Continuity (1)
Continuation 18629301 · Apr 8, 2024
References Cited (65)
US 8688579B1 · Ethington et al. · 2014 [cited by applicant]
US 9137417B2 · Macciola et al. · 2015 [cited by applicant]
US 9672510B2 · Roach et al. · 2017 [cited by applicant]
US 10019772B1 · Smith · 2018 [cited by applicant]
US 10115031B1 · Pashinstev et al. · 2018 [cited by applicant]
US 10339374B1 · Pribble et al. · 2019 [cited by applicant]
US 10635898B1 · Pribble · 2020 [cited by applicant]
US 10803431B2 · Hinski · 2020 [cited by applicant]
US 11068976B1 · Voutour et al. · 2021 [cited by applicant]
US 11321679B1 · Bueche et al. · 2022 [cited by applicant]
US 11900755B1 · Bueche · 2024 [cited by applicant]
US 12039504B1 · Foster · 2024 [cited by examiner]
US 12056434B2 · Bhatia et al. · 2024 [cited by applicant]
US 12106590B1 · Kinsey · 2024 [cited by applicant]
US 12272111B1 · Sureka et al. · 2025 [cited by applicant]
US 20070233615A1 · Tumminaro · 2007 [cited by applicant]
US 20110091092A1 · Nepomniachtchi et al. · 2011 [cited by applicant]
US 20110243459A1 · Deng · 2011 [cited by applicant]
US 20120230577A1 · Calman et al. · 2012 [cited by applicant]
US 20130022231A1 · Nepomniachtchi et al. · 2013 [cited by applicant]
US 20140270464A1 · Nepomniachtchi et al. · 2014 [cited by applicant]
US 20150003666A1 · Wang et al. · 2015 [cited by applicant]
US 20150032631A1 · Hinski · 2015 [cited by examiner]
US 20150063653A1 · Madhani et al. · 2015 [cited by applicant]
US 20150120572A1 · Slade · 2015 [cited by applicant]
US 20150278819A1 · Song et al. · 2015 [cited by applicant]
US 20150294523A1 · Smith · 2015 [cited by applicant]
US 20150302242A1 · Lee et al. · 2015 [cited by applicant]
US 20150309966A1 · Gupta et al. · 2015 [cited by applicant]
US 20160037071A1 · Emmett et al. · 2016 [cited by applicant]
US 20160125613A1 · Shustorovich et al. · 2016 [cited by applicant]
US 20160253569A1 · Eid et al. · 2016 [cited by applicant]
US 20170116494A1 · Isaev · 2017 [cited by applicant]
US 20170185833A1 · Wang et al. · 2017 [cited by applicant]
US 20180211243A1 · Ekpenyong et al. · 2018 [cited by applicant]
US 20180330342A1 · Prakash et al. · 2018 [cited by applicant]
US 20180376193A1 · Sullivan et al. · 2018 [cited by applicant]
US 20190019020A1 · Flament et al. · 2019 [cited by applicant]
US 20190197693A1 · Zagaynov et al. · 2019 [cited by applicant]
US 20190213408A1 · Cali et al. · 2019 [cited by applicant]
US 20200279138A1 · Xu et al. · 2020 [cited by applicant]
US 20200410291A1 · Kriegman et al. · 2020 [cited by applicant]
US 20210124919A1 · Balakrishnan et al. · 2021 [cited by applicant]
US 20210182547A1 · Ayyadevara et al. · 2021 [cited by applicant]
US 20210350516A1 · Tang et al. · 2021 [cited by applicant]
US 20210360149A1 · Mukul · 2021 [cited by examiner]
US 20210365677A1 · Anzenberg · 2021 [cited by applicant]
US 20220224816A1 · Pribble et al. · 2022 [cited by applicant]
US 20220358575A1 · Smith · 2022 [cited by applicant]
US 20220414955A1 · Ota · 2022 [cited by applicant]
US 20230030792A1 · Zheng et al. · 2023 [cited by applicant]
US 20230143239A1 · Yusuf et al. · 2023 [cited by applicant]
US 20230281629A1 · Shevyrev et al. · 2023 [cited by applicant]
US 20230281820A1 · Pizzocchero et al. · 2023 [cited by applicant]
US 20230298370A1 · Nishioka · 2023 [cited by examiner]
US 20240061992A1 · Bhatia et al. · 2024 [cited by applicant]
US 20240176951A1 · Krishnamoorthy · 2024 [cited by applicant]
US 20240177487A1 · Sohoni · 2024 [cited by applicant]
US 20240256955A1 · Goodsitt et al. · 2024 [cited by applicant]
US 20240303658A1 · Cohen · 2024 [cited by examiner]
US 20240428550A1 · Gutierrez Valdes · 2024 [cited by examiner]
CA 3188665A1 · 2024 [cited by applicant]
Chernov, Timofey S., Sergey Allyuhin, and Vladimir V. Arlazarov. “Application of dynamic saliency maps to the video stream recognition systems with image quality assessment.” Eleventh International Conference on Machine… [cited by applicant]
Rybakova et al., “PESAC, the Generalized Framework for RANSAC-Based Methods on SIMD Computing Platforms”, IEEE Access, pp. 82151-82166, 2023, vol. 1. [cited by applicant]
PCT International Search Report and Written Opinion for Application No. PCT/US2025/023145 mailing date Jun. 26, 2025, 34 pages. [cited by applicant]