IP Library › Granted Patent US 12,387,512
Granted Patent B1
US 12,387,512 · App. 18/943,493 · Granted Aug 12, 2025

Machine-learning models for image processing

Inventors: Ashutosh K. Sureka (Irving, TX); Venkata Sesha Kiran Kumar Adimatyam (Irving, TX); Miriam Silver (Tel Aviv, IL); Daniel Funken (Irving, TX)
Assignee: CITIBANK, N.A.
G06V20/95G06V10/24G06V10/44G06V10/70G06V30/413G06V30/414G06V30/418G06V30/42G06Q20/0425
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,512
App. No.
18/943,493
Granted
Aug 12, 2025
Kind
B1
Abstract

Presented herein are systems and methods for the employment of machine learning models for image processing as may be performed by computing devices associated with an end user. A method may include obtaining video data comprising a plurality of frames including a document of a document type. The method may include executing an object recognition engine of a machine-learning architecture using image data of the plurality of frames, the object recognition engine trained to detect edges of documents. The method may include identifying, based on the edge detection, a plurality of boundaries for the document. The method may include validating, based on the plurality of boundaries, the document as the document type. The method may include transmitting via one or more networks, to a computer remote from the computing device, responsive to the validation of the type of document, the image data for the plurality of frames depicting the document.

Claims (43)

1. A method for client-side validation of document-imagery for remote processing, the method comprising:

obtaining, by a camera of a mobile client device associated with an end-user, video data comprising a plurality of frames including a document of a document type of the document;

executing, by the mobile client device, an object recognition engine of a machine-learning architecture to extract a first set of document features from image data of the plurality of frames captured by the camera of the mobile client device;

executing, by the mobile client device, the object recognition engine on the first set of document features to detect the document type of the document based upon the first set of document features, the object recognition engine trained to detect the document type of the document using a set of document features and corresponding training labels indicating the document type of the document having the set of document features;

executing, by the mobile client device, the object recognition engine to extract a second set of document features from the image data, the second set of document features are extracted based upon the document type of the document detected using the first set of document features;

generating, by the mobile client device, a document validation score indicating a likelihood that the document is a valid document based upon the second set of document features extracted from the image data; and

upon validating the document based on determining that the document validation score satisfies a document validation threshold,

generating, by the mobile client device, a packaged image extracted from the image data of at least one frame of the plurality of frames of the video data captured by the mobile client device; and

generating, by the mobile client device, an operation instruction for a backend server, the operation instruction to include the packaged image and mobile client device metadata identifying the mobile client device.

2. The method of claim 1 , wherein validating the document includes identifying, by the mobile client device, a dimension similarity between the document and the document type of the document, based upon comparing a plurality of boundaries corresponding to a plurality of edges of the document in the first set of document features against a predefined dimension having a plurality of expected boundaries for the document type of the document.

3. The method of claim 2 , wherein the identifying the dimension similarity includes matching the plurality of boundaries to the predefined dimension having the plurality of expected boundaries for a plurality of document issuers, wherein each of the plurality of expected boundaries correspond to a document issuer of the plurality of document issuers.

4. The method of claim 1 , further comprising generating, by the mobile client device, an indication of validation for presenting via a user interface of the mobile client device.

5. The method of claim 1 , further comprising generating, by the mobile client device, an indication that the document validation score exceeds a warning threshold for presenting via a user interface of the mobile client device.

6. The method of claim 1 , wherein validating the document includes:

comparing, by the mobile client device, the document validation score against at least one of a second validation threshold corresponding to non-validation and a warning threshold corresponding to an alert trigger; and

generating, by the mobile client device, an output indicator for display at a user interface, based upon comparing the document validation score against the at least one of the second validation threshold or the warning threshold.

7. The method of claim 1 , wherein validating the document includes identifying, by the mobile client device, at least one of a first set of validation criteria corresponding to clerical or imaging errors or a second set of validation criteria corresponding to digital or mechanical manipulation of the document.

8. The method of claim 1 , wherein the machine-learning architecture comprises a classification model configured to classify the document the document type of the document.

9. The method of claim 1 , wherein the machine-learning architecture includes an edge-detection model configured to detect one or more edges of a rectangle bounding the document for validating the document.

10. The method of claim 1 , further comprising generating, by the mobile client device, a spatial transform for the document responsive to detecting a plurality of edges of the document of the first set of document features.

11. The method of claim 1 , wherein validating the document includes:

identifying, by the mobile client device, a subset of the video data including the document; and

comparing, by the mobile client device, the subset against an occupancy threshold.

12. The method of claim 1 , wherein the image data transmitted to the backend server remote from the mobile client device is transmitted in a video feed comprising multiple of the plurality of frames of the video data.

13. A system for client-side validation of document-imagery for remote processing, the system comprising:

a mobile client device associated with an end-user comprising at least one processor and a camera, configured to:

obtain, by the camera, video data comprising a plurality of frames including a document of a document type of the document;

execute an object recognition engine of a machine-learning architecture to extract a first set of document features from image data of the plurality of frames captured by the camera of the mobile client device;

execute the object recognition engine on the first set of document features to detect the document type of the document based upon the first set of document features, the object recognition engine trained to detect the document type using a set of document features and corresponding training labels indicating the document type of the document having the set of document features;

execute the object recognition engine to extract a second set of document features from the image data, the second set of document features are extracted based upon the document type of the document detected using the first set of document features;

generate a document validation score indicating a likelihood that the document is a valid document based upon the second set of document features extracted from the image data; and

upon the mobile client device validating the document based on determining that the document validation score satisfies a document validation threshold,

generate a packaged image extracted from the image data of at least one frame of the plurality of frames of the video data captured by the mobile client device; and

generate an operation instruction for a backend server, the operation instruction to include the packaged image and mobile client device metadata identifying the mobile client device.

14. The system of claim 13 , wherein when validating the document the mobile client device is configured to identify a dimension similarity between the document and the document type of the document, based upon comparing a plurality of boundaries corresponding to a plurality of edges of the document in the first set of document features against a predefined dimension having a plurality of expected boundaries for the document type of the document.

15. The system of claim 13 , wherein the mobile client device is further configured to generate an indication of validation for presenting via a user interface of the mobile client device.

16. The system of claim 13 , wherein the mobile client device is further configured to generate an indication that the document validation score exceeds a warning threshold for presenting via a user interface of the mobile client device.

17. The system of claim 13 , wherein when validating the document the mobile client device is configured to:

compare the document validation score against at least one of a second validation threshold corresponding to non-validation and a warning threshold corresponding to an alert trigger; and

generate an output indicator for display at a user interface, based upon comparing the document validation score against the at least one of the second validation threshold or the warning threshold.

18. The system of claim 13 , wherein when validating the document the mobile client device is configured to identify at least one of a first set of validation criteria corresponding to clerical or imaging errors or a second set of validation criteria corresponding to digital or mechanical manipulation of the document.

19. The system of claim 13 , wherein the mobile client device is further configured to generate a spatial transform for the document responsive to detecting a plurality of edges of the document of the first set of document features.

20. The system of claim 13 , wherein the mobile client device is further configured to transmit the operation instruction to the backend server via one or more networks, the operation instruction including the packaged image and device metadata identifying the mobile client device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 17, 2025
From: SUREKA, ASHUTOSH K.; ADIMATYAM, VENKATA SESHA KIRAN KUMAR; SILVER, MIRIAM; FUNKEN, DANIEL
To: CITIBANK, N.A.
Reel/Frame 070873/0944 →
Continuity (1)
Continuation In Part 18629259 · Apr 8, 2024
References Cited (63)
US 8688579B1 · Ethington · 2014 [cited by examiner]
US 9137417B2 · Macciola et al. · 2015 [cited by applicant]
US 9672510B2 · Roach et al. · 2017 [cited by applicant]
US 10019772B1 · Smith · 2018 [cited by applicant]
US 10115031B1 · Pashinstev et al. · 2018 [cited by applicant]
US 10339374B1 · Pribble et al. · 2019 [cited by applicant]
US 10635898B1 · Pribble · 2020 [cited by applicant]
US 10803431B2 · Hinski · 2020 [cited by applicant]
US 11068976B1 · Voutour et al. · 2021 [cited by applicant]
US 11321679B1 · Bueche et al. · 2022 [cited by applicant]
US 11900755B1 · Bueche · 2024 [cited by applicant]
US 12039504B1 · Foster et al. · 2024 [cited by applicant]
US 12056434B2 · Bhatia et al. · 2024 [cited by applicant]
US 12106590B1 · Kinsey · 2024 [cited by applicant]
US 12272111B1 · Sureka et al. · 2025 [cited by applicant]
US 20070233615A1 · Tumminaro · 2007 [cited by applicant]
US 20110091092A1 · Nepomniachtchi et al. · 2011 [cited by applicant]
US 20110243459A1 · Deng · 2011 [cited by applicant]
US 20120230577A1 · Calman · 2012 [cited by examiner]
US 20130022231A1 · Nepomniachtchi et al. · 2013 [cited by applicant]
US 20140270464A1 · Nepomniachtchi et al. · 2014 [cited by applicant]
US 20150003666A1 · Wang et al. · 2015 [cited by applicant]
US 20150032631A1 · Hinski · 2015 [cited by applicant]
US 20150063653A1 · Madhani et al. · 2015 [cited by applicant]
US 20150120572A1 · Slade · 2015 [cited by applicant]
US 20150278819A1 · Song et al. · 2015 [cited by applicant]
US 20150294523A1 · Smith · 2015 [cited by applicant]
US 20150302242A1 · Lee et al. · 2015 [cited by applicant]
US 20150309966A1 · Gupta et al. · 2015 [cited by applicant]
US 20160037071A1 · Emmett et al. · 2016 [cited by applicant]
US 20160125613A1 · Shustorovich et al. · 2016 [cited by applicant]
US 20160253569A1 · Eid et al. · 2016 [cited by applicant]
US 20170116494A1 · Isaev · 2017 [cited by applicant]
US 20170185833A1 · Wang et al. · 2017 [cited by applicant]
US 20180211243A1 · Ekpenyong et al. · 2018 [cited by applicant]
US 20180330342A1 · Prakash et al. · 2018 [cited by applicant]
US 20180376193A1 · Sullivan et al. · 2018 [cited by applicant]
US 20190019020A1 · Flament et al. · 2019 [cited by applicant]
US 20190197693A1 · Zagaynov et al. · 2019 [cited by applicant]
US 20190213408A1 · Cali et al. · 2019 [cited by applicant]
US 20200279138A1 · Xu et al. · 2020 [cited by applicant]
US 20200410291A1 · Kriegman et al. · 2020 [cited by applicant]
US 20210124919A1 · Balakrishnan · 2021 [cited by examiner]
US 20210350516A1 · Tang et al. · 2021 [cited by applicant]
US 20210360149A1 · Mukul · 2021 [cited by applicant]
US 20210365677A1 · Anzenberg · 2021 [cited by examiner]
US 20220224816A1 · Pribble et al. · 2022 [cited by applicant]
US 20220358575A1 · Smith · 2022 [cited by applicant]
US 20220414955A1 · Ota · 2022 [cited by applicant]
US 20230030792A1 · Zheng et al. · 2023 [cited by applicant]
US 20230143239A1 · Yusuf et al. · 2023 [cited by applicant]
US 20230281629A1 · Shevyrev et al. · 2023 [cited by applicant]
US 20230281820A1 · Pizzocchero et al. · 2023 [cited by applicant]
US 20230298370A1 · Nishioka · 2023 [cited by applicant]
US 20240061992A1 · Bhatia et al. · 2024 [cited by applicant]
US 20240176951A1 · Krishnamoorthy · 2024 [cited by examiner]
US 20240256955A1 · Goodsitt et al. · 2024 [cited by applicant]
US 20240303658A1 · Cohen et al. · 2024 [cited by applicant]
US 20240428550A1 · Gutierrez Valdes et al. · 2024 [cited by applicant]
CA 3188665A1 · 2024 [cited by applicant]
Chernov, Timofey S., Sergey A llyuhin, and Vladimir V. Arlazarov. “Application of dynamic saliency maps to the video stream recognition systems with image quality assessment.” Eleventh International Conference on Machin… [cited by applicant]
Rybakova et al., “PESAC, the Generalized Framework for RANSAC-Based Methods on SIMD Computing Platforms”, IEEE Access, pp. 82151-82166, 2023, vol. 1. [cited by applicant]
PCT International Search Report and Written Opinion for Applicaiton No. PCT/US2025/023145 mailing date Jun. 26, 2025, 34 pages. [cited by applicant]
Cited By (11)
US 12,541,989 US 12,586,402 US 12,608,537 US 12,608,743 US 12,676,026 US 12,682,011 US 12,718,605 US 12,725,268 US 12,725,455 US 12,738,084 US 12,743,759