IP Library Granted Patent US 12,456,322
Granted Patent B2
US 12,456,322 · App. 19/218,062 · Granted Oct 28, 2025

Machine-learning models for image processing

Inventors: Ashutosh K. Sureka (Irving, TX); Venkata Sesha Kiran Kumar Adimatyam (Irving, TX); Miriam Silver (Tel Aviv, IL); Daniel Funken (Irving, TX)
Assignee: CITIBANK, N.A.
G06V30/414G06T7/0002G06V10/44G06V10/761G06T2207/30168G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,456,322
App. No.
19/218,062
Granted
Oct 28, 2025
Kind
B2
Abstract

Presented herein are systems and methods for the employment of machine learning models for image processing as may be performed by computing devices associated with an end user. A method may include obtaining video data comprising a plurality of frames including a document of a document type. The method may include executing an object recognition engine of a machine-learning architecture using image data of the plurality of frames, the object recognition engine trained to detect edges of documents. The method may include identifying, based on the edge detection, a plurality of boundaries for the document. The method may include validating, based on the plurality of boundaries, the document as the document type. The method may include transmitting via one or more networks, to a computer remote from the computing device, responsive to the validation of the type of document, the image data for the plurality of frames depicting the document.

Claims (35)

1 . A method of client-side operation validations for remote processing of document imagery, the method comprising:

obtaining, by a mobile client device associated with an end-user, an operation request via a user interface of the mobile client device;

obtaining, by a camera of the mobile client device, image data including a document and environment imagery around the document;

executing, the mobile client device, an object recognition engine to extract a set of environment features from the environment imagery, the object recognition engine trained for detecting one or more edges of the document and detecting the set of environment features using corresponding training labels indicating expected training environment imagery;

generating, by the mobile client device, an operation validation score based upon an image similarity between the set of environment features and expected environment imagery, the operation validation score indicating a likelihood that the operation request is a valid operation request according to the image similarity between the set of environment features and the expected environment imagery; and

in response to determining that the operation validation score satisfies an operation validation threshold, transmitting, by the mobile client device, an operation instruction for performing the operation request to a backend server.

2 . The method according to claim 1 , further comprising generating, by the mobile client device, the operation instruction including operation information and the image data having the document.

3 . The method according to claim 1 , further comprising obtaining, by the camera of the mobile client device, video data comprising a plurality of frames, including a frame having the image data including the environment imagery around the document.

4 . The method according to claim 3 , wherein the image data of a preceding frame of the video data indicates the expected environment imagery.

5 . The method according to claim 3 , further comprising executing, the mobile client device, the object recognition engine using the video data as input to identify each frame of the plurality of frames having a portion of the image data containing the document.

6 . The method according to claim 1 , further comprising identifying, by the mobile client device, a dimension similarity between the document and a document type of the document, based upon comparing the one or more edges of the document in a set of document dimension features against a predefined set of dimension features for the document type of the document.

7 . The method according to claim 1 , further comprising executing, by the mobile client device, the object recognition engine to extract a set of content features representing content data of the document from the image data.

8 . The method according to claim 7 , further comprising identifying, by the mobile client device, a content similarity between the content data of the document as extracted from the image data and expected content data of a predefined set of content features for a document type of the document.

9 . The method according to claim 7 , further comprising identifying, by the mobile client device, a content similarity between the content data of the document as extracted from the image data and expected content data received via the user interface of the mobile client device.

10 . The method according to claim 1 , further comprising:

generating, by the mobile client device, a quality score for the document using at least the set of environment features of the image data; and

generating, by the mobile client device, an output indicator for display at the user interface of the mobile client device based upon based upon comparing the quality score against a quality threshold.

11 . A system for client-side operation validations for remote processing document imagery, the system comprising:

a mobile client device associated with an end-user comprising at least one processor and a camera, configured to:

obtain an operation request via a user interface of the mobile client device;

obtain, by the camera, image data including a document and environment imagery around the document;

execute an object recognition engine to extract a set of environment features from the environment imagery, the object recognition engine trained for detecting one or more edges of the document and detecting the set of environment features using corresponding training labels indicating expected training environment imagery;

generate an operation validation score based upon an image similarity between the set of environment features and expected environment imagery, the operation validation score indicating a likelihood that the operation request is a valid operation request according to the image similarity between the set of environment features and the expected environment imagery; and

in response to determining that the operation validation score satisfies an operation validation threshold, transmitting an operation instruction for performing the operation request to a backend server.

12 . The system according to claim 11 , wherein the mobile device is further configured to generate the operation instruction including operation information and the image data having the document.

13 . The system according to claim 11 , wherein the mobile device is further configured to obtain, by the camera, video data comprising a plurality of frames, including a frame having the image data including the environment imagery around the document.

14 . The system according to claim 13 , wherein the image data of a preceding frame of the video data indicates the expected environment imagery.

15 . The system according to claim 13 , wherein the mobile device is further configured to execute the object recognition engine using the video data as input to identify each frame of the plurality of frames having a portion of the image data containing the document.

16 . The system according to claim 11 , wherein the mobile device is further configured to identify a dimension similarity between the document and a document type of the document, based upon comparing the one or more edges of the document in a set of document dimension features against a predefined set of dimension features for the document type of the document.

17 . The system according to claim 11 , wherein the mobile device is further configured to execute the object recognition engine to extract a set of content features representing content data of the document from the image data.

18 . The system according to claim 17 , wherein the mobile device is further configured to identify a content similarity between the content data of the document as extracted from the image data and expected content data of a predefined set of content features for a document type of the document.

19 . The system according to claim 17 , wherein the mobile device is further configured to identify a content similarity between the content data of the document as extracted from the image data and expected content data received via the user interface of the mobile client device.

20 . The system according to claim 11 , wherein the mobile device is further configured to:

generate a quality score for the document using at least the set of environment features of the image data; and

generate an output indicator for display at the user interface of the mobile client device based upon based upon comparing the quality score against a quality threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 23, 2025
From: SUREKA, ASHUTOSH K.; ADIMATYAM, VENKATA SESHA KIRAN KUMAR; SILVER, MIRIAM; FUNKEN, DANIEL
To: CITIBANK, N.A.
Reel/Frame 071213/0790 →
Continuity (3)
Continuation 18943565 · Nov 11, 2024
Continuation In Part 18629259 · Apr 8, 2024
Related Publication 20250316108A1 · Oct 9, 2025
References Cited (74)
US 8688579B1 · Ethington et al. · 2014 [cited by applicant]
US 9137417B2 · Macciola et al. · 2015 [cited by applicant]
US 9449312B1 · Wilson et al. · 2016 [cited by applicant]
US 9672510B2 · Roach et al. · 2017 [cited by applicant]
US 10019772B1 · Smith · 2018 [cited by examiner]
US 10115031B1 · Pashinstev et al. · 2018 [cited by applicant]
US 10339374B1 · Pribble · 2019 [cited by examiner]
US 10635898B1 · Pribble · 2020 [cited by applicant]
US 10803431B2 · Hinski · 2020 [cited by applicant]
US 11068976B1 · Voutour · 2021 [cited by examiner]
US 11321679B1 · Bueche et al. · 2022 [cited by applicant]
US 11900755B1 · Bueche · 2024 [cited by applicant]
US 12039504B1 · Foster et al. · 2024 [cited by applicant]
US 12056434B2 · Bhatia et al. · 2024 [cited by applicant]
US 12106590B1 · Kinsey · 2024 [cited by examiner]
US 12272111B1 · Sureka et al. · 2025 [cited by applicant]
US 20070233615A1 · Tumminaro · 2007 [cited by applicant]
US 20110091092A1 · Nepomniachtchi · 2011 [cited by examiner]
US 20110243459A1 · Deng · 2011 [cited by applicant]
US 20120230577A1 · Calman et al. · 2012 [cited by applicant]
US 20130022231A1 · Nepomniachtchi et al. · 2013 [cited by applicant]
US 20130070847A1 · Oami et al. · 2013 [cited by applicant]
US 20140270464A1 · Nepomniachtchi · 2014 [cited by examiner]
US 20150003666A1 · Wang · 2015 [cited by examiner]
US 20150032631A1 · Hinski · 2015 [cited by applicant]
US 20150063653A1 · Madhani · 2015 [cited by examiner]
US 20150120572A1 · Slade · 2015 [cited by applicant]
US 20150278819A1 · Song et al. · 2015 [cited by applicant]
US 20150294523A1 · Smith · 2015 [cited by applicant]
US 20150302242A1 · Lee et al. · 2015 [cited by applicant]
US 20150309966A1 · Gupta et al. · 2015 [cited by applicant]
US 20160037071A1 · Emmett et al. · 2016 [cited by applicant]
US 20160125613A1 · Shustorovich · 2016 [cited by examiner]
US 20160253569A1 · Eid et al. · 2016 [cited by applicant]
US 20170116494A1 · Isaev · 2017 [cited by examiner]
US 20170185833A1 · Wang et al. · 2017 [cited by applicant]
US 20180211243A1 · Ekpenyong et al. · 2018 [cited by applicant]
US 20180330342A1 · Prakash et al. · 2018 [cited by applicant]
US 20180376193A1 · Sullivan et al. · 2018 [cited by applicant]
US 20190019020A1 · Flament et al. · 2019 [cited by applicant]
US 20190197693A1 · Zagaynov et al. · 2019 [cited by applicant]
US 20190213408A1 · Cali · 2019 [cited by examiner]
US 20200279138A1 · Xu et al. · 2020 [cited by applicant]
US 20200410291A1 · Kriegman et al. · 2020 [cited by applicant]
US 20210124919A1 · Balakrishnan et al. · 2021 [cited by applicant]
US 20210182547A1 · Ayyadevara et al. · 2021 [cited by applicant]
US 20210350516A1 · Tang et al. · 2021 [cited by applicant]
US 20210360149A1 · Mukul · 2021 [cited by examiner]
US 20210365677A1 · Anzenberg · 2021 [cited by applicant]
US 20220224816A1 · Pribble et al. · 2022 [cited by applicant]
US 20220358575A1 · Smith · 2022 [cited by examiner]
US 20220414955A1 · Ota · 2022 [cited by applicant]
US 20230030792A1 · Zheng et al. · 2023 [cited by applicant]
US 20230143239A1 · Yusuf et al. · 2023 [cited by applicant]
US 20230281629A1 · Shevyrev et al. · 2023 [cited by applicant]
US 20230281820A1 · Pizzocchero et al. · 2023 [cited by applicant]
US 20230298370A1 · Nishioka · 2023 [cited by applicant]
US 20240061992A1 · Bhatia et al. · 2024 [cited by applicant]
US 20240176951A1 · Krishnamoorthy · 2024 [cited by applicant]
US 20240177487A1 · Sohoni · 2024 [cited by applicant]
US 20240256955A1 · Goodsitt et al. · 2024 [cited by applicant]
US 20240303658A1 · Cohen · 2024 [cited by examiner]
US 20240428550A1 · Gutierrez Valdes et al. · 2024 [cited by applicant]
CA 3188665A1 · 2024 [cited by applicant]
PCT International Search Report and Written Opinion for Application No. PCT/US2025/023145 mailing date Jun. 26, 2025, 34 pages. [cited by applicant]
Kada, Oumayma, et al. “Hologram detection for identity document authentication.” International Conference on Pattern Recognition and Artificial Intelligence. Cham: Springer International Publishing, 2022. (Year: 2022). [cited by applicant]
Koliaskina, L. I., et al. “MIDV-Holo: A dataset for ID document hologram detection in a video stream.” International Conference on Document Analysis and Recognition. Cham: Springer Nature Switzerland, 2023. (Year: 2023). [cited by applicant]
Piatrikova et al., “Digital Verification of Optically Variable Ink Feature on Identity Cards,” 2023 33rd Conference of Open Innovations Association (FRUCT). IEEE, 2023. (Year: 2023). [cited by applicant]
Pouliquen, Glen, et al. “Weakly Supervised Training for Hologram Verification in Identity Documents.” International Conference on Document Analysis and Recognition. Cham: Springer Nature Switzerland, 2024. (Year: 2024). [cited by applicant]
Soukup et al., “Mobile gram verification with deep learning,” IPSJ Transactions on Computer Vision and Applications 9 (2017): 1-6. (Year: 2017). [cited by applicant]
Chernov, Timofey S., Sergey A Ilyuhin, and Vladimir V. Arlazarov. “Application of dynamic saliency maps to the video stream recognition systems with image quality assessment.” Eleventh International Conference on Machin… [cited by applicant]
Rybakova et al., “PESAC, the Generalized Framework for RANSAC-Based Methods on SIMD Computing Platforms”, IEEE Access, pp. 82151-82166, 2023, vol. 1. [cited by applicant]
EPO Extended European Search Report for Application No. 25168952.7 mailing date Sep. 3, 2025, 7 pages. [cited by applicant]
EPO Extended European Search Report for Application No. 25168960.0 mailing date Sep. 3, 2025, 7 pages. [cited by applicant]