IP Library › Granted Patent US 12,277,792
Granted Patent B2
US 12,277,792 · App. 17/977,506 · Granted Apr 15, 2025

Localized anomaly detection in digital documents using machine learning techniques

Inventors: Atul Kumar (Bangalore, IN); Anamika Chatterjee (Kolkata, IN); Saurabh Jha (Austin, TX)
Assignee: Dell Products L.P.
G06V30/418G06T9/00G06V30/19093G06T3/4046G06V10/764
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,277,792
App. No.
17/977,506
Granted
Apr 15, 2025
Kind
B2
Abstract

Methods, apparatus, and processor-readable storage media for localized anomaly detection in digital documents are provided herein. An example computer-implemented method includes identifying at least one image in a digital document; processing the identified at least one image using at least an image compression algorithm; applying a first machine learning model to the at least one processed image, wherein the first machine learning model is trained to detect whether the at least one image comprises one or more modifications; in response to detecting that the at least one image comprises at least one modification, applying a second machine learning model to identify a location in the at least one image corresponding to the at least one modification; and generating an indication that identifies the location of the at least one modification in the at least one image.

Claims (45)

1. A computer-implemented method comprising:

identifying at least one image in a digital document;

processing the identified at least one image using at least an image compression algorithm;

applying a first machine learning model to the processed at least one image, wherein the first machine learning model is trained to detect whether the processed at least one image comprises one or more modifications;

in response to detecting that the processed at least one image comprises at least one modification, applying a second machine learning model to identify a location in the processed at least one image corresponding to the at least one modification; and

generating an indication that identifies the location of the at least one modification in the processed at least one image;

wherein the method is performed by at least one processing device comprising a processor coupled to a memory.

2. The computer-implemented method of claim 1 , wherein the first machine learning model comprises an unparameterized linear transformation process.

3. The computer-implemented method of claim 1 , wherein the image compression algorithm comprises:

generating a compressed version of the processed at least one image; and

calculating a difference between pixels of the processed at least one image and pixels of the compressed version of the processed at least one image.

4. The computer-implemented method of claim 1 , wherein the digital document is saved in a portable document format, and wherein the identifying comprises extracting the processed at least one image from the digital document in an image format.

5. The computer-implemented method of claim 1 , wherein the second machine learning model comprises a convolutional neural network that uses empty pixel values to increase an input scope of neurons for the convolutional neural network.

6. The computer-implemented method of claim 1 , wherein the second machine learning model is trained using a contiguous block dropout process.

7. The computer-implemented method of claim 1 , wherein the one or more modifications relate to one or more of: at least one additional image and text.

8. The computer-implemented method of claim 1 , wherein the digital document corresponds to a request for one or more of a process and a service.

9. The computer-implemented method of claim 8 , further comprising:

preventing the request from being processed in response to detecting that the processed at least one image comprises the at least one modification.

10. A non-transitory processor-readable storage medium having stored therein program code of one or more software programs, wherein the program code when executed by at least one processing device causes the at least one processing device:

to identify at least one image in a digital document;

to process the identified at least one image using at least an image compression algorithm;

to apply a first machine learning model to the at least one image, wherein the first machine learning model is trained to detect whether the processed at least one image comprises one or more modifications;

in response to detecting that the processed at least one image comprises at least one modification, to apply a second machine learning model to identify a location in the processed at least one image corresponding to the at least one modification; and

to generate an indication that identifies the location of the at least one modification in the processed at least one image.

11. The non-transitory processor-readable storage medium of claim 10 , wherein the first machine learning model comprises an unparameterized linear transformation process.

12. The non-transitory processor-readable storage medium of claim 10 , wherein the image compression algorithm comprises:

generating a compressed version of the processed at least one image; and

calculating a difference between pixels of the processed at least one image and pixels of the compressed version of the processed at least one image.

13. The non-transitory processor-readable storage medium of claim 10 , wherein the digital document is saved in a portable document format, and wherein the identifying comprises extracting the processed at least one image from the digital document in an image format.

14. The non-transitory processor-readable storage medium of claim 10 , wherein the second machine learning model comprises a convolutional neural network that uses empty pixel values to increase an input scope of neurons for the convolutional neural network.

15. The non-transitory processor-readable storage medium of claim 10 , wherein the second machine learning model is trained using a contiguous block dropout process.

16. An apparatus comprising:

at least one processing device comprising a processor coupled to a memory;

the at least one processing device being configured:

to identify at least one image in a digital document;

to process the identified at least one image using at least an image compression algorithm;

to apply a first machine learning model to the at least one image, wherein the first machine learning model is trained to detect whether the processed at least one image comprises one or more modifications;

in response to detecting that the processed at least one image comprises at least one modification, to apply a second machine learning model to identify a location in the processed at least one image corresponding to the at least one modification; and

to generate an indication that identifies the location of the at least one modification in the processed at least one image.

17. The apparatus of claim 16 , wherein the first machine learning model comprises an unparameterized linear transformation process.

18. The apparatus of claim 16 , wherein the image compression algorithm comprises:

generating a compressed version of the processed at least one image; and

calculating a difference between pixels of the processed at least one image and pixels of the compressed version of the processed at least one image.

19. The apparatus of claim 16 , wherein the digital document is saved in a portable document format, and wherein the identifying comprises extracting the processed at least one image from the digital document in an image format.

20. The apparatus of claim 16 , wherein the second machine learning model comprises a convolutional neural network that uses empty pixel values to increase an input scope of neurons for the convolutional neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2022
From: KUMAR, ATUL; CHATTERJEE, ANAMIKA; JHA, SAURABH
To: DELL PRODUCTS L.P.
Reel/Frame 061598/0812 →
Continuity (1)
Related Publication 20240144712A1 · May 2, 2024
References Cited (7)
US 11526990B1 · Tamir · 2022 [cited by examiner]
US 20170308734A1 · Chalom · 2017 [cited by examiner]
US 20230073357A1 · Tsunoda · 2023 [cited by examiner]
US 20230360430A1 · Qian · 2023 [cited by examiner]
US 20230386191A1 · Zhu · 2023 [cited by examiner]
Van Beusekom, Joost, et al., “Distortion measurement for automatic document verification,” 2011 International Conference on Document Analysis and Recognition (pp. 289-293), IEEE, Sep. 18, 2011. [cited by applicant]
Dosovitskiy, Alexey, et al. “An image is worth 16x16 words: Transformers for image recognition at scale.” arXiv preprint arXiv:2010.11929, available at: https://arxiv.org/pdf/2010.11929.pdf (last accessed Oct. 31, 2022)… [cited by applicant]