IP Library Granted Patent US 11,392,800
Granted Patent B2
US 11,392,800 · App. 16/919,840 · Granted Jul 19, 2022

Computer vision systems and methods for blind localization of image forgery

Inventors: Aurobrata Ghosh (Pondicherry, IN); Zheng Zhong (Jersey City, NJ); Terrance E. Boult (Colorado Springs, CO); Maneesh Kumar Singh (Princeton, NJ)
Assignee: Insurance Services Office, Inc.
G06K9/6265G06K9/6257G06N3/04G06N3/08G06V10/30G06V10/50G06V20/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,392,800
App. No.
16/919,840
Granted
Jul 19, 2022
Kind
B2
Abstract

Computer vision systems and methods for localizing image forgery are provided. The system generates a constrained convolution via a plurality of learned rich filters. The system trains a convolutional neural network with the constrained convolution and a plurality of images of a dataset to learn a low level representation of each image among the plurality of images. The low level representation is indicative of a statistical signature of at least one source camera model of each image. The system can determine a splicing manipulation localization by the trained convolutional neural network.

Claims (51)

1. A computer vision system for localizing image forgery comprising:

a memory; and

a processor in communication with the memory, the processor:

generating a constrained convolution using a plurality of learned rich filters by computing residuals for the plurality of learned rich filters, each residual being a difference between a predicted value for a central pixel defined over a pixel neighborhood and a scaled value of a pixel;

training a neural network with the constrained convolution and a plurality of images of a dataset to learn a low-level representation indicative of a statistical signature of at least one source camera model for each image among the plurality of images, and

localizing an attribute of an image of the dataset by the trained neural network.

2. The system of claim 1 , wherein the processor:

extracts at least one noise residual pattern from each image among the plurality of images via the constrained convolution,

determines a spatial distribution of the extracted at least one noise residual pattern, and

suppresses semantic edges present in each image among the plurality of images by applying a probabilistic regularization,

generating a constrained convolution via a plurality of learned rich filters by computing residuals for the plurality of learned rich filters, each residual being a difference between a predicted value for a central pixel defined over a pixel neighborhood and a scaled value of a pixel;

training a neural network with the constrained convolution and a plurality of images of a dataset to learn a low-level representation indicative of a statistical signature of at least one source camera model for each image among the plurality of images, and

localizing an attribute of an image of the dataset by the trained neural network.

3. The system of claim 2 , wherein the processor trains the neural network with a complete loss function based on a cross-entropy loss function over the dataset, the probabilistic regularization, and a rich filter constraint penalty.

4. The system of claim 1 , wherein the processor localizes the attribute of the image of the dataset by the trained neural network by:

subdividing the image into a plurality of patches,

determining a hundred-dimensional feature vector for each patch, and

segmenting the plurality of patches by applying an expectation maximization algorithm to each patch to fit a two component Gaussian mixture model to each feature vector.

5. The system of claim 1 , wherein the neural network is an 18 layer deep Convolutional Neural Network (CNN).

6. The system of claim 1 , wherein the dataset is a Dresden Image dataset.

7. The system of claim 1 , wherein the localized adversarial perturbation of the image is a splicing manipulation.

8. A method for localizing image forgery by a computer vision system, comprising the steps of:

subdividing the image into a plurality of patches,

determining a hundred-dimensional feature vector for each patch, and

segmenting the plurality of patches by applying an expectation maximization algorithm to each patch to fit a two component Gaussian mixture model to each feature vector.

9. The method of claim 8 , further comprising:

extracting at least one noise residual pattern from each image among the plurality of images via the constrained convolution,

determining a spatial distribution of the extracted at least one noise residual pattern, and

suppressing semantic edges present in each image among the plurality of images by applying a probabilistic regularization.

10. The method of claim 9 , further comprising training the neural network with a complete loss function based on a cross-entropy loss function over the dataset, the probabilistic regularization, and a rich filter constraint penalty.

11. The method of claim 8 , further comprising localizing the attribute of the image of the dataset by the trained neural network by

subdividing the image into a plurality of patches,

determining a hundred-dimensional feature vector for each patch, and

segmenting the plurality of patches by applying an expectation maximization algorithm to each patch to fit a two component Gaussian mixture model to each feature vector.

12. The method of claim 8 , wherein the neural network is an 18 layer deep Convolutional Neural Network (CNN).

13. The method of claim 8 , wherein the localized adversarial perturbation of the image is a splicing manipulation.

14. A non-transitory computer readable medium having instructions stored thereon for localizing image forgery by a computer vision system which, when executed by a processor, causes the processor to carry out the steps of:

generating a constrained convolution via a plurality of learned rich filters by computing residuals for the plurality of learned rich filters, each residual being a difference between a predicted value for a central pixel defined over a pixel neighborhood and a scaled value of a pixel;

training a neural network with the constrained convolution and a plurality of images of a dataset to learn a low-level representation indicative of a statistical signature of at least one source camera model for each image among the plurality of images, and

localizing an attribute of an image of the dataset by the trained neural network.

15. The non-transitory computer readable medium of claim 14 , the processor further carrying out the steps of:

extracting at least one noise residual pattern from each image among the plurality of images via the constrained convolution,

determining a spatial distribution of the extracted at least one noise residual pattern, and

suppressing semantic edges present in each image among the plurality of images by applying a probabilistic regularization.

16. The non-transitory computer readable medium of claim 15 , the processor further carrying out the step of training the neural network with a complete loss function based on a cross-entropy loss function over the dataset, the probabilistic regularization, and a rich filter constraint penalty.

17. The non-transitory computer readable medium of claim 14 , the processor localizing the attribute of the image of the dataset by the trained neural network by carrying out the steps of:

subdividing the image into a plurality of patches,

determining a hundred-dimensional feature vector for each patch, and

segmenting the plurality of patches by applying an expectation maximization algorithm to each patch to fit a two component Gaussian mixture model to each feature vector.

18. The non-transitory computer readable medium of claim 14 , wherein the neural network is an 18 layer deep Convolutional Neural Network (CNN).

19. The non-transitory computer readable medium of claim 14 , wherein the localized adversarial perturbation of the image is a splicing manipulation.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 18, 2023
From: BOULT, TERRENCE E., DR.
To: THE REGENTS OF THE UNIVERSITY OF COLORADO
Reel/Frame 063362/0366 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2022
From: GHOSH, AUROBRATA; ZHONG, ZHENG; SINGH, MANEESH KUMAR
To: INSURANCE SERVICES OFFICE, INC.
Reel/Frame 060193/0498 →
Continuity (2)
Provisional Application 62869712 · Jul 2, 2019
Related Publication 20210004648A1 · Jan 7, 2021