IP Library › Granted Patent US 11,481,683
Granted Patent B1
US 11,481,683 · App. 16/888,589 · Granted Oct 25, 2022

Machine learning models for direct homography regression for image rectification

Inventors: Kunwar Yashraj Singh (Redmond, WA); Joaquin Zepeda Salvatierra (Mercer Island, WA); Erhan Bas (Sammamish, WA); Vijay Mahadevan (Los Angeles, CA); Jonathan Wu (Seattle, WA); Rahul Bhotika (Bellevue, WA)
Assignee: Amazon Technologies, Inc.
G06N20/00G06F17/16G06N5/04G06T3/0006
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,481,683
App. No.
16/888,589
Filed
May 29, 2020
Granted
Oct 25, 2022
Kind
B1
Art Unit
2643
USPC
706/12
Abstract

Techniques for creating machine learning models for direct homography regression for image rectification are described. In certain embodiments, a training service trains an algorithm on a source view of a training image and a homography matrix of the training image into a machine learning model that generates a normalized homography matrix for an input of the source view. The normalized homography matrix may then be utilized to generate a target view of an image input into the machine learning model. The target view of the image may be used in a document processing pipeline for document images captured using cameras.

Claims (40)

1. A computer-implemented method comprising:

receiving, by a machine learning service of a provider network, a source view of a training image and a homography matrix of the training image;

training, by the machine learning service of the provider network, an algorithm on the source view of the training image and the homography matrix of the training image into a machine learning model that generates a normalized homography matrix for an input of a source view;

receiving an inference request for a source view of an input image from a user located outside the provider network;

generating, by the machine learning model, a normalized homography matrix based at least in part on the source view of the input image;

generating an inference response based at least in part on the normalized homography matrix for the source view of the input image; and

transmitting the inference response to a client application or to a storage location.

2. The computer-implemented method of claim 1 , wherein the generating the inference response based at least in part on the normalized homography matrix for the source view of the input image comprises performing a perspective correction on the input image with the normalized homography matrix for the source view of the input image to generate a target view of the input image.

3. The computer-implemented method of claim 1 , wherein the generating the inference response based at least in part on the normalized homography matrix for the source view of the input image comprises determining a non-normalized homography matrix for the source view of the input image from the normalized homography matrix.

4. A computer-implemented method comprising:

receiving a source view of a training image and homography data of the training image;

training an algorithm on the source view of the training image and the homography data of the training image into a machine learning model that generates normalized homography data for an input of a source view;

receiving an inference request for a source view of an input image;

generating, by the machine learning model, normalized homography data based at least in part on the source view of the input image;

generating an inference response based at least in part on the normalized homography data for the source view of the input image; and

transmitting the inference response to a client application or to a storage location.

5. The computer-implemented method of claim 4 , wherein the generating the inference response based at least in part on the normalized homography data for the source view of the input image comprises performing a perspective correction on the input image with the normalized homography data for the source view of the input image to generate a target view of the input image.

6. The computer-implemented method of claim 5 , wherein the inference response includes the target view of the input image.

7. The computer-implemented method of claim 4 , wherein the generating the inference response based at least in part on the normalized homography data for the source view of the input image comprises determining non-normalized homography data for the source view of the input image from the normalized homography data.

8. The computer-implemented method of claim 7 , wherein the inference response includes the non-normalized homography data for the source view of the input image.

9. The computer-implemented method of claim 4 , wherein the generating the normalized homography data based at least in part on the source view of the input image does not rely on detecting one or more key points of the input image.

10. The computer-implemented method of claim 4 , wherein the normalized homography data based at least in part on the source view of the input image includes rotation, translation, scale, perspective, and shear components.

11. The computer-implemented method of claim 4 , further comprising performing a transform on the source view of the training image to create a centered source view of the training image, wherein the training the algorithm is on the centered source view of the training image.

12. The computer-implemented method of claim 11 , wherein the performing the transform comprises locating an origin of a coordinate system of the normalized homography data at an approximate center of the centered source view.

13. The computer-implemented method of claim 4 , wherein the homography data of the training image is normalized homography data of the training image, and the training the algorithm includes training the algorithm on the normalized homography data of the training image.

14. The computer-implemented method of claim 4 , wherein the homography data of the training image is non-normalized homography data of the training image, and the training the algorithm includes training the algorithm on the non-normalized homography data of the training image.

15. A system comprising:

a first one or more electronic devices to implement a storage service in a multi-tenant provider network, the storage service to store a source view of a training image and homography data of the training image; and

a second one or more electronic devices to implement a machine learning service in the multi-tenant provider network, the machine learning service including instructions that upon execution cause the machine learning service to perform operations comprising:

receiving the source view of the training image and the homography data of the training image;

training an algorithm on the source view of the training image and the homography data of the training image into a machine learning model that generates normalized homography data for an input of a source view;

receiving an inference request for a source view of an input image;

generating, by the machine learning model, normalized homography data based at least in part on the source view of the input image;

generating an inference response based at least in part on the normalized homography data for the source view of the input image; and

transmitting the inference response to a client application or to a storage location.

16. The system of claim 15 , wherein the instructions upon execution cause the machine learning service to perform operations wherein the generating the inference response based at least in part on the normalized homography data for the source view of the input image comprises performing a perspective correction on the input image with the normalized homography data for the source view of the input image to generate a target view of the input image.

17. The system of claim 15 , wherein the instructions upon execution cause the machine learning service to perform operations wherein the generating the inference response based at least in part on the normalized homography data for the source view of the input image comprises determining non-normalized homography data for the source view of the input image from the normalized homography data.

18. The system of claim 15 , wherein the instructions upon execution cause the machine learning service to perform operations wherein the generating the normalized homography data based at least in part on the source view of the input image does not rely on detecting one or more key points of the input image.

19. The system of claim 15 , wherein the instructions upon execution cause the machine learning service to perform operations further comprising performing a transform on the source view of the training image to create a centered source view of the training image, wherein the training the algorithm is on the centered source view of the training image.

20. The system of claim 15 , wherein the homography data of the training image is normalized homography data of the training image, and the training the algorithm includes training the algorithm on the normalized homography data of the training image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2020
From: SINGH, KUNWAR YASHRAJ; ZEPEDA SALVATIERRA, JOAQUIN; BAS, ERHAN; MAHADEVAN, VIJAY; WU, JONATHAN; BHOTIKA, RAHUL
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 053779/0963 →
Cited By (1)
US 12,591,847