IP Library Granted Patent US 10,783,394
Granted Patent B2
US 10,783,394 · App. 16/006,728 · Granted Sep 22, 2020

Equivariant landmark transformation for landmark localization

Inventors: Pavlo Molchanov (San Jose, CA); Stephen Walter Tyree (St. Louis, MO); Jan Kautz (Lexington, MA); Sina Honari (Hampstead, CA)
Assignee: NVIDIA Corporation
G06K9/46G06K9/00234G06K9/00248G06K9/00281G06K9/00302G06K9/4628G06K9/6256G06K9/6279G06K9/66G06N3/0454G06N3/08G06N3/084G06N3/088G06N3/0445
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,783,394
App. No.
16/006,728
Granted
Sep 22, 2020
Kind
B2
Abstract

A method, computer readable medium, and system are disclosed to generate coordinates of landmarks within images. The landmark locations may be identified on an image of a human face and used for emotion recognition, face identity verification, eye gaze tracking, pose estimation, etc. A transform is applied to input image data to produce transformed input image data. The transform is also applied to predicted coordinates for landmarks of the input image data to produce transformed predicted coordinates. A neural network model processes the transformed input image data to generate additional landmarks of the transformed input image data and additional predicted coordinates for each one of the additional landmarks. Parameters of the neural network model are updated to reduce differences between the transformed predicted coordinates and the additional predicted coordinates.

Claims (47)

1. A computer-implemented method, comprising:

applying a transform to input image data to produce transformed input image data;

applying the transform to predicted coordinates for landmarks of the input image data to produce transformed predicted coordinates;

processing, by a neural network model, the transformed input image data to generate additional landmarks of the transformed input image data and additional predicted coordinates for each one of the additional landmarks;

updating parameters of the neural network model to produce a trained neural network model by reducing differences between the transformed predicted coordinates and the additional predicted coordinates; and

providing the trained neural network model for use in identifying landmark locations within unlabeled images.

2. The method of claim 1 , further comprising:

processing, by the neural network model, the input image data to generate pixel-level likelihood estimates for landmarks in the input image data; and

computing, by a soft-argmax function, the predicted coordinates of each landmark based on the pixel-level feature maps.

3. The method of claim 1 , further comprising:

processing, by the neural network model, the transformed input image data to generate pixel-level likelihood estimates for the additional landmarks; and

computing, by a soft-argmax function, the additional predicted coordinates based on the pixel-level likelihood estimates.

4. The method of claim 1 , further comprising updating parameters of the neural network model to reduce differences between the predicted coordinates and ground truth coordinates corresponding ground truth landmarks in the input image data.

5. The method of claim 1 , wherein the neural network model comprises at least one convolutional layer.

6. The method of claim 1 , wherein the transform is an affine transform.

7. A system, comprising:

a memory storing input image data; and

a processor coupled to the memory and configured to:

apply a transform to input image data to produce transformed input image data;

apply the transform to predicted coordinates for landmarks of the input image data to produce transformed predicted coordinates;

implement a neural network model for processing the transformed input image data to generate additional landmarks of the transformed input image data and additional predicted coordinates for each one of the additional landmarks;

update parameters of the neural network model to produce a trained neural network model by reducing differences between the transformed predicted coordinates and the additional predicted coordinates; and

provide the trained neural network model for use in identifying landmark locations within unlabeled images.

8. The system of claim 7 , wherein the processor is further configured to:

process, by the neural network model, the input image data to generate pixel-level likelihood estimates for landmarks in the input image data; and

compute, by a soft-argmax function, the predicted coordinates of each landmark based on the pixel-level feature maps.

9. The system of claim 7 , wherein the processor is further configured to:

process, by the neural network model, the transformed input image data to generate pixel-level likelihood estimates for the additional landmarks; and

compute, by a soft-argmax function, the additional predicted coordinates based on the pixel-level likelihood estimates.

10. The system of claim 7 , wherein the processor is further configured to update parameters of the neural network model to reduce differences between the predicted coordinates and ground truth coordinates corresponding ground truth landmarks in the input image data.

11. The system of claim 7 , wherein the neural network model comprises at least one convolutional layer.

12. The system of claim 7 , wherein the transform is an affine transform.

13. A non-transitory, computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to:

apply a transform to input image data to produce transformed input image data;

apply the transform to predicted coordinates for landmarks of the input image data to produce transformed predicted coordinates;

process, by a neural network model, the transformed input image data to generate additional landmarks of the transformed input image data and additional predicted coordinates for each one of the additional landmarks;

update parameters of the neural network model to produce a trained neural network model by reducing differences between the transformed predicted coordinates and the additional predicted coordinates; and

provide the trained neural network model for use in identifying landmark locations within unlabeled images.

14. The non-transitory, computer-readable storage medium of claim 13 , wherein the processor is further configured to:

process, by the neural network model, the input image data to generate pixel-level likelihood estimates for landmarks in the input image data; and

compute, by a soft-argmax function, the predicted coordinates of each landmark based on the pixel-level feature maps.

15. The non-transitory, computer-readable storage medium of claim 13 , wherein the processor is further configured to:

process, by the neural network model, the transformed input image data to generate pixel-level likelihood estimates for the additional landmarks; and

compute, by a soft-argmax function, the additional predicted coordinates based on the pixel-level likelihood estimates.

16. The non-transitory, computer-readable storage medium of claim 13 , wherein the processor is further configured to update parameters of the neural network model to reduce differences between the predicted coordinates and ground truth coordinates corresponding ground truth landmarks in the input image data.

17. The non-transitory, computer-readable storage medium of claim 13 , wherein the neural network model comprises at least one convolutional layer.

18. The non-transitory, computer-readable storage medium of claim 13 , wherein the transform is an affine transform.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 10, 2018
From: MOLCHANOV, PAVLO; TYREE, STEPHEN WALTER; KAUTZ, JAN; HONARI, SINA
To: NVIDIA CORPORATION
Reel/Frame 046832/0046 →
Continuity (2)
Provisional Application 62522520 · Jun 20, 2017
Related Publication 20180365512A1 · Dec 20, 2018
Cited By (4)
US 12,430,550 US 12,633,093 US 12,694,658 US 12,718,527