IP Library Granted Patent US 12,008,464
Granted Patent B2
US 12,008,464 · App. 15/815,635 · Granted Jun 11, 2024

Neural network based face detection and landmark localization

Inventors: Haoxiang Li (San Jose, CA); Zhe Lin (Fremont, CA); Jonathan Brandt (Santa Cruz, CA); Xiaohui Shen (San Jose, CA)
Assignee: ADOBE INC.
G06N3/08G06F3/04812G06F18/24143G06N3/045G06T15/04G06T15/205G06V10/454G06V10/764G06V10/82G06V40/165G06V40/171
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,008,464
App. No.
15/815,635
Granted
Jun 11, 2024
Kind
B2
Abstract

Approaches are described for determining facial landmarks in images. An input image is provided to at least one trained neural network that determines a face region (e.g., bounding box of a face) of the input image and initial facial landmark locations corresponding to the face region. The initial facial landmark locations are provided to a 3D face mapper that maps the initial facial landmark locations to a 3D face model. A set of facial landmark locations are determined from the 3D face model. The set of facial landmark locations are provided to a landmark location adjuster that adjusts positions of the set of facial landmark locations based on the input image. The input image is presented on a user device using the adjusted set of facial landmark locations.

Claims (40)

1. A method comprising:

jointly predicting, by a neural network, (1) a scale change or offset vector representing an adjustment to an initial bounding box of a face in an input image, the adjustment to the initial bounding box defining an adjusted bounding box of the face and (2) initial facial landmark locations of the face by outputting a representation of both the scale change or offset vector for the initial bounding box and the initial facial landmark locations from a common fully-connected layer of the neural network, each of the initial facial landmark locations corresponding to a two-dimensional point in the input image;

generating refined facial landmark locations in the input image from the initial facial landmark locations; and

causing presentation of a representation of the input image using the refined facial landmark locations.

2. The method of claim 1 , wherein generating the refined facial landmark locations comprises adjusting a first number of facial landmarks at the initial facial landmark locations to a second number of facial landmarks at the refined facial landmark locations, wherein the first number of facial landmarks is different than the second number of facial landmarks.

3. The method of claim 1 , wherein generating the refined facial landmark locations from the initial facial landmark locations comprises:

mapping the initial facial landmark locations to a 3D face model;

determining a set of facial landmark locations from the 3D face model; and

adjusting positions of the set of facial landmark locations in the input image to generate the refined facial landmark locations.

4. The method of claim 1 , further comprising classifying, using a face region classifier neural network, a candidate region of the input image as containing the face and providing a representation of a corresponding boundary of the candidate region as the representation of the first boundary, to the neural network, based on the candidate region being classified as containing the face.

5. The method of claim 1 , wherein predicting the adjusted bounding box comprises predicting confidence scores for calibration patterns representing components of an adjustment to an initial bounding box of the face.

6. The method of claim 1 , wherein causing the presentation of the representation of the input image using the refined facial landmark locations comprises applying image processing to the face represented by the refined facial landmark locations.

7. The method of claim 1 , wherein generating the refined facial landmark locations comprises predicting, for each of the refined facial landmark locations, a respective two-dimensional point in the input image.

8. One or more non-transitory computer-readable media having a plurality of executable instructions embodied thereon, which, when executed by one or more processors, cause the one or more processors to perform a method comprising:

jointly predicting, by a neural network, an adjusted bounding box of a face in an input image and initial facial landmark locations by outputting a representation of both the adjusted bounding box and the initial facial landmark locations from a common fully-connected layer of the neural network, each of the initial facial landmark locations corresponding to a two-dimensional point in the input image;

mapping initial facial landmark locations to a 3D face model;

determining a second set of predicted facial landmark locations in the input image from the 3D face model;

identifying refined facial landmark locations in the input image by adjusting positions of the second set of predicted facial landmark locations; and

causing presentation of a representation of the input image using the refined facial landmark locations.

9. The one or more non-transitory computer-readable media of claim 8 , wherein the at least one neural network comprises:

a face detector configured to classify a candidate region of the input image as containing the face; and

a subsequent joint calibration and alignment neural network configured to jointly predict, based on the candidate region being classified as containing the face, the adjusted bounding box of the face and the initial facial landmark locations of the face.

10. The one or more non-transitory computer-readable media of claim 8 , wherein determining the second set of predicted facial landmark locations comprises:

projecting the 3D face model into two-dimensions based on a pose of a face in the input image; and

determining the second set of predicted facial landmark locations from the projected 3D face model.

11. The one or more non-transitory computer-readable media of claim 8 , wherein determining the second set of predicted facial landmark locations comprises identifying a different number of facial landmarks locations than the initial facial landmark locations.

12. The one or more non-transitory computer-readable media of claim 8 , wherein the neural network is configured to predict the adjusted bounding box as a scale change or offset vector applied to an initial bounding box of the face.

13. The one or more non-transitory computer-readable media of claim 8 , wherein causing the presentation of the representation of the input image comprises applying image processing to the face represented by the refined facial landmark locations.

14. A system comprising:

a model manager configured to use one or more hardware processors to:

use a joint calibration and alignment neural network to jointly predict (1) an adjusted bounding box of a face in an input image represented by a predicted scale change or offset vector for an initial bounding box of the face and (2) initial facial landmark locations of the face, each of the initial facial landmark locations corresponding to a two-dimensional point in the input image; and

provide the initial facial landmark locations to a landmark location refiner configured to generate refined facial landmark locations in the input image from the initial facial landmark locations; and

a presentation component configured to use the one or more hardware processors to cause presentation of a representation of the input image using the refined facial landmark locations.

15. The system of claim 14 , further comprising an image processor configured to use the one or more hardware processors to process the input image using the refined facial landmark locations to generate a processed input image, wherein the presentation component is configured to cause the presentation of the processed input image.

16. The system of claim 14 , wherein the joint calibration and alignment neural network includes a common fully connected layer configured to generate a representation of both the adjusted bounding box and the initial facial landmark locations.

17. The system of claim 14 , wherein the landmark location refiner comprises a landmark location adjuster configured to use the one or more hardware processors to adjust positions of a set of facial landmark locations corresponding to the initial facial landmark locations to generate the refined facial landmark locations.

18. The system of claim 14 , wherein the landmark location refiner is further configured to use the one or more hardware processors to:

map the initial facial landmark locations to a 3D face model;

determine a set of facial landmark locations from the 3D face model; and

adjust positions of the set of facial landmark locations in the input image to generate the refined facial landmark locations.

Assignments (2)
CHANGE OF NAME Recorded Nov 29, 2018
From: ADOBE SYSTEMS INCORPORATED
To: ADOBE INC.
Reel/Frame 047687/0115 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 17, 2017
From: LI, HAOXIANG; LIN, ZHE; SHEN, XIAOHUI; BRANDT, JONATHAN
To: ADOBE SYSTEMS INCORPORATED
Reel/Frame 044160/0122 →
Continuity (1)
Related Publication 20190147224A1 · May 16, 2019
Cited By (1)
US 12,444,232