IP Library Granted Patent US 12,299,913
Granted Patent B2
US 12,299,913 · App. 17/823,033 · Granted May 13, 2025

Image processing framework for performing object depth estimation

Inventors: Kuang-Man Huang (Zhubei, TW); Michel Adib Sarkis (San Diego, CA)
Assignee: QUALCOMM Incorporated
G06T7/50G06T7/11G06T7/60G06T17/00G06T2200/08G06T2207/20132G06T2207/30201G06T2210/12G06T2210/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,299,913
App. No.
17/823,033
Granted
May 13, 2025
Kind
B2
Abstract

Disclosed are techniques for processing image data. In some aspects, a three-dimensional model can be determined corresponding to an object in an input image. Based on the three-dimensional model, an estimated focal length can be determined that corresponds to the input image. An estimated depth associated with the object in the input image can be calculated based on the estimated focal length and an input image focal length.

Claims (68)

1. A method for processing image data, the method comprising:

determining a three-dimensional model corresponding to an object in an input image;

determining, based on the three-dimensional model, an estimated focal length corresponding to the input image;

calculating an estimated depth associated with the object in the input image based on the estimated focal length and an input image focal length;

generating a projected two-dimensional image at least in part by projecting three-dimensional vertices of three-dimensional model to the input image;

determining an object bounding box in the projected two-dimensional image;

generating an attention mask based on the object bounding box;

determining, using a neural network, a depth bias value based on the attention mask and the input image; and

calculating, based on the estimated depth and the depth bias value, a corrected depth associated with the object in the input image.

2. The method of claim 1 , wherein the object corresponds to a face.

3. The method of claim 2 , further comprising:

determining a frontal face using the three-dimensional model; and

generating the projected two-dimensional image at least in part by projecting the frontal face to the input image.

4. The method of claim 1 , wherein the object bounding box is a face bounding box, and wherein the face bounding box is based on a two-dimensional distance between eyes of a face determined from the projected two-dimensional image.

5. The method of claim 4 , wherein the two-dimensional distance is based on a three-dimensional distance between the eyes determined from the three-dimensional model.

6. The method of claim 4 , wherein a center of the face bounding box corresponds to a nose of the face.

7. The method of claim 1 , wherein the attention mask is a binary mask including a first pixel value associated with each of a first plurality of pixels inside of the object bounding box and a second pixel value associated with each of a second plurality of pixels outside of the object bounding box.

8. The method of claim 1 , wherein the depth bias value is further based on one or more parameters from the three-dimensional model.

9. The method of claim 1 , wherein the depth bias value is based on a lower-resolution version of the input image.

10. The method of claim 1 , wherein the input image corresponds to a cropped image of the object from an original input image, and wherein the estimated depth is associated with the object in the original input image.

11. The method of claim 1 , further comprising:

determining a size of the object bounding box in the input image; and

determining the input image focal length based on the size of the object bounding box, a size of the input image, and a camera focal length, wherein the camera focal length is associated with a camera used to capture the input image.

12. An apparatus for processing image data, comprising:

at least one memory; and

at least one processor coupled to the at least one memory, the at least one processor configured to:

determine a three-dimensional model corresponding to an object in an input image;

determine, based on the three-dimensional model, an estimated focal length corresponding to the input image;

calculate an estimated depth associated with the object in the input image based on the estimated focal length and an input image focal length;

generate a projected two-dimensional image at least in part by projecting three-dimensional vertices of three-dimensional model to the input image;

determine an object bounding box in the projected two-dimensional image;

generate an attention mask based on the object bounding box;

determine, using a neural network, a depth bias value based on the attention mask and the input image; and

calculate, based on the estimated depth and the depth bias value, a corrected depth associated with the object in the input image.

13. The apparatus of claim 12 , wherein the object corresponds to a face.

14. The apparatus of claim 13 , wherein the at least one processor is configured to:

determine a frontal face using the three-dimensional model; and

generate the projected two-dimensional image at least in part by projecting the frontal face to the input image.

15. The apparatus of claim 12 , wherein the object bounding box is a face bounding box, and wherein the face bounding box is based on a two-dimensional distance between eyes of a face determined from the projected two-dimensional image.

16. The apparatus of claim 15 , wherein the two-dimensional distance is based on a three-dimensional distance between the eyes determined from the three-dimensional model.

17. The apparatus of claim 15 , wherein a center of the face bounding box corresponds to a nose of the face.

18. The apparatus of claim 12 , wherein the attention mask is a binary mask including a first pixel value associated with each of a first plurality of pixels inside of the object bounding box and a second pixel value associated with each of a second plurality of pixels outside of the object bounding box.

19. The apparatus of claim 12 , wherein the depth bias value is further based on one or more parameters from the three-dimensional model.

20. The apparatus of claim 12 , wherein the depth bias value is based on a lower-resolution version of the input image.

21. The apparatus of claim 12 , wherein the input image corresponds to a cropped image of the object from an original input image, and wherein the estimated depth is associated with the object in the original input image.

22. The apparatus of claim 12 , wherein the at least one processor is configured to:

determine a size of the object bounding box in the input image; and

determine the input image focal length based on the size of the object bounding box, a size of the input image, and a camera focal length, wherein the camera focal length is associated with a camera used to capture the input image.

23. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to:

determine a three-dimensional model corresponding to an object in an input image;

determine, based on the three-dimensional model, an estimated focal length corresponding to the input image;

calculate an estimated depth associated with the object in the input image based on the estimated focal length and an input image focal length;

generate a projected two-dimensional image at least in part by projecting three-dimensional vertices of three-dimensional model to the input image;

determine an object bounding box in the projected two-dimensional image;

generate an attention mask based on the object bounding box;

determine, using a neural network, a depth bias value based on the attention mask and the input image; and

calculate, based on the estimated depth and the depth bias value, a corrected depth associated with the object in the input image.

24. The non-transitory computer-readable medium of claim 23 , wherein the object corresponds to a face.

25. The non-transitory computer-readable medium of claim 24 , wherein the instructions, when executed by the one or more processors, cause the one or more processors to:

determine a frontal face using the three-dimensional model; and

generate the projected two-dimensional image at least in part by projecting the frontal face to the input image.

26. The non-transitory computer-readable medium of claim 23 , wherein the object bounding box is a face bounding box, and wherein the face bounding box is based on a two-dimensional distance between eyes of a face determined from the projected two-dimensional image.

27. The non-transitory computer-readable medium of claim 23 , wherein the instructions, when executed by the one or more processors, cause the one or more processors to:

determine a size of the object bounding box in the input image; and

determine the input image focal length based on the size of the object bounding box, a size of the input image, and a camera focal length, wherein the camera focal length is associated with a camera used to capture the input image.

28. The non-transitory computer-readable medium of claim 26 , wherein the two-dimensional distance is based on a three-dimensional distance between the eyes determined from the three-dimensional model.

29. The non-transitory computer-readable medium of claim 26 , wherein a center of the face bounding box corresponds to a nose of the face.

30. The non-transitory computer-readable medium of claim 23 , wherein the attention mask is a binary mask including a first pixel value associated with each of a first plurality of pixels inside of the object bounding box and a second pixel value associated with each of a second plurality of pixels outside of the object bounding box.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 23, 2022
From: HUANG, KUANG-MAN; SARKIS, MICHEL ADIB
To: QUALCOMM INCORPORATED
Reel/Frame 061199/0494 →
Continuity (2)
Provisional Application 63249514 · Sep 28, 2021
Related Publication 20230093827A1 · Mar 30, 2023
References Cited (17)
US 10027883B1 · Kuo · 2018 [cited by examiner]
US 20150055821A1 · Fotland · 2015 [cited by examiner]
US 20170069056A1 · Sachs et al. · 2017 [cited by applicant]
US 20180061018A1 · Ahn · 2018 [cited by examiner]
US 20200160070A1 · Sholingar · 2020 [cited by examiner]
US 20210166049A1 · Tariq · 2021 [cited by examiner]
WO WO2019147024A1 · 2019 [cited by examiner]
WO WO2021194490A1 · 2021 [cited by examiner]
Song et al, Beyond Trade-Off: Accelerate FCN-Based Face Detector with Higher Accuracy, 2018, IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1-10. (Year: 2018). [cited by examiner]
Zhang et al, Feature map masking based single-stage face detection, 2020, IEEE International Joint Conference on Biometrics, pp. 1-8. (Year: 2020). [cited by examiner]
Kar et al, Amodal Completion and Size Constancy in Natural Scenes, 2015, IEEE International Conference on Computer Vision, pp. 1-10. (Year: 2015). [cited by examiner]
Chen et al, Computational efficient deep neural network with difference attention maps for facial action unit detection, 2020, arXiv: 2011.12082v2, pp. 1-37. (Year: 2020). [cited by examiner]
Borghi et al, Face-from-Depth for Head Pose Estimation on Depth Images, 2018, IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 596-609. (Year: 2018). [cited by examiner]
Flores A., et al., “Camera Distance from Face Images”, Jul. 29, 2013, SAT 2015 18th International Conference, Austin, TX, USA, Sep. 24-27, 2015, [Lecture Notes In Computer Science, Lect. Notes Computer], Springer, Berli… [cited by applicant]
He L., et al., “Learning Depth from Single Images with Deep Neural Network Embedding Focal Length”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Mar. 27, 2018, 14 pages, X… [cited by applicant]
International Search Report and Written Opinion—PCT/US2022/075668—ISA/EPO—Dec. 20, 2022. [cited by applicant]
Xiao S., et al., “Recurrent 3D-2D Dual Learning for Large-Pose Facial Landmark Detection”, 2017 IEEE International Conference On Computer Vision (ICCV), IEEE, Oct. 22, 2017, pp. 1642-1651, XP033283024, DOI:10.1109/ICCV.… [cited by applicant]