IP Library › Granted Patent US 12,244,918
Granted Patent B2
US 12,244,918 · App. 18/199,826 · Granted Mar 4, 2025

Adaptive multi-scale face and body detector

Inventors: Hamidreza Vaezi Joze (Redmond, WA); Zehua Wei (Seattle, WA)
Assignee: Microsoft Technology Licensing, LLC
H04N23/611G06F18/2163G06V10/32H04N23/69H04N23/695
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,244,918
App. No.
18/199,826
Granted
Mar 4, 2025
Kind
B2
Abstract

Systems and methods are provided for determining faces and bodies of people in an image by adaptively scaling images and by iteratively using a deep neural network for inferencing. A camera captures an image including faces and bodies of people. A face/body determiner determines faces and bodies of people appearing in the image by resizing the image into a predetermined pixel dimension as input to the deep neural network. A region cropper determines a crop region associated with a low level of confidence in detecting faces and bodies that are too small to determine with an acceptable level of confidence. The region cropper resizes the crop region into the predetermined pixel dimension as input to the deep neural network. The face and body determiner determines other faces and bodies appearing in the resized crop region. An aggregator aggregates locations of the determined faces and bodies in the image.

Claims (40)

1. A computer-implemented method, comprising:

obtaining an image;

determining a first object in the image using a first machine learning model, wherein the first machine learning model is trained to detect the first object and generate a level of confidence of detecting the first object as at least a part of a body in the image;

determining, based on a first level of confidence of detecting a second object in the image using the first machine learning model, a region within the image, wherein the region includes the second object, and the first object and the second object are distinct and without overlapping;

determining, based on a second level of confidence associated with detecting the second object as at least a part of another body in the region using a second machine learning model, the second object in the region; and

automatically updating, based on the second object, a zoom setting of a camera to zoom and follow the second object.

2. The computer-implemented method according to claim 1 , wherein the first object includes either a face or a body of a person.

3. The computer-implemented method according to claim 1 , wherein the first and machine learning model and the second machine learning model are identical.

4. The computer-implemented method according to claim 2 , wherein the first machine learning model includes a deep neural network.

5. The computer-implemented method according to claim 2 , wherein a location of the region is one of predetermined set of grid regions in the image.

6. The computer-implemented method according to claim 1 , wherein a size of the image is greater than a size of the region, and wherein the size of the image represents a set of number of pixels in horizontal and vertical directions as a pixel dimension of the image.

7. The computer-implemented method according to claim 6 , further comprising:

aggregating respective locations and sizes of the first object and the second object in the image, wherein the aggregating includes non-maximum suppression.

8. A system for determining objects in an image, the system comprising:

a processor; and

a memory storing computer-executable instructions that when executed by the processor cause the system to:

determine a first object in the image using a first machine learning model, wherein the first machine learning model is trained to detect the first object and generate a level of confidence of detecting the first object as at least a part of a body in the image;

determine, based on a first level of confidence of detecting a second object in the image using the first machine learning model, a region within the image, wherein the region includes the second object, and the first object and the second object are distinct and without overlapping;

determine, based on a second level of confidence associated with detecting the second object as at least a part of another body in the region using a second machine learning model, the second object in the region; and

automatically update, based on the second object, a setting of a camera, wherein the setting includes a zoom level of the camera to zoom and follow the second object; and

capture, based on the updated setting of the camera, another image.

9. The system of claim 8 , wherein the first object includes either a face or a body of a person.

10. The system of claim 9 , wherein the first machine learning model includes a deep neural network.

11. The system of claim 9 , wherein a location of the region is one of predetermined set of grid regions in the image.

12. The system of claim 9 , wherein the first machine learning model and the second machine learning model are identical.

13. The system of claim 9 , wherein a size of the image is greater than a size of the region and wherein the size of the image represents a set of number of pixels in horizontal and vertical directions as a pixel dimension of the image.

14. The system of claim 9 , the computer-executable instructions when executed by the processor further cause the system to:

aggregate respective locations and sizes of the first object and the second object in the image, wherein the aggregating includes non-maximum suppression.

15. A computer-implemented method, comprising:

capturing an image using a camera;

determining a first face in the image using a first machine learning model, wherein the first machine learning model is trained to detect a face of a person in an image and generate a level of confidence of detecting the face as at least a part of a body in the image;

determining, based on a first level of confidence of detecting a second face in the image using the first machine learning model, a region within the image, wherein the region includes the second face, and the first face and the second face are distinct and without overlapping;

determining, based on a second level of confidence associated with detecting the second face as at least a part of another body in the region using a second machine learning model, the second face in the region; and

automatically updating, based on the second face, a setting of the camera to zoom and follow the second face; and

capturing, based on the updated setting of the camera, another image.

16. The computer-implemented method of claim 15 , wherein the first machine learning model includes a deep neural network.

17. The computer-implemented method of claim 15 , wherein a location of the region is one of predetermined set of grid regions in the image.

18. The computer-implemented method of claim 15 , wherein the first machine learning model and the second machine learning model are identical.

19. The computer-implemented method of claim 15 , wherein a size of the image is greater than a size of the region, and wherein the size of the image represents a set of number of pixels in horizontal and vertical directions as a pixel dimension of the image.

20. The computer-implemented method of claim 15 , wherein the setting includes at least one of a position or a zoom level of the camera.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2023
From: VAEZI JOZE, HAMIDREZA; WEI, ZEHUA
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 063708/0514 →
Continuity (2)
Continuation 17512780 · Oct 28, 2021
Related Publication 20230291993A1 · Sep 14, 2023
References Cited (5)
US 8478347B2 · Kim · 2013 [cited by examiner]
US 11482045B1 · Kim · 2022 [cited by examiner]
US 20190130165A1 · Seshadri · 2019 [cited by examiner]
US 20210319254A1 · Yu · 2021 [cited by examiner]
US 20220076018A1 · Geiss · 2022 [cited by examiner]