IP Library Granted Patent US 11,775,834
Granted Patent B2
US 11,775,834 · App. 17/294,573 · Granted Oct 3, 2023

Joint upper-body and face detection using multi-task cascaded convolutional networks

Inventors: Hai Xu (Beijing, CN); Xi Lu (Beijing, CN); Yongkang Fan (Beijing, CN); Wenxue He (Beijing, CN)
Assignee: Polycom, LLC
G06N3/082G06N3/045G06N3/08G06V10/454G06V10/82G06V40/10G06V40/161H04N7/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,775,834
App. No.
17/294,573
Granted
Oct 3, 2023
Kind
B2
Abstract

A videoconferencing endpoint is described that uses a cascading sequence of convolutional neural networks to perform face detection and upper body detection of participants in a videoconference at the endpoint, where at least one member of the sequence of neural networks performs upper body detection, and where the final member of the sequence of neural networks performs face detection based on the results of the upper body detection. The models of the neural networks are trained on both large datasets of faces well as images that have been distorted by a wide-angle camera of the videoconferencing endpoint.

Claims (50)

1. A method of detecting faces and upper bodies of participants in a videoconference, comprising:

receiving video data from a camera of a videoconferencing endpoint;

performing upper body detection in the video data using a first neural network; and

performing face detection by a second neural network in areas of the video data identified by the upper body detection of the first neural network.

2. The method of claim 1 , further comprising:

performing upper body detection in the video data using a third neural network,

wherein performing upper body detection in the video data using the first neural network comprises performing upper body detection in areas of the video data identified as possible upper body areas by the third neural network.

3. The method of claim 1 , wherein receiving video data from the camera of the videoconferencing endpoint comprises receiving video data that is distorted by the camera of the videoconferencing endpoint.

4. The method of claim 1 , wherein the first neural network and the second neural network employ models that have been trained on both undistorted video images and distorted video images.

5. The method of claim 1 , wherein performing upper body detection comprises producing upper body bounding box information.

6. The method of claim 1 , wherein performing face detection comprises:

considering areas for face detection having a lower probability threshold than used by the first neural network for upper body detection.

7. The method of claim 1 , further comprising:

performing head detection by a fourth neural network in areas identified as upper bodies by the first neural network,

wherein performing face detection comprises performing face detection by the second neural network in areas identified as heads by the fourth neural network.

8. A video conferencing endpoint, comprising:

a housing;

a camera, disposed in the housing;

a processing unit, disposed in the housing and coupled to the camera;

a memory, disposed in the housing and coupled to the processing unit and the camera, in which are stored instructions for performing face detection and upper body detection, comprising instructions that when executed cause the processing unit to:

receive video data from the camera;

perform upper body detection in the video data using a first neural network; and

perform face detection by a second neural network in areas of the video data identified by the upper body detection of the first neural network.

9. The videoconferencing endpoint of claim 8 ,

wherein the instructions further comprise instructions that when executed cause the processing unit to:

perform upper body detection in the third video data using a third neural network, and

wherein the instructions that when executed cause the processing unit to perform upper body detection using the first neural network comprise instructions that when executed cause the processing unit to:

perform upper body detection in areas of the video data identified as possible upper body areas by the third neural network.

10. The videoconferencing endpoint of claim 8 , wherein the camera is a wide-angle camera producing distorted images.

11. The videoconferencing endpoint of claim 8 , wherein the first neural network and the second neural network employ models that have been trained on both undistorted video data and distorted video data.

12. The videoconferencing endpoint of claim 8 , wherein the instructions that when executed cause the processing unit to perform upper body detection comprise instructions that when executed cause the first neural network to generate upper body bounding box information.

13. The videoconferencing endpoint of claim 8 , wherein the instructions that when executed cause the processing unit to perform face detection comprise instructions to adjust a probability threshold consider areas for face detection having a lower probability of being an upper body than identified by the first neural network as containing upper bodies.

14. The videoconferencing endpoint of claim 8 , wherein the instructions further comprise instructions that when executed cause the processing unit to:

perform head detection by a fourth neural network in areas identified as upper bodies by the first neural network,

wherein the instructions that when executed cause the processing unit to perform face detection comprise instructions that when executed cause the processing unit to perform face detection in areas identifies as heads by the fourth neural network.

15. A non-transitory machine readable medium including instructions, that when executed cause a processing unit of a videoconferencing endpoint to perform the methods of:

receiving video data from a camera of a videoconferencing endpoint;

performing upper body detection in the video data using a first neural network; and

performing face detection by a second neural network in areas of the video data identified by the upper body detection of the first neural network.

16. The non-transitory machine readable medium of claim 15 , the method further comprising:

performing upper body detection in the video data using a third neural network,

wherein performing upper body detection in the video data using the first neural network comprises performing upper body detection in areas of the video data identified as possible upper body areas by the third neural network.

17. The non-transitory machine readable medium of claim 15 , wherein receiving video data from the camera of the videoconferencing endpoint comprises receiving video data that is distorted by the camera of the videoconferencing endpoint.

18. The non-transitory machine readable medium of claim 15 , wherein the first neural network and the second neural network employ models that have been trained on both undistorted video images and distorted video images.

19. The non-transitory machine readable medium of claim 15 , wherein performing upper body detection comprises producing upper body bounding box information.

20. The non-transitory machine readable medium of claim 15 , wherein performing face detection comprises:

considering areas for face detection having a lower probability threshold than used by the first neural network for upper body detection.

21. The non-transitory machine readable medium of claim 15 , the method further comprising:

performing head detection by a fourth neural network in areas identified as upper bodies by the first neural network,

wherein performing face detection comprises performing face detection by the second neural network in areas identified as heads by the fourth neural network.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY ADDRESS PREVIOUSLY RECORDED AT REEL: 063115 FRAME: 0558. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 21, 2023
From: POLYCOM, LLC.
To: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
Reel/Frame 066175/0381 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 20, 2023
From: XU, HAI; LU, XI; FAN, YONGKANG; HE, WENXUE
To: POLYCOM, INC.
Reel/Frame 064325/0052 →
CORRECTIVE ASSIGNMENT TO CORRECT THE CONVEYING AND RECEIVING PARTY NAMES PREVIOUSLY RECORDED AT REEL: 062699 FRAME: 0203. ASSIGNOR(S) HEREBY CONFIRMS THE CHANGE OF NAME. Recorded Mar 16, 2023
From: POLYCOM, INC.
To: POLYCOM, LLC
Reel/Frame 063115/0558 →
CHANGE OF NAME Recorded Feb 9, 2023
From: POLYCOMM, INC.
To: POLYCOMM, LLC
Reel/Frame 062699/0203 →
Continuity (1)
Related Publication 20210409645A1 · Dec 30, 2021