IP Library Granted Patent US 11,587,321
Granted Patent B2
US 11,587,321 · App. 16/949,084 · Granted Feb 21, 2023

Enhanced person detection using face recognition and reinforced, segmented field inferencing

Inventors: Rommel Gabriel Childress, Jr. (Cedar Park, TX); Alain Elon Nimri (Austin, TX); Stephen Paul Schaefer (Cedar Park, TX)
Assignee: PLANTRONICS, INC.
G06V20/49G06V40/10G06V40/161H04N7/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,587,321
App. No.
16/949,084
Granted
Feb 21, 2023
Kind
B2
Abstract

The frame or image of a video stream of a videoconference is divided into a series of segments for analysis. There is a primary grid, which covers the entire frame, and an alternate grid, which is shifted from the primary grid. Each segment is small enough to allow a neural network to efficiently operate on the segment without requiring downsampling. By operating on full resolution images, a participant can be identified at a greater distance from the camera. The entire frame is analyzed at a lower frequency, such as once per five seconds, but each segment containing a participant in the conference is scanned at a higher frequency, such as once per second, to maintain responsiveness to participant movement but also allow the full resolution operation.

Claims (38)

1. A method for detecting participants in a videoconferencing video stream, the method comprising:

obtaining a video image from the video stream;

dividing the entire video image into a first plurality of segments based on a primary grid;

dividing the video image into a second plurality of segments based on an alternate grid that is offset from the primary grid;

performing first participant detection on the video image based on the first plurality of segments individually, wherein the first participant detection is indicative of a first count of participants in the video image; and

performing second participant detection on the video image based on the second plurality of segments individually, wherein the second participant detection is indicative of a second count of participants that is higher than the first count of participants, and wherein at least one participant detected by the second participant detection is not detected by the first participant detection based on the at least one participant being located next to an edge or a boundary of a segment in the first plurality of segments.

2. The method of claim 1 , wherein the video image has a resolution, and wherein the resolution of the first plurality of segments and the second plurality of segments is not reduced from the resolution of the video image.

3. The method of claim 1 , wherein the video image has a resolution, and wherein the resolution of the first plurality of segments and the second plurality of segments is reduced from the resolution of the video image.

4. The method of claim 1 , wherein the performing the first participant detection on the video image based on the first plurality of segments individually and the performing the second participant detection on the video image based on the second plurality of segments individually are performed to provide participant detection of the video image in a first period.

5. The method of claim 4 , wherein the performing the first participant detection on the video image based on first plurality of segments individually is performed for a segment containing a detected participant or having contained a detected participant in a second period, the second period being shorter than the first period.

6. The method of claim 1 , wherein the performing the first participant detection on the video image based on the first plurality of segments individually and the performing the second participant detection on the video image based on the second plurality of segments individually comprises performing face detection on each segment.

7. The method of claim 6 , wherein the performing the first participant detection on the video image based on the first plurality of segments individually and the performing the second participant detection on the video image based on the second plurality of segments individually comprises performing body detection on each segment if a face is not detected.

8. A videoconferencing device comprising:

a camera interface for receiving a videoconferencing video stream;

a processor for executing programs and operations to perform videoconferencing operations; and

memory coupled to the processor for storing programs executed by the processor, the memory storing programs executed by the processor to perform participant detection of the videoconferencing video stream by:

obtaining a video image from the video stream,

dividing the entire video image into a first plurality of segments based on a primary grid,

dividing the video image into a second plurality of segments based on an alternate grid that is offset from the primary grid,

performing first participant detection on the video image based on the first plurality of segments individually, wherein the first participant detection is indicative of a first count of participants in the video image, and

performing second participant detection on the video image based on the second plurality of segments individually, wherein the second participant detection is indicative of a second count of participants that is higher than the first count of participants, and wherein at least one participant detected by the second participant detection is not detected by the first participant detection based on the at least one participant being located next to an edge or a boundary of a segment in the first plurality of segments.

9. The videoconferencing device of claim 8 , wherein the video image has a resolution, and wherein the resolution of the first plurality of segments and the second plurality of segments is not reduced from the resolution of the video image.

10. The videoconferencing device of claim 8 , wherein the video image has a resolution, and wherein the resolution of the first plurality of segments and the second plurality of segments is reduced from the resolution of the video image.

11. The videoconferencing device of claim 8 , wherein the performing the first participant detection on the video image based on the first plurality of segments individually and the performing the second participant detection on the video image based on the second plurality of segments individually are performed to provide participant detection of the video image in a first period.

12. The videoconferencing device of claim 11 , wherein the performing the first participant detection on the video image based on the first plurality of segments individually is performed for a segment containing a detected participant or having contained a detected participant in a second period, the second period being shorter than the first period.

13. The videoconferencing device of claim 8 , wherein the performing the first participant detection on the video image based on the first plurality of segments individually and the performing the second participant detection on the video image based on the second plurality of segments individually comprises performing face detection on each segment.

14. The videoconferencing device of claim 13 , wherein the performing the first participant detection on the video image based on the first plurality of segments individually and the performing the second participant detection on the video image based on the second plurality of segments individually comprises performing body detection on each segment if a face is not detected.

15. A non-transitory processor readable memory containing programs that when executed cause a processor to perform the following method of detecting participants in a videoconferencing video stream, the method comprising:

obtaining a video image from the video stream;

dividing the entire video image into a first plurality of segments based on a primary grid;

dividing the video image into a second plurality of segments based on an alternate grid that is offset from the primary grid;

performing first participant detection on the video image based on the first plurality of segments individually, wherein the first participant detection is indicative of a first count of participants in the video image; and

performing second participant detection on the video image based on the second plurality of segments individually, wherein the second participant detection is indicative of a second count of participants that is higher than the first count of participants, and wherein at least one participant detected by the second participant detection is not detected by the first participant detection based on the at least one participant being located next to an edge or a boundary of a segment in the first plurality of segments.

16. The non-transitory processor readable memory of claim 15 , wherein the video image has a resolution, and wherein the resolution of the first plurality of segments and the second plurality of segments is not reduced from the resolution of the video image.

17. The non-transitory processor readable memory of claim 15 , wherein the performing the first participant detection on the video image based on the first plurality of segments individually and the performing the second participant detection on the video image based on the second plurality of segments individually are performed to provide participant detection of the video image in a first period.

18. The non-transitory processor readable memory of claim 17 , wherein the performing the first participant detection on the video image based on the first plurality of segments individually is performed for a segment containing a detected participant or having contained a detected participant in a second period, the second period being shorter than the first period.

19. The non-transitory processor readable memory of claim 15 , wherein the performing the first participant detection on the video image based on the first plurality of segments individually and the performing the second participant detection on the video image based on the second plurality of segments individually comprises performing face detection on each segment.

20. The non-transitory processor readable memory of claim 19 , wherein the performing the first participant detection on the video image based on the first plurality of segments individually and the performing the second participant detection on the video image based on the second plurality of segments individually comprises performing body detection on each segment if a face is not detected.

Assignments (4)
NUNC PRO TUNC ASSIGNMENT Recorded Nov 13, 2023
From: PLANTRONICS, INC.
To: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
Reel/Frame 065549/0065 →
RELEASE OF PATENT SECURITY INTERESTS Recorded Aug 30, 2022
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: PLANTRONICS, INC.; POLYCOM, INC.
Reel/Frame 061356/0366 →
SUPPLEMENTAL SECURITY AGREEMENT Recorded Oct 6, 2021
From: PLANTRONICS, INC.; POLYCOM, INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION
Reel/Frame 057723/0041 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 14, 2020
From: CHILDRESS, ROMMEL GABRIEL, JR; NIMRI, ALAIN ELON; SCHAEFER, STEPHEN PAUL
To: PLANTRONICS, INC.
Reel/Frame 054044/0669 →
Continuity (2)
Provisional Application 63009340 · Apr 13, 2020
Related Publication 20210319233A1 · Oct 14, 2021