IP Library Granted Patent US 11,676,369
Granted Patent B2
US 11,676,369 · App. 17/129,145 · Granted Jun 13, 2023

Context based target framing in a teleconferencing environment

Inventors: Rommel Gabriel Childress, Jr. (Cedar Park, TX); Alain Elon Nimri (Austin, TX); Stephen Paul Schaefer (Cedar Park, TX); David Young (Austin, TX)
Assignee: Plantronics, Inc.
G06V10/764G06V10/82G06V20/46G06V40/161G06V40/171H04N7/04H04N7/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,676,369
App. No.
17/129,145
Granted
Jun 13, 2023
Kind
B2
Abstract

A method for determining camera framing in a teleconferencing system comprises a process loop which includes acquiring an audio-visual frame from a captured a video data frame; detecting objects and extracting image features of the objects within the video data frame, ingesting the audio-visual frame into a context-based audio-visual map in an intelligent manner, and selecting targets from within the map for inclusion in an audio-video stream for transmission to a remote endpoint.

Claims (58)

1. A video framing method in a teleconferencing system, the method comprising:

receiving a first audio-video stream;

deriving a first sound source location from a first plurality of audio frames of the first audio-video stream;

updating a first talker weight value of a first participant when the first sound source location corresponds to the first participant;

deriving a second sound source location from a second plurality of audio frames of the first audio-video stream;

updating a second talker weight value of a second participant when the second sound source location corresponds to the second participant;

updating a conversation weight value of the first participant and the second participant when the first talker weight value and the second talker weight value exceed a first threshold and a difference between the first talker weight value and the second talker weight value is less than a predetermined amount;

selecting a first video sub-frame depicting the first participant for transmission in a second audio-video stream to a remote endpoint when the first talker weight value exceeds a second threshold; and

selecting a second video sub-frame depicting the first participant and the second participant for transmission in the second audio-video stream to the remote endpoint when the conversation weight value exceeds a third threshold.

2. The method of claim 1 , further comprising:

selecting a third video sub-frame depicting the second participant for transmission in the second audio-video stream to the remote endpoint when the second talker weight value exceeds the second threshold.

3. The method of claim 1 , further comprising:

updating the conversation weight value of the first participant and the second participant when the first participant makes one or more hand gestures directed toward the second participant.

4. The method of claim 1 , further comprising:

updating the conversation weight value of the first participant and the second participant when the first participant looks toward the second participant.

5. The method of claim 4 , wherein updating the conversation weight value of the first participant and the second participant when the first participant looks toward the second participant comprises estimating at least one of a head pose or an eye gaze of the first participant.

6. The method of claim 5 , wherein estimating at least one of the head pose or the eye gaze of the first participant comprises analyzing, using an artificial neural network, a plurality of video sub-frames depicting the first participant.

7. The method of claim 6 , wherein the artificial neural network is a convolutional neural network.

8. A non-transitory computer readable medium storing instructions executable by a processor, wherein the instructions comprise instructions to:

receive a first audio-video stream at a first endpoint;

derive a first sound source location from a first plurality of audio frames of the first audio-video stream;

update a first talker weight value of a first participant when the first sound source location corresponds to the first participant;

derive a second sound source location from a second plurality of audio frames of the first audio-video stream;

update a second talker weight value of a second participant when the second sound source location corresponds to the second participant;

update a conversation weight value of the first participant and the second participant when the first talker weight value and the second talker weight value exceed a first threshold and a difference between the first talker weight value and the second talker weight value is less than a predetermined amount;

select a first video sub-frame depicting the first participant for transmission in a second audio-video stream to a second endpoint when the first talker weight value exceeds a second threshold; and

select a second video sub-frame depicting the first participant and the second participant for transmission in the second audio-video stream to the second endpoint when the conversation weight value exceeds a third threshold.

9. The non-transitory computer readable medium of claim 8 , wherein the instructions further comprise instructions to:

select a third video sub-frame depicting the second participant for transmission in the second audio-video stream to the second endpoint when the second talker weight value exceeds the second threshold.

10. The non-transitory computer readable medium of claim 8 , wherein the instructions further comprise instructions to:

update the conversation weight value of the first participant and the second participant when the first participant makes one or more hand gestures directed toward the second participant.

11. The non-transitory computer readable medium of claim 8 , wherein the instructions further comprise instructions to:

update the conversation weight value of the first participant and the second participant when the first participant looks toward the second participant.

12. The non-transitory computer readable medium of claim 11 , wherein the instructions to update the conversation weight value of the first participant and the second participant when the first participant looks toward the second participant comprise instructions to evaluate at least one of a head pose or an eye gaze of the first participant.

13. The non-transitory computer readable medium of claim 12 , wherein the instructions to evaluate at least one of the head pose or the eye gaze of the first participant further comprise instructions to analyze a plurality of video sub-frames depicting the first participant, using an artificial neural network.

14. The non-transitory computer readable medium of claim 13 , wherein the instructions to estimate at least one of the head pose or the eye gaze of the first participant further comprise instructions to analyze the plurality of video sub-frames depicting the first participant, using a convolutional neural network.

15. A teleconferencing endpoint, comprising:

a network interface;

a camera system;

a microphone system;

a processor, the processor coupled to the network interface, the camera system, and the microphone system; and

a memory, the memory storing instructions executable by the processor, wherein the instructions comprise instructions to:

capture a first audio-video stream using the camera system and the microphone system;

derive a first sound source location from a first plurality of audio frames of the first audio-video stream;

update a first talker weight value of a first participant when the first sound source location corresponds to the first participant;

derive a second sound source location from a second plurality of audio frames of the first audio-video stream;

update a second talker weight value of a second participant when the second sound source location corresponds to the second participant;

update a conversation weight value of the first participant and the second participant when the first talker weight value and the second talker weight value exceed a first threshold and a difference between the first talker weight value and the second talker weight value is less than a predetermined amount;

select a first video sub-frame depicting the first participant for transmission in a second audio-video stream to a second teleconferencing endpoint through the network interface, when the first talker weight value exceeds a second threshold; and

select a second video sub-frame depicting the first participant and the second participant for transmission in the second audio-video stream to the second teleconferencing endpoint when the conversation weight value exceeds a third threshold.

16. The teleconferencing endpoint of claim 15 , wherein the instructions further comprise instructions to:

select a third video sub-frame depicting the second participant for transmission in the second audio-video stream to the second teleconferencing endpoint when the second talker weight value exceeds the second threshold.

17. The teleconferencing endpoint of claim 15 , wherein the instructions further comprise instructions to:

update the conversation weight value of the first participant and the second participant when the first participant makes one or more hand gestures directed toward the second participant.

18. The teleconferencing endpoint of claim 15 , wherein the instructions further comprise instructions to:

update the conversation weight value of the first participant and the second participant when the first participant looks toward the second participant.

19. The teleconferencing endpoint of claim 18 , wherein the instructions to update the conversation weight value of the first participant and the second participant when the first participant looks toward the second participant comprise instructions to evaluate at least one of a head pose or an eye gaze of the first participant.

20. The teleconferencing endpoint of claim 19 , wherein the instructions to evaluate at least one of the head pose or the eye gaze of the first participant further comprise instructions to analyze, using an artificial neural network, a plurality of video sub-frames depicting the first participant.

Assignments (4)
NUNC PRO TUNC ASSIGNMENT Recorded Nov 13, 2023
From: PLANTRONICS, INC.
To: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
Reel/Frame 065549/0065 →
RELEASE OF PATENT SECURITY INTERESTS Recorded Aug 30, 2022
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: PLANTRONICS, INC.; POLYCOM, INC.
Reel/Frame 061356/0366 →
SUPPLEMENTAL SECURITY AGREEMENT Recorded Oct 6, 2021
From: PLANTRONICS, INC.; POLYCOM, INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION
Reel/Frame 057723/0041 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 21, 2020
From: CHILDRESS, ROMMEL GABRIEL, JR.; NIMRI, ALAIN ELON; SCHAEFER, STEPHEN PAUL; YOUNG, DAVID
To: PLANTRONICS, INC.
Reel/Frame 054714/0058 →
Continuity (2)
Continuation 16773421 · Jan 27, 2020
Related Publication 20210235040A1 · Jul 29, 2021