IP Library › Granted Patent US 10,516,852
Granted Patent B2
US 10,516,852 · App. 15/981,299 · Granted Dec 24, 2019

Multiple simultaneous framing alternatives using speaker tracking

Inventors: Christian Fjelleng Theien (Asker, NO); Rune Øistein Aas (Lysaker, NO); Kristian Tangeland (Oslo, NO)
Assignee: Cisco Technology, Inc.
H04N7/152H04N7/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,516,852
App. No.
15/981,299
Granted
Dec 24, 2019
Kind
B2
Abstract

In one embodiment, a video conference endpoint may detect a one or more participants within a field of view of a camera of the video conference endpoint. The video conference endpoint may determine one or more alternative framings of an output of the camera of the video conference endpoint based on the detected one or more participants. The video conference endpoint may send the output of the camera of the video conference endpoint to one or more far-end video conference endpoints participating in a video conference with the video conference endpoint. The video conference endpoint may send data descriptive of the one or more alternative framings of the output of the camera to the far-end video conference endpoints. The far-end video conference endpoints may utilize the data to display one of the one or more alternative framings.

Claims (52)

1. A method comprising:

detecting, by a primary video conference endpoint, one or more participants within a field of view of a camera of the primary video conference endpoint;

determining, by the primary video conference endpoint, one or more alternative framings of an output of the camera of the primary video conference endpoint based on the detected one or more participants;

sending, by the primary video conference endpoint, the output of the camera of the primary video conference endpoint to one or more secondary video conference endpoints; and

sending, by the primary video conference endpoint, data descriptive of the one or more alternative framings of the output of the camera of the primary video conference endpoint to the one or more secondary video conference endpoints, wherein the one or more secondary video conference endpoints can utilize the data to alter the output of the camera of the primary video conference endpoint to display one of the one or more alternative framings.

2. The method of claim 1 , wherein the output of the camera of the primary video conference endpoint is a high resolution video stream.

3. The method of claim 1 , wherein the sending the output includes sending the output of the camera of the primary video conference endpoint to one or more of the secondary video conference endpoints via a first channel.

4. The method of claim 3 , wherein the sending the data includes sending the data of the one or more alternative framings to one or more of the secondary video conference endpoints as metadata via a secondary metadata channel.

5. The method of claim 3 , wherein the sending the data includes sending the data of the one or more alternative framings to one or more of the secondary video conference endpoints as metadata with the output of the camera of the primary video conference endpoint via the first channel.

6. The method of claim 1 , further comprising:

detecting, by the primary video conference endpoint, a first participant of the one or more participants as an active speaker.

7. The method of claim 6 , wherein determining one or more alternative framings of an output of the camera of the primary video conference endpoint based on the detected one or more participants further comprises:

determining a wide framing that includes each of the one or more participants;

determining a tight framing that includes each of the one or more participants;

determining a first speaker framing that includes the active speaker and surroundings of the active speaker;

determining a second speaker framing that is a close-up of the active speaker; and

determining a thumbnail framing of each of the one or more participants.

8. An apparatus comprising:

a network interface unit that enables communication over a network by a primary video conference endpoint; and

a processor coupled to the network interface unit, the processor configured to:

detect one or more participants within a field of view of a camera of the primary video conference endpoint;

determine one or more alternative framings of an output of the camera of the primary video conference endpoint based on the detected one or more participants;

send the output of the camera of the primary video conference endpoint to one or more secondary video conference endpoints; and

send data descriptive of the one or more alternative framings of the output of the camera of the primary video conference endpoint to the one or more secondary video conference endpoints, wherein the one or more secondary video conference endpoints can utilize the data to alter the output of the camera of the primary video conference endpoint to display the one or more alternative framings.

9. The apparatus of claim 8 , wherein the output of the camera of the primary video conference endpoint is a high resolution video stream.

10. The apparatus of claim 8 , wherein the output of the camera of the primary video conference endpoint is sent to one or more of the secondary video conference endpoints via a first channel.

11. The apparatus of claim 10 , wherein the data of the one or more alternative framings is sent to one or more of the secondary video conference endpoints as metadata via a secondary metadata channel.

12. The apparatus of claim 10 , wherein the data of the one or more alternative framings is sent to one or more of the secondary video conference endpoints as metadata with the output of the camera of the primary video conference endpoint via the first channel.

13. The apparatus of claim 8 , wherein the processor is further configured to:

detect a first participant of the one or more participants as an active speaker.

14. The apparatus of claim 13 , wherein the processor, when determining one or more alternative framings of an output of the camera of the primary video conference endpoint based on the detected one or more participants, is further configured to:

determine a wide framing that includes each of the one or more participants;

determine a tight framing that includes each of the one or more participants;

determine a first speaker framing that includes the active speaker and surroundings of the active speaker;

determine a second speaker framing that is a close-up of the active speaker; and

determine a thumbnail framing of each of the one or more participants.

15. A non-transitory processor readable medium storing instructions that, when executed by a processor, cause the processor to:

detect one or more participants within a field of view of a camera of a primary video conference endpoint;

determine one or more alternative framings of an output of the camera of the primary video conference endpoint based on the detected one or more participants;

send the output of the camera of the primary video conference endpoint to one or more secondary video conference endpoints; and

send data descriptive of the one or more alternative framings of the output of the camera of the primary video conference endpoint to the one or more secondary video conference endpoints, wherein the one or more secondary video conference endpoints can utilize the data to alter the output of the camera of the primary video conference endpoint to display one of the one or more alternative framings.

16. The non-transitory processor readable medium of claim 15 , wherein the output of the camera of the primary video conference endpoint is a high resolution video stream.

17. The non-transitory processor readable medium of claim 15 , wherein the output of the camera of the primary video conference endpoint is sent to one or more of the secondary video conference endpoints via a first channel.

18. The non-transitory processor readable medium of claim 17 , wherein the data of the one or more alternative framings is sent to one or more of the secondary video conference endpoints as metadata via a secondary metadata channel.

19. The non-transitory processor readable medium of claim 15 , further comprising instructions that, when executed by the processor, cause the processor to:

detect a first participant of the one or more participants as an active speaker.

20. The non-transitory processor readable medium of claim 19 , further comprising instructions that, when executed by the processor, cause the processor, when determining one or more alternative framings of an output of the camera of the primary video conference endpoint based on the detected one or more participants, to:

determine a wide framing that includes each of the one or more participants;

determine a tight framing that includes each of the one or more participants;

determine a first speaker framing that includes the active speaker and surroundings of the active speaker;

determine a second speaker framing that is a close-up of the active speaker; and

determine a thumbnail framing of each of the one or more participants.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 16, 2018
From: THEIEN, CHRISTIAN FJELLENG; AAS, RUNE ØISTEIN; TANGELAND, KRISTIAN
To: CISCO TECHNOLOGY, INC.
Reel/Frame 045822/0050 →
Continuity (1)
Related Publication 20190356883A1 · Nov 21, 2019
Cited By (5)
US 12,333,854 US 12,342,100 US 12,593,008 US 12,615,347 US 12,719,709