IP Library Granted Patent US 9,641,585
Granted Patent B2
US 9,641,585 · App. 14/733,485 · Granted May 2, 2017

Automated video editing based on activity in video conference

Inventors: Ingrid Kvaal (Oslo, NO); Vigleik Norheim (Oslo, NO); Kristian Tangeland (Oslo, NO)
Assignee: Cisco Technology, Inc.
H04L65/605G11B27/031H04N7/147H04N7/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,641,585
App. No.
14/733,485
Granted
May 2, 2017
Kind
B2
Abstract

In one embodiment, a method includes receiving at a network device, video and activity data for a video conference, automatically processing the video at the network device based on the activity data, and transmitting edited video from the network device. Processing comprises identifying active locations in the video and editing the video to display each of the active locations before a start of activity at the location and switch between the active locations. An apparatus and logic are also disclosed herein.

Claims (31)

1. A method comprising:

receiving at a network device, video and activity data for a video conference, the network device positioned within a communication path between endpoints of the video conference;

automatically processing the video at the network device receiving the video and activity data based on the activity data, wherein processing comprises identifying active locations in the video and editing the video to display each of said active locations before a start of activity at said location and switch between said active locations before a start of activity; and

transmitting the edited video from the network device as a near real-time video with a small delay to account for automated editing of the video;

wherein the video comprises raw video from one or more endpoints and contains said active locations and wherein the edited video displays said active locations before a start of said activity at said active locations and includes a speaker's initial facial reaction before speaking.

2. The method of claim 1 wherein the activity data comprises event logs generated in the video conference.

3. The method of claim 1 wherein processing the video comprises identifying a first location in the video and editing the video to display said first location before a start of activity at said first location and switching the video to display a second location before a start of activity at said second location, wherein each of said locations comprise an active speaker.

4. The method of claim 1 wherein identifying said active locations comprises identifying locations of active speakers in the video.

5. The method of claim 4 wherein the activity data comprises audio data associated with the active speakers.

6. The method of claim 4 wherein the video is edited to simultaneously display at least two active speakers when a change in speaker occurs in less than a predetermined time.

7. The method of claim 4 wherein the active speaker is talking more than a predetermined time and further comprising editing the video to temporarily display an image other than the active speaker.

8. The method of claim 1 wherein receiving video comprises receiving raw video images from a plurality of image capturing devices at an endpoint.

9. The method of claim 1 wherein receiving video comprises receiving raw video images from a plurality of endpoints participating in the video conference.

10. The method of claim 1 wherein the activity data comprises speaker tracking data.

11. The method of claim 1 wherein transmitting the edited video comprises transmitting the edited video to an endpoint in the video conference with a delay in a live broadcast.

12. The method of claim 1 wherein transmitting the edited video comprises transmitting the edited video after the video conference has ended.

13. The method of claim 1 wherein the video comprises multiple input streams of the video conference covering all participants at all endpoints for the duration of the video conference.

14. The method of claim 13 further comprising editing the video to display all participants for a specified period of time at a start and end of the video conference.

15. An apparatus comprising:

an interface for receiving video and activity data for a video conference;

a processor for automatically processing the video based on the activity data, wherein processing comprises using the activity data to identify active locations in the video and editing the video to display each of said active locations before a start of activity at said location and switch between said active locations before a start of activity, and transmitting the edited video as a near real-time video with a small delay to account for said automated editing of the video; and

memory for storing the video and the activity data;

wherein the video comprises raw video from one or more endpoints and contains said active locations and wherein the edited video displays said active locations before a start of said activity at said active locations and includes a speaker's initial facial reaction before speaking.

16. The apparatus of claim 15 wherein the activity data comprises event logs generated in the video conference.

17. The apparatus of claim 15 wherein identifying said active locations comprises identifying locations of active speakers in the video.

18. The apparatus of claim 15 wherein the activity data comprises speaker tracking data.

19. The apparatus of claim 15 wherein transmitting the edited video comprises transmitting the edited video to an endpoint in the video conference with a delay in a live broadcast.

20. Logic encoded on one or more non-transitory computer readable media for execution and when executed operable to:

automatically process video based on activity data, wherein processing comprises identifying active locations in the video and editing the video to display each of said active locations before a start of activity at said location and switch between said active locations before a start of activity; and

transmit the edited video from the network device as a near real-time video with a small delay to account for automated editing of the video;

wherein the video and activity data are collected for a video conference and the video comprises raw video from one or more endpoints and contains said active locations and wherein the edited video displays said active locations before a start of said activity at said active locations and includes a speaker's initial facial reaction before speaking.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 8, 2015
From: KVAAL, INGRID; NORHEIM, VIGLEIK; TANGELAND, KRISTIAN
To: CISCO TECHNOLOGY, INC.
Reel/Frame 035804/0356 →
Continuity (1)
Related Publication 20160359941A1 · Dec 8, 2016