IP Library › Granted Patent US 10,949,674
Granted Patent B2
US 10,949,674 · App. 16/298,549 · Granted Mar 16, 2021

Video summarization using semantic information

Inventors: Myung Hwangbo (Lake Oswego, OR); Krishna Kumar Singh (Davis, CA); Teahyung Lee (Chandler, AZ); Omesh Tickoo (Portland, OR)
Assignee: Intel Corporation
G06K9/00751G06K9/00718G06K9/00765G06K9/46G06K9/628G06N3/0454G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,949,674
App. No.
16/298,549
Granted
Mar 16, 2021
Kind
B2
Abstract

An apparatus for video summarization using sematic information is described herein. The apparatus includes a controller, a scoring mechanism, and a summarizer. The controller is to segment an incoming video stream into a plurality of activity segments, wherein each frame is associated with an activity. The scoring mechanism is to calculate a score for each frame of each activity, wherein the score is based on a plurality of objects in each frame. The summarizer is to summarize the activity segments based on the score for each frame.

Claims (48)

1. An electronic device comprising:

an image capture sensor;

memory to store video segments;

wireless communication circuitry to transmit data; and

processor circuitry to:

process a first image of a first video segment from the image capture sensor with a neural network to determine a first score for the first image, the neural network to detect actions associated with images, the actions associated with labels;

determine a second score for the first video segment based on first scores for corresponding images in the first video segment; and

determine, based on the second score, whether to retain the first video segment in the memory.

2. The electronic device of claim 1 , wherein the neural network is to output a confidence that the first image is associated with a first one of the labels.

3. The electronic device of claim 1 , wherein the first video segment has a duration of at least five seconds.

4. The electronic device of claim 1 , wherein the labels correspond to a group of labels having a size of at least hundreds of labels.

5. The electronic device of claim 1 , wherein the electronic device is a wearable device.

6. The electronic device of claim 1 , wherein the wireless communication circuitry includes at least one of WiFi hardware, Bluetooth hardware or cellular hardware.

7. The electronic device of claim 1 , wherein the neural network is a convolutional neural network.

8. The electronic device of claim 1 , wherein the neural network is trained to detect the actions associated with the images.

9. The electronic device of claim 1 , further including:

a display;

a microphone; and

at least one of a keyboard, a touchpad or a touchscreen.

10. At least one memory device comprising computer readable instructions that, when executed, cause at least one processor of an electronic device to at least:

process a first image of a first video segment from an image capture sensor of the electronic device with a neural network to determine a first score for the first image, the neural network to detect an action associated with at least one image, the action associated with a label;

determine a second score for the first video segment based on first scores for corresponding images in the first video segment; and

determine, based on the second score, whether to retain the first video segment in the electronic device.

11. The at least one memory device of claim 10 , wherein the neural network is to output a confidence that the first image is associated with the label.

12. The at least one memory device of claim 10 , wherein the first video segment has a duration of at least five seconds.

13. The at least one memory device of claim 10 , wherein the label is one of a group of labels having a size of at least hundreds of labels.

14. The at least one memory device of claim 10 , wherein the neural network is a convolutional neural network trained to detect the action.

15. An electronic device comprising:

means for capturing images;

means for storing video segments;

means for transmitting data; and

means for determining whether to retain a first video segment, the means for determining to:

process a first image of the first video segment with a neural network to determine a first score for the first image, the neural network to detect actions associated with images, the actions associated with labels;

determine a second score for the first video segment based on first scores for corresponding images in the first video segment; and

determine, based on the second score, whether to retain the first video segment.

16. The electronic device of claim 15 , wherein the neural network is to output a confidence that the first image is associated with a first one of the labels.

17. The electronic device of claim 15 , wherein the first video segment has a duration of at least five seconds.

18. The electronic device of claim 15 , wherein the labels correspond to a group of labels having a size of at least hundreds of labels.

19. The electronic device of claim 15 , wherein the means for transmitting is to transmit via a WiFi network.

20. The electronic device of claim 15 , wherein the neural network is a convolutional neural network to detect the actions associated with the images.

21. A method comprising:

processing, by executing an instruction with at least one processor of an electronic device, a first image of a first video segment with a neural network to determine a first score for the first image, the neural network to detect actions associated with images, the actions associated with labels, the first video segment from an image capture sensor of the electronic device;

determining, by executing an instruction with the at least one processor, a second score for the first video segment based on first scores for corresponding images in the first video segment; and

determining, by executing an instruction with the at least one processor, whether to retain the first video segment in the electronic device based on the second score.

22. The method of claim 21 , wherein the neural network is to output a confidence that the first image is associated with a first one of the labels.

23. The method of claim 21 , wherein the first video segment has a duration of at least five seconds.

24. The method of claim 21 , wherein the labels correspond to a group of labels having a size of at least hundreds of labels.

25. The method of claim 21 , wherein the neural network is a convolutional neural network trained to detect the actions associated with the images.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE ATTORNEY DOCKET NUMBER TO P92596-C1 PREVIOUSLY RECORDED ON REEL 048789 FRAME 0642. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNORS INTEREST. Recorded Apr 9, 2019
From: HWANGBO, MYUNG; SINGH, KRISHNA KUMAR; LEE, TEAHYUNG; TICKOO, OMESH
To: INTEL CORPORATION
Reel/Frame 048838/0568 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 4, 2019
From: HWANGBO, MYUNG; SINGH, KRISHNA KUMAR; LEE, TEAHYUNG; TICKOO, OMESH
To: INTEL CORPORATION
Reel/Frame 048789/0642 →
Continuity (2)
Continuation 14998322 · Dec 24, 2015
Related Publication 20200012864A1 · Jan 9, 2020
Cited By (1)
US 12,277,768