IP Library Granted Patent US 10,460,196
Granted Patent B2
US 10,460,196 · App. 15/232,533 · Granted Oct 29, 2019

Salient video frame establishment

Inventors: Anmol Dhawan (Ghaziabad, IN); Varun Maini (New Delhi, IN); Srinivasa Madhava Phaneen Angara (Noida, IN); Amol Jindal (Patiala, IN)
Assignee: Adobe Inc.
G06K9/4671G06K9/00751G06K9/623
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,460,196
App. No.
15/232,533
Granted
Oct 29, 2019
Kind
B2
Abstract

Salient video frame establishment is described. In one or more example embodiments, salient frames of a video are established based on multiple photos. An image processing module is capable of analyzing both video frames and photos, both of which may include entities, such as faces or objects. Frames of a video are decoded and analyzed in terms of attributes of the video. Attributes include, for example, scene boundaries, facial expressions, brightness levels, and focus levels. From the video frames, the image processing module determines candidate frames based on the attributes. The image processing module analyzes multiple photos to ascertain multiple relevant entities based on the presence of entities in the multiple photos. Relevancy of an entity can depend, for instance, on a number of occurrences. The image processing module establishes multiple salient frames from the candidate frames based on the multiple relevant entities. Salient frames can be displayed.

Claims (59)

1. In a digital medium environment to extract multiple salient frames from a video based at least partially on entities present in one or more photos, a method implemented by at least one computing device, the method comprising:

obtaining, by the at least one computing device, a video including multiple frames;

obtaining, by the at least one computing device, multiple photos which are extrinsic to the video;

ascertaining, by the at least one computing device, multiple relevant entities in the video based on the multiple photos extrinsic to the video;

determining, by the at least one computing device, multiple candidate frames from the multiple frames of the video;

establishing, by the at least one computing device, multiple salient frames, in part, by:

filtering the multiple candidate frames based on the multiple relevant entities based on the multiple photos which are extrinsic to the video, and

computing multiple salient scores for the multiple candidate frames, each respective salient score corresponding to a respective candidate frame, each respective salient score is based on an image quality indicator of the each of the respective candidate frame and a relevancy score computed for at least one entity appearing in the respective candidate frame; and

controlling, by the at least one computing device, presentation of the multiple salient frames via a user interface.

2. The method as described in claim 1 , wherein the obtaining of the multiple photos comprises retrieving the multiple photos from an image library based on a time associated with the video and a temporal threshold.

3. The method as described in claim 1 , wherein the image quality indicator comprises at least one of a frame focus level or a frame brightness level.

4. The method as described in claim 1 , wherein the ascertaining comprises:

detecting a relevant entity in at least one photo of the multiple photos;

recognizing the detected relevant entity; and

assigning an entity identifier to the recognized relevant entity.

5. The method as described in claim 4 , wherein:

the ascertaining further comprises:

determining an occurrence value for the recognized relevant entity across the multiple photos; and

associating the occurrence value with the entity identifier; and

the establishing further comprises establishing the multiple salient frames based on the occurrence value associated with the entity identifier.

6. The method as described in claim 1 , wherein the multiple relevant entities comprise at least one of relevant faces or relevant objects.

7. The method as described in claim 1 , wherein:

the ascertaining comprises computing a relevancy score for each relevant entity of the multiple relevant entities based on the multiple photos; and

the establishing further comprises:

ranking the multiple candidate frames based on the multiple salient scores, and

selecting the multiple salient frames based on the ranking of the multiple candidate frames.

8. The method as described in claim 1 , wherein:

the establishing further comprises ranking the multiple salient frames; and

the controlling further comprises causing the multiple salient frames to be displayed based on the ranking of the multiple salient frames.

9. At least one computing device operative in a digital medium environment to extract frames from a video based at least partially on entities present in one or more photos, the at least one computing device comprising:

a processing system and at least one computer-readable storage medium including:

a relevant entity ascertainment module configured to ascertain multiple relevant entities based on multiple photos, wherein the multiple photos are extrinsic to the video;

a candidate frame determination module configured to determine multiple candidate frames from multiple frames of the video;

a salient frame establishment module configured to establish multiple salient frames, at least in part, by:

filtering the multiple candidate frames based on the multiple relevant entities, and

computing multiple salient scores for the multiple candidate frames, each respective salient score corresponding to a respective candidate frame, each respective salient score is based on an image quality indicator of the respective candidate frame and a relevancy score computed for at least one entity appearing in the respective candidate frame; and

a salient frame output module configured to control presentation of the multiple salient frames via a user interface.

10. The at least one computing device described in claim 9 , wherein the relevant entity ascertainment module is configured to compute the relevancy score for each respective corresponding relevant entity of the multiple relevant entities.

11. The at least one computing device as described in claim 10 , wherein the relevant entity ascertainment module is configured to compute each relevancy score based on an occurrence value that depends on a number of occurrences of the respective corresponding relevant entity across the multiple photos.

12. The at least one computing device as described in claim 9 , wherein the salient frame output module includes a collage creation module configured to create a static collage using at least a portion of the multiple salient frames.

13. The at least one computing device as described in claim 12 , wherein the collage creation module is configured to create the static collage using at least the portion of the multiple salient frames responsive to user input directed to a presentation of the multiple salient frames.

14. At least one computing device operative in a digital medium environment to extract frames from a video based at least partially on entities present in one or more photos, the at least one computing device including hardware components comprising a processing system, one or more computer-readable storage media storing computer-readable instructions that are executable by the processing system to perform operations comprising:

computing a relevancy score for each respective one of multiple relevant entities based on a presence of a corresponding relevant entity in at least one photo of multiple photos, wherein the multiple photos are extrinsic to the video;

computing a salient score for each respective one of multiple candidate frames from multiple frames of a video, each respective salient score corresponding to a respective candidate frame, each respective salient score is based on an image quality indicator of the respective candidate frame and incorporating the respective relevancy score responsive to an appearance of the corresponding relevant entity in the respective candidate frame;

establishing multiple salient frames of the multiple candidate frames using a ranking of the multiple candidate frames that is based on the respective salient scores of the multiple candidate frames; and

causing at least a portion of the multiple salient frames to be presented via a user interface.

15. The at least one computing device as described in claim 14 , wherein:

the one or more attributes of the video comprise at least one of a per-frame focus level indicator, a per-frame brightness level indicator, or a length of time a recognized entity appears in the video; and

the presence of the corresponding relevant entity comprises at least one of a number of occurrences across the multiple photos, a proportional spatial coverage over at least one photo of the multiple photos, or a positional presence in at least one photo of the multiple photos.

16. The method as described in claim 1 , wherein at least one of the multiple scenes includes consecutive frames, each of which displays an object such that the displayed object appears at an angle in one of the consecutive frames and another angle in at least another one of the consecutive frames.

17. The at least one computing device as described in claim 9 , wherein at least one of the multiple scenes includes consecutive frames, each of which displays an object such that the displayed object appears at an angle in one of the consecutive frames and another angle in at least another one of the consecutive frames.

18. The at least one computing device as described in claim 9 , wherein the relevant entity ascertaining module ascertains multiple relevant entities by:

detecting a relevant entity in at least one photo of the multiple photos;

recognizing the detected relevant entity; and

assigning an entity identifier to the recognized relevant entity.

19. The method as described in claim 18 , further comprising:

determining an occurrence value for the recognized relevant entity across the multiple photos; and

associating the occurrence value with the entity identifier.

20. The at least one computing device as described in claim 14 , wherein the multiple relevant entities comprise at least one of relevant faces or relevant objects.

Assignments (2)
CHANGE OF NAME Recorded Jan 18, 2019
From: ADOBE SYSTEMS INCORPORATED
To: ADOBE INC.
Reel/Frame 048097/0414 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2016
From: DHAWAN, ANMOL; MAINI, VARUN; ANGARA, SRINIVASA MADHAVA PHANEEN; JINDAL, AMOL
To: ADOBE SYSTEMS INCORPORATED
Reel/Frame 039387/0567 →
Continuity (1)
Related Publication 20180046879A1 · Feb 15, 2018
Cited By (1)
US 12,681,897