IP Library › Granted Patent US 11,263,489
Granted Patent B2
US 11,263,489 · App. 16/616,533 · Granted Mar 1, 2022

Techniques for dense video descriptions

Inventors: Yurong Chen (Beijing, CN); Jianguo Li (Beijing, CN); Zhou Su (Beijing, CN); Zhiqiang Shen (Beijing, CN)
Assignee: INTEL CORPORATION
G06K9/6259G06F40/169G06K9/00718G06K9/00744G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,263,489
App. No.
16/616,533
Granted
Mar 1, 2022
Kind
B2
Abstract

Techniques and apparatus for generating dense natural language descriptions for video content are described. In one embodiment, for example, an apparatus may include at least one memory and logic, at least a portion of the logic comprised in hardware coupled to the at least one memory, the logic to receive a source video comprising a plurality of frames, determine a plurality of regions for each of the plurality of frames, generate at least one region-sequence connecting the determined plurality of regions, apply a language model to the at least one region-sequence to generate description information comprising a description of at least a portion of content of the source video. Other embodiments are described and claimed.

Claims (81)

1. An apparatus, comprising:

at least one memory; and

logic, at least a portion of the logic comprised in hardware coupled to the at least one memory, the logic to:

receive a source video comprising a plurality of frames,

determine a plurality of regions for the plurality of frames,

generate at least one region-sequence connecting the determined plurality of regions based on at least one selection criterion, the at least one selection criterion comprises an informativeness selection criterion, a coherency selection criterion or a divergence selection criterion, and

apply a language model to the at least one region-sequence to generate description information comprising a description of at least a portion of content of the source video.

2. The apparatus of claim 1 , the logic to generate a captioned video comprises at least one of the plurality of frames annotated with the at least one region-sequence and the description information.

3. The apparatus of claim 1 , the description information comprises a natural language description of at least one of the plurality of regions.

4. The apparatus of claim 1 , the logic is further to determine the at least one region-sequence based on at least one selection criterion, the at least one selection criterion comprises the informativeness selection criterion configured to maximize information in the at least one region-sequence.

5. The apparatus of claim 1 , the logic is further to determine the at least one region-sequence based on at least one selection criterion, the at least one selection criterion comprises the coherency selection criterion configured to maximize a cosine similarity between the plurality of regions of the at least one-region sequence.

6. The apparatus of claim 1 the logic is further to determine the at least one region-sequence based on at least one selection criterion, the at least one selection criterion comprises the divergency selection criterion configured to maximally separate the plurality of regions of the at least one region-sequence in terms of divergence.

7. The apparatus of claim 1 , the language model comprises a sequence-to-sequence learning framework comprising a plurality of long-short term memory networks (LSTMs).

8. An apparatus, comprising:

at least one memory; and

logic, at least a portion of the logic comprised in hardware coupled to the at least one memory, the logic to:

receive a source video comprising a plurality of frames,

determine a plurality of regions for the plurality of frames,

generate at least one region-sequence connecting the determined plurality of regions, and

apply a language model to the at least one region-sequence to generate description information comprising a description of at least a portion of content of the source video,

the logic is further to process at least one frame via a computational model to generate a response map comprising at least one anchor point representing at least one region of the plurality of regions.

9. An apparatus, comprising:

at least one memory; and

logic, at least a portion of the logic comprised in hardware coupled to the at least one memory, the logic to:

receive a source video comprising a plurality of frames,

determine a plurality of regions for the plurality of frames,

generate at least one region-sequence connecting the determined plurality of regions, and

apply a language model to the at least one region-sequence to generate description information comprising a description of at least a portion of content of the source video,

the logic is further to process at least one frame via a computational model comprising a convolutional neural network (CNN), the CNN comprises a lexical-fully convolutional neural network (lexical-FCN), in order to determine at least one of the plurality of regions.

10. An apparatus, comprising:

at least one memory; and

logic, at least a portion of the logic comprised in hardware coupled to the at least one memory, the logic to:

receive a source video comprising a plurality of frames,

determine a plurality of regions for the plurality of frames,

generate at least one region-sequence connecting the determined plurality of regions, and

apply a language model to the at least one region-sequence to generate description information comprising a description of at least a portion of content of the source video,

the logic is further to process at least one frame via a computational model comprising a convolutional neural network (CNN) trained with a multi-instance multi-label learning (MIMLL) process in order to determine at least one of the plurality of regions.

11. An apparatus, comprising:

at least one memory; and

logic, at least a portion of the logic comprised in hardware coupled to the at least one memory, the logic to:

receive a source video comprising a plurality of frames,

determine a plurality of regions for the plurality of frames,

generate at least one region-sequence connecting the determined plurality of regions, and

apply a language model to the at least one region-sequence to generate description information comprising a description of at least a portion of content of the source video,

the logic is further to process at least one frame via a computational model comprising a convolutional neural network (CNN) trained with a multi-instance multi-label learning (MIMLL) process to generate a lexical-fully convolutional neural network (lexical-FCN) in order to determine at least one of the plurality of regions.

12. A method, comprising:

receiving a source video comprising a plurality of frames;

determining a plurality of regions for each of the plurality of frames;

generating at least one region-sequence connecting the determined plurality of regions based on at least one selection criterion, the at least one selection criterion comprises an informativeness selection criterion, a coherency selection criterion or a divergence selection criterion; and

applying a language model to the at least one region-sequence to generate description information comprising a description of at least a portion of content of the source video.

13. The method of claim 12 , further comprising generating a captioned video comprising at least one of the plurality of frames annotated with the at least one region-sequence and the description information.

14. The method of claim 12 , wherein determining the at least one region-sequence comprises determining the at least one region-sequence based on at least one selection criterion, the at least one selection criterion comprises the informativeness selection criterion configured to maximize information in the at least one region-sequence.

15. The method of claim 12 , wherein determining the at least one region-sequence comprises determining the at least one region-sequence based on at least one selection criterion, the at least one selection criterion comprises the coherency selection criterion configured to maximize a cosine similarity between the plurality of regions of the at least one-region sequence.

16. The method of claim 12 , wherein determining the at least one region-sequence comprises determining the at least one region-sequence based on at least one selection criterion, the at least one selection criterion comprises the divergency selection criterion configured to maximally separate the plurality of regions of the at least one region-sequence in terms of divergence.

17. A method, comprising:

receiving a source video comprising a plurality of frames;

determining a plurality of regions for each of the plurality of frames;

generating at least one region-sequence connecting the determined plurality of regions; and

applying a language model to the at least one region-sequence to generate description information comprising a description of at least a portion of content of the source video;

wherein determining at least one of the plurality of regions comprises processing at least one frame via a computational model comprising a convolutional neural network (CNN), the CNN comprises a lexical-fully convolutional neural network (lexical-FCN).

18. A method, comprising:

receiving a source video comprising a plurality of frames;

determining a plurality of regions for each of the plurality of frames;

generating at least one region-sequence connecting the determined plurality of regions; and

applying a language model to the at least one region-sequence to generate description information comprising a description of at least a portion of content of the source video;

wherein determining at least one of the plurality of regions comprises processing at least one frame via a computational model comprising a convolutional neural network (CNN) trained with a multi-instance multi-label learning (MIMLL) process to generate a lexical-fully convolutional neural network (lexical-FCN).

19. A non-transitory computer-readable storage medium that stores executable computer instructions for execution by processing circuitry of a computing device, the instructions to cause the computing device to:

receive a source video comprising a plurality of frames;

determine a plurality of regions for each of the plurality of frames;

generate at least one region-sequence connecting the plurality of regions based on at least one selection criterion, the at least one selection criterion comprises an informativeness selection criterion, a coherency selection criterion or a divergence selection criterion; and

apply a language model to the at least one region-sequence to generate description information comprising a description of at least a portion of content of the source video.

20. The non-transitory computer-readable storage medium of claim 19 , the executable computer instructions to cause the computing device to generate a captioned video comprising at least one of the plurality of frames annotated with the at least one region-sequence and the description information.

21. The non-transitory computer-readable storage medium of claim 19 , the executable computer instructions to cause the computing device to determine the at least one region-sequence based on at least one selection criterion, the at least one selection criterion comprises the informativeness selection criterion configured to maximize information in the at least one region-sequence.

22. The non-transitory computer-readable storage medium of claim 19 , the executable computer instructions to cause the computing device to determine the at least one region-sequence based on at least one selection criterion, the at least one selection criterion comprises the coherency selection criterion configured to maximize a cosine similarity between the plurality of regions of the at least one-region sequence.

23. The non-transitory computer-readable storage medium of claim 19 , the executable computer instructions to cause the computing device to determine the at least one region-sequence based on at least one selection criterion, the at least one selection criterion comprises the divergency selection criterion configured to maximally separate the plurality of regions of the at least one region-sequence in terms of divergence.

24. A non-transitory computer-readable storage medium that stores executable computer instructions for execution by processing circuitry of a computing device, the instructions to cause the computing device to:

receive a source video comprising a plurality of frames;

determine a plurality of regions for each of the plurality of frames;

generate at least one region-sequence connecting the plurality of regions; and

apply a language model to the at least one region-sequence to generate description information comprising a description of at least a portion of content of the source video;

wherein determining at least one of the plurality of regions comprises process at least one frame via a computational model comprising a convolutional neural network (CNN) trained with a multi-instance multi-label learning (MIMLL) process to generate a lexical-fully convolutional neural network (lexical-FCN).

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 25, 2019
From: LI, JIANGUO; SHEN, ZHIQIANG; SU, ZHOU; CHEN, YURONG
To: INTEL CORPORATION
Reel/Frame 051100/0605 →
Continuity (1)
Related Publication 20210142115A1 · May 13, 2021