IP Library › Granted Patent US 11,068,722
Granted Patent B2
US 11,068,722 · App. 16/342,084 · Granted Jul 20, 2021

Method for analysing media content to generate reconstructed media content

Inventors: Francesco Cricri (Tampere, FI); Mikko Honkala (Espoo, FI); Emre Baris Aksu (Tampere, FI); Xingyang Ni (Tampere, FI)
Assignee: Nokia Technologies Oy
G06K9/00744G06K9/00664G06K9/00711G06K9/4628G06K9/6234G06N3/04G06N3/0445G06N3/0454
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,068,722
App. No.
16/342,084
Granted
Jul 20, 2021
Kind
B2
Abstract

The invention relates to a method, an apparatus and a computer program product for analyzing media content. The method comprises receiving media content; performing feature extraction of the media content at a plurality of convolution layers to produce a plurality of layer-specific feature maps; transmitting from the plurality of convolution layers a corresponding layer-specific feature map to a corresponding de-convolution layer of a plurality of de-convolution layers via a recurrent connection between the plurality of convolution layers and the plurality of de-convolution layers; and generating a reconstructed media content based on the plurality of feature maps.

Claims (34)

1. A method, comprising:

receiving a media content;

performing feature extraction of the media content at a plurality of convolution layers to produce a plurality of layer-specific feature maps;

transmitting from the plurality of convolution layers a corresponding layer-specific feature map to a corresponding de-convolution layer of a plurality of de-convolution layers via a direct recurrent connection between a convolution layer of the plurality of convolution layers and the corresponding deconvolution layer of the plurality of de-convolution layers; and

generating, with the plurality of de-convolution layers, a reconstructed media content based on the plurality of layer-specific feature maps.

2. The method according to claim 1 , wherein the media content comprises video frames, and wherein the reconstructed media content comprises predicted future video frames.

3. The method according to claim 1 , wherein the recurrent connection comprises a long short-term memory network.

4. The method according to claim 1 , further comprising providing the reconstructed media content and corresponding original media content to a discriminator system, wherein the discriminator system is configured to determine whether the reconstructed media content is real.

5. The method according to claim 4 , wherein the discriminator system comprises a plurality of discriminators corresponding to the plurality of de-convolution layers.

6. The method according to claim 5 , further comprising providing, to respective discriminators of the plurality of discriminators, the reconstructed media content from a corresponding de-convolution layer of a plurality of additional de-convolution layers.

7. An apparatus comprising processor circuitry, at least one non-transitory memory including computer program code, wherein the at least one memory and the computer program code, with the processor circuitry, configured to cause the apparatus to perform at least the following:

receive media content;

perform feature extraction of the media content at a plurality of convolution layers to produce a plurality of layer-specific feature maps;

transmit from the plurality of convolution layers a corresponding layer-specific feature map to a corresponding de-convolution layer of a plurality of de-convolution layers via a direct recurrent connection between a convolution layer of the plurality of convolution layers and the corresponding de-convolution layer of the plurality of de-convolution layers; and

generate, with the plurality of de-convolution layers, reconstructed media content based on the plurality of layer-specific feature maps.

8. The apparatus according to claim 7 , wherein the media content comprises video frames, and wherein the reconstructed media content comprises predicted future video frames.

9. The apparatus according to claim 7 , wherein the recurrent connection comprises a long short-term memory network.

10. The apparatus according to claim 7 , wherein the apparatus is further caused to provide the reconstructed media content and corresponding original media content to a discriminator system, wherein the discriminator system is configured to determine whether the reconstructed media content is real.

11. The apparatus according to claim 10 , wherein the discriminator system comprises a plurality of discriminators corresponding to the plurality of de-convolution layers.

12. The apparatus according claim 11 , wherein the apparatus is further caused to provide, to respective discriminators of the plurality of discriminators, the reconstructed media content from a corresponding de-convolution layer of a plurality of additional de-convolution layers.

13. A computer program product embodied on a non-transitory computer readable medium, comprising computer program code, which when executed with at least one processor, cause an apparatus or a system to:

receive media content;

perform feature extraction of the media content at a plurality of convolution layers to produce a plurality of layer-specific feature maps;

transmit from the plurality of convolution layers a corresponding layer-specific feature map to a corresponding de-convolution layer of a plurality of de-convolution layers via a direct recurrent connection between a convolution layer of the plurality of convolution layers and the corresponding de-convolution layer of the plurality of de-convolution layers; and

generate, with the plurality of de-convolution layers, reconstructed media content based on the plurality of layer-specific feature maps.

14. The computer program product according to claim 13 , wherein the media content comprises video frames, and wherein the reconstructed media content comprises predicted future video frames.

15. The computer program product according to claim 13 , wherein the recurrent connection comprises a long short-term memory network.

16. The computer program product according to any of the claim 13 , wherein the apparatus or the system is further caused to provide the reconstructed media content and corresponding original media content to a discriminator system, wherein the discriminator system is configured to determine whether the reconstructed media content is real.

17. The computer program product according to claim 16 , wherein the discriminator system comprises a plurality of discriminators corresponding to the plurality of de-convolution layers.

18. The computer program product according claim 17 , wherein the apparatus or the system is further caused to provide, to respective discriminators of the plurality of discriminators, the reconstructed media content from a corresponding de-convolution layer of a plurality of additional de-convolution layers.

19. The method according to claim 1 , wherein the transmitting from the plurality of convolution layers the corresponding layer-specific feature map to the corresponding de-convolution layer comprises:

transmitting, from a first convolution layer of the plurality of convolution layers, a first layer-specific feature map to a first de-convolution layer of the plurality of de-convolution layers via a first recurrent connection; and

transmitting, from a different second convolution layer of the plurality of convolution layers, a second layer-specific feature map to a second de-convolution layer of the plurality of de-convolution layers via a second recurrent connection.

20. The method according to claim 1 , further comprising transmitting, from the plurality of convolution layers, direct information to the corresponding de-convolution layer of the plurality of de-convolution layers via the direct recurrent connection between the convolution layer of the plurality of convolution layers and the corresponding de-convolution layer of the plurality of de-convolution layers.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2019
From: CRICRI, FRANCESCO; HONKALA, MIKKO; AKSU, EMRE BARIS; NI, XINGYANG
To: NOKIA TECHNOLOGIES OY
Reel/Frame 048886/0319 →
Priority Claims (1)
GB 1618160 · Oct 27, 2016 · national
Continuity (1)
Related Publication 20190251360A1 · Aug 15, 2019
Cited By (1)
US 12,651,143