IP Library Granted Patent US 11,290,746
Granted Patent B2
US 11,290,746 · App. 17/052,342 · Granted Mar 29, 2022

Method and device for multi-view video decoding and method and device for image processing

Inventors: Joel Jung (Chatillon, FR); Pavel Nikitin (Chatillon, FR); Patrick Boissonade (Chatillon, FR)
Assignee: ORANGE
H04N19/597H04N19/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,290,746
App. No.
17/052,342
Granted
Mar 29, 2022
Kind
B2
Abstract

A method and a device for decoding a data stream representative of a multi-view video. Syntax elements are obtained from at least one part of the stream data, and used to reconstruct at least one image of a view of the video. Then, at least one item of metadata in a predetermined form is obtained from at least one obtained syntax element, and provided to an image processing module. Also provided are a method and a device for processing images configured to read said at least one item of metadata in the predetermined form and use it to generate at least one image of a virtual view from a reconstructed view of the multi-view video.

Claims (84)

1. A decoding method for decoding a data stream representative of a multi-view video, implemented by a decoding device, the method comprising:

obtaining syntax elements from at least one part of the stream data;

reconstructing at least one image of a view of the video from the syntax elements obtained;

obtaining at least one item of metadata in a predetermined form from at least one syntax element, said at least one item of metadata being associated with said at least one reconstructed image, and said at least one item of metadata corresponding to an item of information included in the group consisting of:

camera parameters,

decoded and scaled motion vectors,

a partitioning of the reconstructed image,

a reference image used by a block of an image of the reconstructed view,

the coding modes of an image of the reconstructed view,

quantization parameter values of an image of the reconstructed view,

prediction residual values of an image of the reconstructed view,

a map representative of movement in an image of the reconstructed view,

a map representative of presence of occlusions in an image of the reconstructed view,

a map representative of confidence values associated with a depth map; and

providing said at least one item of metadata to an image synthesis module, said image synthesis module being configured to synthesize at least one virtual view, distinct from the views of the multi-view video, from said at least one reconstructed image and said at least one item of metadata.

2. The decoding method according to claim 1 , wherein obtaining at least one item of metadata further comprises calculating said at least one item of metadata from at least one part of the syntax elements.

3. The decoding method according to claim 1 , wherein said at least one item of metadata is not used for reconstructing the at least one image.

4. The decoding method according to claim 1 , wherein the predetermined form corresponds to an indexed table in which at least one item of metadata is stored in association with an index.

5. The decoding method according to claim 1 , wherein said at least one item of metadata is obtained based on a granularity level specified in the decoding device.

6. The decoding method according to claim 1 , further comprising receiving by the decoding device a request from the image synthesis module indicating at least one item of metadata required by the image synthesis module.

7. The decoding method according to claim 6 , wherein the request comprises at least one index indicating the required item of metadata among a predetermined list of available metadata.

8. A device for decoding a data stream representative of a multi-view video, wherein the device comprises:

a processor; and

a non-transitory computer-readable medium comprising instructions stored thereon which when executed by the processor configure the decoding device to:

obtain syntax elements from at least one part of the stream data;

reconstruct at least one image of a view of the video from the syntax elements obtained;

obtain at least one item of metadata in a predetermined form from at least one syntax element, said at least one item of metadata being associated with said at least one reconstructed image, and said at least one item of metadata corresponding to an item of information included in the group consisting of:

camera parameters,

decoded and scaled motion vectors,

a partitioning of the reconstructed image,

a reference image used by a block of an image of the reconstructed view,

the coding modes of an image of the reconstructed view,

quantization parameter values of an image of the reconstructed view,

prediction residual values of an image of the reconstructed view,

a map representative of movement in an image of the reconstructed view,

a map representative of presence of occlusions in an image of the reconstructed view,

a map representative of confidence values associated with a depth map; and

provide said at least one item of metadata to an image synthesis module, said image synthesis module being configured to synthesize at least one virtual view, distinct from the views of the multi-view video, from said at least one reconstructed image and said at least one item of metadata.

9. An image synthesis method implemented by a synthesis device, the method comprising:

generating at least one image of a virtual view, from at least one image of a view of a multi-view video decoded by a decoding device, said virtual view being distinct from views of the multi-view video, by:

reading at least one item of metadata in a predetermined form, said at least one item of metadata being obtained by the decoding device from at least one syntax element obtained from a data stream representative of a multi-view video, and said at least one item of metadata corresponding to an item of information included in the group consisting of:

camera parameters,

decoded and scaled motion vectors,

a partitioning of the reconstructed image,

a reference image used by a block of an image of the reconstructed view,

the coding modes of an image of the reconstructed view,

quantization parameter values of an image of the reconstructed view,

prediction residual values of an image of the reconstructed view,

a map representative of movement in an image of the reconstructed view,

a map representative of presence of occlusions in an image of the reconstructed view,

a map representative of confidence values associated with a depth map; and

generating said at least one image comprising using said at least one read item of metadata and said decoded image.

10. The image synthesis method according to claim 9 , further comprising sending to the decoding device a request indicating at least one item of metadata required to generate the image.

11. An image synthesis device configured to generate at least one image of a virtual view, from at least one image of a view of a multi-view video decoded by a decoding device, said virtual view being distinct from views of the multi-view video, the image synthesis device comprising:

a processor; and

a non-transitory computer-readable medium comprising instructions stored thereon which when executed by the processor configure the image synthesis device to:

read at least one item of metadata in a predetermined form, said at least one item of metadata being obtained by the decoding device from at least one syntax element obtained from a data stream representative of a multi-view video, and said at least one item of metadata corresponding to an item of information included in the group consisting of:

camera parameters,

decoded and scaled motion vectors,

a partitioning of the reconstructed image,

a reference image used by a block of an image of the reconstructed view,

the coding modes of an image of the reconstructed view,

quantization parameter values of an image of the reconstructed view,

prediction residual values of an image of the reconstructed view,

a map representative of the movement in an image of the reconstructed view,

a map representative of the presence of occlusions in an image of the reconstructed view,

a map representative of confidence values associated with a depth map; and

generate said at least one image using said at least one read item of metadata and said decoded image.

12. A non-transitory computer-readable medium comprising a computer program stored thereon comprising instructions for implementing a decoding method when said program is executed by a processor of a decoding device, wherein the instructions configure the decoding device to:

decode a data stream representative of a multi-view video by:

obtain syntax elements from at least one part of the stream data;

reconstruct at least one image of a view of the video from the syntax elements obtained;

obtain at least one item of metadata in a predetermined form from at least one syntax element, said at least one item of metadata being associated with said at least one reconstructed image, and said at least one item of metadata corresponding to an item of information included in the group consisting of:

camera parameters,

decoded and scaled motion vectors,

a partitioning of the reconstructed image,

a reference image used by a block of an image of the reconstructed view,

the coding modes of an image of the reconstructed view,

quantization parameter values of an image of the reconstructed view,

prediction residual values of an image of the reconstructed view,

a map representative of movement in an image of the reconstructed view,

a map representative of presence of occlusions in an image of the reconstructed view,

a map representative of confidence values associated with a depth map; and

provide said at least one item of metadata to an image synthesis module, said image synthesis module being configured to synthesize at least one virtual view, distinct from the views of the multi-view video, from said at least one reconstructed image and said at least one item of metadata.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 16, 2021
From: JUNG, JOEL; NIKITIN, PAVEL; BOISSONADE, PATRICK
To: ORANGE
Reel/Frame 055943/0376 →
Priority Claims (1)
FR 1853829 · May 3, 2018 · national
Continuity (1)
Related Publication 20210243472A1 · Aug 5, 2021