IP Library Granted Patent US 12666072
Granted Patent B2
US 12666072 · App. 18/573,015 · Granted Jun 23, 2026

Method for constructing a depth image from a multiview video, method for decoding a data stream representative of a multiview video, encoding method, devices, system, terminal equipment, signal and computer programs corresponding thereto

Inventors: Félix Henry (Châtillon Cedex, FR); Patrick Garus (Châtillon Cedex, FR)
Assignee: ORANGE
H04N19/517H04N19/119H04N19/139H04N19/176H04N19/46H04N19/597
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12666072
App. No.
18/573,015
Granted
Jun 23, 2026
Kind
B2
Abstract

A method for constructing a depth image associated with a view of a multiview video, called current view, from a data stream representative of the video. The stream includes information representative of the motion vectors of a texture image associated with the current view with respect to at least one reference texture image, the texture image having been divided into blocks. The method includes: obtaining the motion vectors from the information encoded in the stream; when at least one motion vector has been obtained for at least one block, called current block, of the texture image, motion-compensating a block of the depth image, co-located with the current block, from the at least one motion vector and at least one available reference depth image, the reference depth image being associated with the same view as the reference texture image.

Claims (64)

1 . A method implemented by at least one device and comprising:

constructing at least one block of a depth image associated with a view of a multiview video, called current view, said depth image not being encoded in a data stream representative of said video, said stream comprising information representative of motion vectors of a texture image associated with said current view with respect to at least one reference texture image, said texture image having been divided into blocks, wherein the constructing comprises:

obtaining said motion vectors from the information encoded in the stream; and

in response to at least one motion vector having been obtained for at least one block, called current block, of the texture image, constructing a block of the depth image, co-located with the current block, by motion compensation of at least one available reference depth image, pointed to by said at least one motion vector, said reference depth image being associated with a same view as said reference texture image.

2 . The method according to claim 1 , wherein the method comprises obtaining a motion compensation flag from an item of information encoded in the stream, said flag being associated with said block of the depth image and deciding to implement said motion compensation in response to the flag being set to a predetermined value.

3 . The method according to claim 1 , wherein the method comprises obtaining an identifier of the reference texture image by decoding an item of information encoded in the data stream and obtaining the reference depth image from said identifier.

4 . The method according to claim 1 , wherein the method further comprises:

decoding the encoded information representative of the motion vectors of the texture image associated with said current view; and

constructing the at least one block of the depth image associated with the current view at least from the decoded motion vectors.

5 . The method according to claim 4 , wherein the method comprises decoding an encoded item of information representative of a motion compensation flag of said at least one block of said depth image, said construction being implemented for said block in response to the flag being set to a predetermined value.

6 . A method implemented by at least one device and comprising:

encoding a data stream representative of a multi-view video, wherein the encoding comprises:

determining motion vectors of a texture image associated with a view of the multiview video, called current view, with respect to a reference texture image, said texture image having been divided into blocks;

encoding the motion vectors in the data stream;

obtaining a depth image associated with said current view, captured by a depth camera, called captured depth image;

constructing at least one motion-compensated block of a depth image associated with the current view, from at least one motion vector determined for at least one block of the texture image co-located with the block, called current block, wherein the constructing comprises:

obtaining, from the data stream, the motion vectors encoded in the stream; and

in response to at least one motion vector having been obtained for the current block, motion-compensating a block of the depth image, from said at least one motion vector and at least one available reference depth image, said reference depth image being associated with a same view as said reference texture image;

evaluating the motion-compensated block of said constructed depth image by comparison with the co-located block of the captured depth image, a compensation error being obtained; and

encoding an item of information representative of a motion compensation flag of said at least one block of said depth image depending on a predetermined error criterion, said flag being set to a predetermined value when the error criterion is satisfied.

7 . A device comprising:

at least one processor; and

at least one non-transitory computer readable medium comprising instructions stored thereon which when executed by the at least one processor implement constructing at least one block of a depth image associated with a view of a multiview video, called current view, said depth image not being encoded in a data stream representative of said video, said stream comprising encoded information representative of motion vectors of a texture image associated with said current view with respect to at least one reference texture image, said texture image having been divided into blocks, wherein the constructing comprises:

obtaining said motion vectors from the information encoded in the stream; and

in response to at least one motion vector having been obtained for at least one block, called current block, of the texture image, constructing a block of the depth image, co-located with the current block, by motion compensation of at least one available reference depth image, pointed to by said at least one motion vector, said reference depth image being associated with a same view as said reference texture image.

8 . The device according to claim 7 , wherein the instructions further configure the at least one processor to:

decode the encoded information representative of the motion vectors of the texture image associated with said current view; and

construct the at least one block of the depth image associated with the current view.

9 . A device comprising:

at least one processor; and

at least one non-transitory computer readable medium comprising instructions stored thereon which when executed by the at least one processor implement encoding a data stream representative of a multi-view video, wherein the encoding comprises:

determining motion vectors of a texture image associated with a view of the multiview video, called current view, with respect to a reference texture image, said texture image having been divided into blocks;

encoding the motion vectors in the data stream;

obtaining a depth image associated with said current view, captured by a depth camera, called captured depth image;

constructing at least one motion-compensated block of a depth image associated with the current view, or obtaining the at least one motion-compensated block having been constructed, from at least one motion vector determined for at least one block of the texture image co-located with the block by:

obtaining the motion vectors from the information encoded in the data stream; and

in response to at least one motion vector having been obtained for the current block, motion-compensating a block of the depth image, from said at least one motion vector and at least one available reference depth image, said reference depth image being associated with a same view as said reference texture image;

evaluating the motion-compensated block of said constructed depth image by comparison with the co-located block of the captured depth image, a compensation error being obtained; and

encoding an item of information representative of a motion compensation flag of said at least one block of said depth image depending on a predetermined error criterion, said flag being set to a predetermined value when the error criterion is satisfied.

10 . A system for free navigation in a multiview video of a scene, comprising:

a device comprising:

at least one processor; and

at least one non-transitory computer readable medium comprising instructions stored thereon which, when executed by the at least one processor, implement a method comprising:

decoding a data stream representative of the multi-view video, said stream comprising encoded information representative of motion vectors of a texture image of a current view with respect to a reference texture image, said texture image having been divided into blocks, no depth image co-located with the texture image being encoded in the data stream, and

in response to at least one motion vector having been obtained for at least one block, called current block, of the texture image, constructing a block of a depth image, co-located with the current block, by motion compensation of at least one available reference depth image, pointed to by said at least one motion vector, said reference depth image being associated with a same view as said reference texture image; and

a module configured to synthesize a view according to a viewpoint chosen by a user from the decoded texture images and the constructed depth images.

11 . Terminal equipment configured to receive an encoded data stream representative of a multiview video, the terminal equipment comprising:

a device comprising for free navigation in said multi-view video:

at least one processor; and

at least one non-transitory computer readable medium comprising instructions stored thereon which, when executed by the at least one processor, implement a method comprising:

decoding a data stream representative of the multi-view video, said stream comprising encoded information representative of motion vectors of a texture image of a current view with respect to a reference texture image, said texture image having been divided into blocks, no depth image co-located with the texture image being encoded in the data stream, and

in response to at least one motion vector having been obtained for at least one block, called current block, of the texture image, constructing a block of a depth image, co-located with the current block, by motion compensation of at least one available reference depth image, pointed to by said at least one motion vector, said reference depth image being associated with a same view as said reference texture image; and

a module configured to synthesize a view according to a viewpoint chosen by a user from the decoded texture images and the constructed depth images.

12 . A non-transitory computer readable medium comprising instructions of a computer program stored thereon for implementing a method according to claim 1 , when executed by a processor of the at least one device.

13 . A non-transitory computer readable medium comprising instructions of a computer program stored thereon for implementing a method according to claim 6 , when executed by a processor of the at least one device.

14 . A method implemented by at least one device and comprising:

decoding a data stream representative of a multiview video, said stream comprising encoded information representative of motion vectors of a texture image of a current view with respect to a reference texture image, said texture image having been divided into blocks, no depth image co-located with the texture image being encoded in the data stream, wherein the decoding comprises:

decoding the encoded information representative of the motion vectors of the texture image associated with said current view; and

transmitting the decoded information to a device comprising:

at least one processor; and

at least one non-transitory computer readable medium comprising instructions stored thereon which, when executed by the at least one processor, implement constructing at least one block of a depth image associated with said current view, wherein the constructing comprises:

obtaining the motion vectors; and

in response to at least one motion vector having been obtained for at least one block, called current block, of the texture image, constructing a block of the depth image, co-located with the current block, by motion compensation of at least one available reference depth image, pointed to by said at least one motion vector, said reference depth image being associated with a same view as said reference texture image.

15 . A non-transitory computer readable medium comprising instructions of a computer program stored thereon for implementing a method according to claim 14 , when executed by a processor of the at least one device.