METHOD AND APPARATUS FOR ENCODING AND DECODING A LARGE FIELD OF VIEW VIDEO
A method and an apparatus for coding a large field of view video into a bitstream are disclosed. At least one picture of said large field of view video is represented as a surface, said surface being projected onto at least one 2D picture using a projection function. For at least one current block of said at least one 2D picture, at least one neighbor block of said 2D picture not spatially adjacent to said current block in said 2D picture is determined from said projection function, and said at least one neighbor block is spatially adjacent to said current block on said surface. Said current block is then encoded using at least said determined neighbor block. Corresponding decoding method and apparatus are also disclosed.
1 . A method for coding a large field of view video into a bitstream, at least one picture of said large field of view video being represented as a surface, said surface being projected onto at least one 2D picture using a projection function, said method comprising, for at least one current block of said at least one 2D picture coded according to a current intra prediction mode m:
determining from said projection function, at least one neighbor block of said 2D picture, called first neighbor block C, not spatially adjacent to said current block in said 2D picture, said at least one neighbor block being spatially adjacent to said current block on said surface,
determining a list of most probable modes based on an intra prediction mode m_C of said first neighbor block C and further based on at least an intra prediction mode m_A of a second neighbor block A and an intra prediction mode m_B of a third neighbor block B, said second and third neighbor blocks being spatially adjacent to said current block in said 2D picture;
encoding said current intra prediction mode from said list of most probable modes.
2 . The method of claim 1 , wherein said determining a list of most probable modes comprises:
if m_A and m_B are different, determining the list as follows:
if m_C is equal to either m_A or m_B, the list of most probable modes comprises m_A and m_B and an additional intra prediction mode, said additional intra prediction mode being equal to a planar mode in the case where neither m_A nor m_B is a planar mode, being equal to a DC mode in the case where m_A or m_B is a planar mode but neither m_A nor m_B is a DC mode, being equal to a vertical intra prediction mode otherwise;
otherwise, the list of most probable modes comprises m_A, m_B and m_C;
if m_A and m_B are equal, determining the list as follows:
if m_C is equal to m_A, the list of most probable modes comprises m_A and two adjacent angular modes of m_A in the case where m_A is different from planar and DC modes, otherwise the list of most probable modes comprises planar mode, DC mode and vertical mode,
otherwise, the list of most probable modes comprises m_A and m_C and an additional intra prediction mode, said additional intra prediction mode being equal to a planar mode in the case where neither m_A nor m_C is a planar mode, being equal to a DC mode in the case where m_A or m_C is a planar mode but neither m_A nor m_C is a DC mode, being equal to a vertical intra prediction mode otherwise.
3 . The method of claim 1 , wherein encoding said current intra prediction mode comprises:
encoding a flag indicating whether said current intra prediction mode is equal to one mode of said list of most probable modes;
encoding an index identifying the most probable mode of said list equal to said current intra prediction mode in the case where said current intra prediction mode is equal to one mode of said list of most probable modes and encoding an index identifying the current intra prediction mode otherwise.
4 . The method according to claim 1 , further comprising coding an item of information relating to said projection function.
5 . The method according to claim 1 , wherein said 3D surface is a sphere and said projection function is an equi-rectangular projection.
6 . An apparatus for coding a large field of view video into a bitstream, at least one picture of said large field of view video being represented as a surface, said surface being projected onto at least one 2D picture using a projection function, said apparatus comprising one or more processors configured to:
determine from said projection function, for at least one current block of said at least one 2D picture coded according to a current intra prediction mode m, at least one neighbor block of said 2D picture, called first neighbor block C, not spatially adjacent to said current block in said 2D picture, said at least one neighbor block being spatially adjacent to said current block on said surface,
determine a list of most probable modes based on an intra prediction mode m_C of said first neighbor block C and further based on at least an intra prediction mode m_A of a second neighbor block A and on an intra prediction mode m_B of a third neighbor block B, said second and third neighbor blocks being spatially adjacent to said current block in said 2D picture;
encode said current intra prediction mode from said list of most probable modes.
7 . The apparatus of claim 6 , wherein the list of most probable modes is determined as follows:
if m_A and m_B are different:
if m_C is equal to either m_A or m_B, the list of most probable modes comprises m_A and m_B and an additional intra prediction mode, said additional intra prediction mode being equal to a planar mode in the case where neither m_A nor m_B is a planar mode, being equal to a DC mode in the case where m_A or m_B is a planar mode but neither m_A nor m_B is a DC mode, being equal to a vertical intra prediction mode otherwise;
otherwise, the list of most probable modes comprises m_A, m_B and m_C;
if m_A and m_B are equal:
if m_C is equal to m_A, the list of most probable modes comprises m_A and two adjacent angular modes of m_A in the case where m_A is different from planar and DC modes, otherwise the list of most probable modes comprises planar mode, DC mode and vertical mode,
otherwise, the list of most probable modes comprises m_A and m_C and an additional intra prediction mode, said additional intra prediction mode being equal to a planar mode in the case where neither m_A nor m_C is a planar mode, being equal to a DC mode in the case where m_A or m_C is a planar mode but neither m_A nor m_C is a DC mode, being equal to a vertical intra prediction mode otherwise.
8 . The apparatus according to claim 6 , wherein encoding said current intra prediction mode comprises:
encoding a flag indicating whether said current intra prediction mode is equal to one mode of said list of most probable modes;
encoding an index identifying the most probable mode of said list equal to said current intra prediction mode in the case where said current intra prediction mode is equal to one mode of said list of most probable modes and encode an index identifying the current intra prediction mode otherwise.
9 . The apparatus according to claim 6 , wherein said encoding of said current intra prediction mode further comprises encoding an item of information relating to said projection function.
10 . The apparatus according to claim 6 , wherein said 3D surface is a sphere and said projection function is an equi-rectangular projection.
11 . A method for decoding a bitstream representative of a large field of view video, at least one picture of said large field of view video being represented as a surface, said surface being projected onto at least one 2D picture using a projection function, said method comprising, for at least one current block of said at least one 2D picture coded according to a current intra prediction mode m:
determining from said projection function, at least one neighbor block of said 2D picture, called first neighbor block C, not spatially adjacent to said current block in said 2D picture, said at least one neighbor block being spatially adjacent to said current block on said surface,
determining a list of most probable modes based on an intra prediction mode m_C of said first neighbor block C and further based on at least an intra prediction mode m_A of a second neighbor block A and on an intra prediction mode m_B of a third neighbor block B, said second and third neighbor blocks being spatially adjacent to said current block in said 2D picture; and
decoding said current intra prediction mode from said list of most probable modes.
12 . The method of claim 11 , wherein said determining a list of most probable modes comprises:
if m_A and m_B are different, determining the list as follows:
if m_C is equal to either m_A or m_B, the list of most probable modes comprises m_A and m_B and an additional intra prediction mode, said additional intra prediction mode being equal to a planar mode in the case where neither m_A nor m_B is a planar mode, being equal to a DC mode in the case where m_A or m_B is a planar mode but neither m_A nor m_B is a DC mode, being equal to a vertical intra prediction mode otherwise;
otherwise, the list of most probable modes comprises m_A, m_B and m_C;
if m_A and m_B are equal, determining the list as follows:
if m_C is equal to m_A, the list of most probable modes comprises m_A and two adjacent angular modes of m_A in the case where m_A is different from planar and DC modes, otherwise the list of most probable modes comprises planar mode, DC mode and vertical mode,
otherwise, the list of most probable modes comprises m_A and m_C and an additional intra prediction mode, said additional intra prediction mode being equal to a planar mode in the case where neither m_A nor m_C is a planar mode, being equal to a DC mode in the case where m_A or m_C is a planar mode but neither m_A nor m_C is a DC mode, being equal to a vertical intra prediction mode otherwise.
13 . The method of claim 11 , wherein decoding said current intra prediction mode comprises:
decoding a flag indicating whether said current intra prediction mode is equal to one mode of said list of most probable modes;
decoding an index identifying the most probable mode of said list equal to said current intra prediction mode in the case where said current intra prediction mode is equal to one mode of said list of most probable modes and decoding an index identifying the current intra prediction mode otherwise.
14 . The method according to claim 11 , further comprising decoding an item of information relating to said projection function.
15 . The method according to claim 11 , wherein said 3D surface is a sphere and said projection function is an equi-rectangular projection.
16 . An apparatus for decoding a bitstream representative of a large field of view video, at least one picture of said large field of view video being represented as a surface, said surface being projected onto at least one 2D picture using a projection function, said apparatus comprising one or more processors configured to:
determine from said projection function, for at least one current block of said at least one 2D picture coded according to a current intra prediction mode m, at least one neighbor block of said 2D picture, called first neighbor block C, not spatially adjacent to said current block in said 2D picture, said at least one neighbor block being spatially adjacent to said current block on said surface,
determine a list of most probable modes based on an intra prediction mode m_C of said first neighbor block C and further based on at least an intra prediction mode m_A of a second neighbor block A and on an intra prediction mode m_B of a third neighbor block B, said second and third neighbor blocks being spatially adjacent to said current block in said 2D picture; and
decode said current intra prediction mode from said list of most probable modes.
17 . The apparatus of claim 16 , wherein the list of most probable modes is determined as follows:
if m_A and m_B are different:
if m_C is equal to either m_A or m_B, the list of most probable modes comprises m_A and m_B and an additional intra prediction mode, said additional intra prediction mode being equal to a planar mode in the case where neither m_A nor m_B is a planar mode, being equal to a DC mode in the case where m_A or m_B is a planar mode but neither m_A nor m_B is a DC mode, being equal to a vertical intra prediction mode otherwise;
otherwise, the list of most probable modes comprises m_A, m_B and m_C;
if m_A and m_B are equal:
if m_C is equal to m_A, the list of most probable modes comprises m_A and two adjacent angular modes of m_A in the case where m_A is different from planar and DC modes, otherwise the list of most probable modes comprises planar mode, DC mode and vertical mode,
otherwise, the list of most probable modes comprises m_A and m_C and an additional intra prediction mode, said additional intra prediction mode being equal to a planar mode in the case where neither m_A nor m_C is a planar mode, being equal to a DC mode in the case where m_A or m_C is a planar mode but neither m_A nor m_C is a DC mode, being equal to a vertical intra prediction mode otherwise.
18 . The apparatus of claim 17 , wherein decoding of said current intra prediction mode comprises:
decoding a flag indicating whether said current intra prediction mode is equal to one mode of said list of most probable modes;
decoding an index identifying the most probable mode of said list equal to said current intra prediction mode in the case where said current intra prediction mode is equal to one mode of said list of most probable modes and decode an index identifying the current intra prediction mode otherwise.
19 . The apparatus according to claim 16 , wherein decoding said current intra prediction mode comprises decoding an item of information relating to said projection function.
20 . The apparatus according to claim 16 , wherein said 3D surface is a sphere and said projection function is an equi-rectangular projection.
21 . An immersive rendering device comprising an apparatus for decoding a bitstream representative of a large field of view video according to claim 16 .
22 . A system for immersive rendering of a large field of view video encoded into a bitstream, comprising at least:
a network interface for receiving said bitstream from a data network,
an apparatus for decoding said bitstream according to claim 16 ,
an immersive rendering device.