Method and apparatus of view synthesis prediction in 3D video coding
A method and apparatus for a three-dimensional encoding or decoding system incorporating view synthesis prediction (VSP) with reduced computational complexity and/or memory access bandwidth are disclosed. The system applies the VSP process to the texture data only and applies non-VSP process to the depth data. Therefore, when a current texture block in a dependent view is coded according to VSP by backward warping the current texture block to the reference picture using an associated depth block and the motion parameter inheritance (MPI) mode is selected for the corresponding depth block in the dependent view, the corresponding depth block in the dependent view is encoded or decoded using non-VSP inter-view prediction based on motion information inherited from the current texture block.
1. A method for three-dimensional or multi-view video encoding, the method comprising:
receiving a reference picture in a reference view;
receiving input data associated with a first texture block and a second texture block in a dependent view;
deriving a first disparity vector (DV) from a set of neighboring blocks of the first texture block;
locating a first depth block from a reference depth map in the reference view according to the first DV and a location of the first texture block;
generating view synthesis prediction (VSP) data for the first texture block by backward warping the first texture block to the reference picture using the first depth block,
wherein the first depth block is located from a reference depth map in the dependent view according to a location of the first texture block and a selected disparity vector (DV),
wherein the selected DV is derived using a Neighboring Block Disparity Vector (NBDV) process,
wherein the selected DV is selected based on a first available DV from a set of neighboring blocks of the current texture block,
wherein a selection process for the selected DV is determined adaptively in a sequence level, picture level, slice level, largest coding unit (LCU) level, coding unit (CU) level, prediction unit (PU) level, Macroblock level, or sub-block level, and
wherein the selected DV that is derived using the NBDV process is utilized to fetch the depth block for VSP data generation, wherein the depth block is accessed by the selected DV only once for each texture block;
encoding the first texture block using the VSP data;
deriving a refined DV from a maximum value of a second depth block located according to a second DV derived from a set of neighboring blocks of the second texture block;
deriving an inter-view Merge candidate using the refined DV and a location of the second texture block to locate a refined depth block from the reference depth map;
encoding the second texture block using the inter-view Merge candidate; and
encoding a corresponding depth block in the dependent view using non-VSP interview prediction based on motion information inherited from the first texture block while the first texture block is encoded using the VSP data,
wherein the corresponding depth block is collocated with the first texture block.
2. The method of claim 1 , wherein the selected DV is derived using a Depth oriented Neighboring Block Disparity Vector (DoNBDV) process, wherein a derived DV is selected based on a first available DV from a set of neighboring blocks of the current texture block, a selected depth block is located from the reference depth map according to the derived DV and the location of the current texture block, and the selected DV is derived from a maximum value of the selected depth block.
3. The method of claim 1 , wherein a syntax element is used to indicate the selection process for the selected DV.
4. The method of claim 1 , wherein the selection process for the selected DV is implicitly decided at encoder side and decoder side.
5. The method of claim 1 , wherein the current texture block is divided into texture sub-blocks and each sub-block is predicted by sub-block VSP data generated by backward warping said each texture sub-block to the reference picture using the associated depth block.
6. The method of claim 1 , wherein the current texture block corresponds to a prediction unit (PU).
7. The method of claim 1 , wherein a derived DV is selected based on a first available DV from a set of neighboring blocks of the current texture block, a selected depth block is located from a reference depth map in the reference view according to the derived DV and the location of the current texture block, a refined DV is derived from a maximum value of the selected depth block, the refined DV and the location of the current texture block is used to locate a refined depth block from the reference depth map for deriving an interview Merge candidate.
8. The method of claim 1 , wherein encoding the corresponding depth block using non-VSP inter-view prediction based on motion information inherited from the current texture block when a motion parameter inheritance (MPI) mode is selected to code the corresponding depth block.
9. A system comprising:
one or more electronic circuits configured to:
receive a reference picture in a reference view;
receive input data associated with a first texture block and a second texture block in a dependent view;
derive a first disparity vector (DV) from a set of neighboring blocks of the first texture block;
locate a first depth block from a reference depth map in the reference view according to the first DV and a location of the first texture block;
generate view synthesis prediction (VSP) data for the first texture block by backward warping the first texture block to the reference picture using the first depth block,
wherein the first depth block is located from a reference depth map in the dependent view according to a location of the first texture block and a selected disparity vector (DV),
wherein the selected DV is derived using a Neighboring Block Disparity Vector (NBDV) process, wherein the selected DV is selected based on a first available DV from a set of neighboring blocks of the current texture block,
wherein a selection process for the selected DV is determined adaptively in a sequence level, picture level, slice level, largest coding unit (LCU) level, coding unit (CU) level, prediction unit (PU) level, Macroblock level, or sub-block level, and
wherein the selected DV that is derived using the NBDV process is utilized to fetch the depth block for VSP data generation, wherein the depth block is accessed by the selected DV only once for each texture block;
encode the first texture block using the VSP data;
derive a refined DV from a maximum value of a second depth block located according to a second DV derived from a set of neighboring blocks of the second texture block;
derive an inter-view Merge candidate using the refined DV and a location of the second texture block to locate a refined depth block from the reference depth map;
encode the second texture block using the inter-view Merge candidate; and
encode a corresponding depth block in the dependent view using non-VSP inter-view prediction based on motion information inherited from the first texture block while the first texture block is encoded using the VSP data,
wherein the corresponding depth block is collocated with the first texture block.
10. An apparatus comprising:
processing circuitry configured to:
receive a reference picture in a reference view;
receive input data associated with a first texture block and a second texture block in a dependent view;
derive a first disparity vector (DV) from a set of neighboring blocks of the first texture block;
locate a first depth block from a reference depth map in the reference view according to the first DV and a location of the first texture block;
generate view synthesis prediction (VSP) data for the first texture block by backward warping the first texture block to the reference picture using the first depth block,
wherein the first depth block is located from a reference depth map in the dependent view according to a location of the first texture block and a selected disparity vector (DV),
wherein the selected DV is derived using a Neighboring Block Disparity Vector (NBDV) process,
wherein the selected DV is selected based on a first available DV from a set of neighboring blocks of the current texture block,
wherein a selection process for the selected DV is determined adaptively in a sequence level, picture level, slice level, largest coding unit (LCU) level, coding unit (CU) level, prediction unit (PU) level, Macroblock level, or sub-block level, and
wherein the selected DV that is derived using the NBDV process is utilized to fetch the depth block for VSP data generation,
wherein the depth block is accessed by the selected DV only once for each texture block;
encode the first texture block using the VSP data;
derive a refined DV from a maximum value of a second depth block located according to a second DV derived from a set of neighboring blocks of the second texture block;
derive an inter-view Merge candidate using the refined DV and a location of the second texture block to locate a refined depth block from the reference depth map;
encode the second texture block using the inter-view Merge candidate; and
encode a corresponding depth block in the dependent view using non-VSP inter-view prediction based on motion information inherited from the first texture block while the first texture block is encoded using the VSP data,
wherein the corresponding depth block is collocated with the first texture block.
11. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors of a computing device, cause the computing device to:
receive a reference picture in a reference view;
receive input data associated with a first texture block and a second texture block in a dependent view;
derive a first disparity vector (DV) from a set of neighboring blocks of the first texture block;
locate a first depth block from a reference depth map in the reference view according to the first DV and a location of the first texture block;
generate view synthesis prediction (VSP) data for the first texture block by backward warping the first texture block to the reference picture using the first depth block,
wherein the first depth block is located from a reference depth map in the dependent view according to a location of the first texture block and a selected disparity vector (DV),
wherein the selected DV is derived using a Neighboring Block Disparity Vector (NBDV) process,
wherein the selected DV is selected based on a first available DV from a set of neighboring blocks of the current texture block,
wherein a selection process for the selected DV is determined adaptively in a sequence level, picture level, slice level, largest coding unit (LCU) level, coding unit (CU) level, prediction unit (PU) level, Macroblock level, or sub-block level, and
wherein the selected DV that is derived using the NBDV process is utilized to fetch the depth block for VSP data generation, wherein the depth block is accessed by the selected DV only once for each texture block;
encode the first texture block using the VSP data;
derive a refined DV from a maximum value of a second depth block located according to a second DV derived from a set of neighboring blocks of the second texture block;
derive an inter-view Merge candidate using the refined DV and a location of the second texture block to locate a refined depth block from the reference depth map;
encode the second texture block using the inter-view Merge candidate; and
encode a corresponding depth block in the dependent view using non-VSP inter-view prediction based on motion information inherited from the first texture block while the first texture block is encoded using the VSP data,
wherein the corresponding depth block is collocated with the first texture block.