IP Library Granted Patent US 9,860,563
Granted Patent B2
US 9,860,563 · App. 15/257,447 · Granted Jan 2, 2018

Hybrid video coding supporting intermediate view synthesis

Inventors: Thomas Wiegand (Berlin, DE); Karsten Mueller (Berlin, DE); Philipp Merkle (Berlin, DE)
Assignee: GE VIDEO COMPRESSION, LLC
H04N19/597H04N13/0011H04N19/105H04N19/109H04N19/11H04N19/124H04N19/17H04N19/172H04N19/177H04N19/30H04N19/46H04N19/513H04N19/521H04N19/567H04N19/61H04N19/80H04N2013/0081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,860,563
App. No.
15/257,447
Granted
Jan 2, 2018
Kind
B2
Abstract

Hybrid video decoder supporting intermediate view synthesis of an intermediate view video from a first- and a second-view video which are predictively coded into a multi-view data signal with frames of the second-view video being spatially subdivided into sub-regions and the multi-view data signal having a prediction mode is provided, having: an extractor configured to respectively extract, from the multi-view data signal, for sub-regions of the frames of the second-view video, a disparity vector and a prediction residual; a predictive reconstructor configured to reconstruct the sub-regions of the frames of the second-view video, by generating a prediction from a reconstructed version of a portion of frames of the first-view video using the disparity vectors and a prediction residual for the respective sub-regions; and an intermediate view synthesizer configured to reconstruct first portions of the intermediate view video.

Claims (52)

1. A decoder for decoding multi-view data, comprising:

an extractor configured to:

receive multi-view data comprising data related to a first-view video and a second-view video,

obtain, from the multi-view data, a disparity vector associated with a sub-region of a first frame of the second-view video, wherein the disparity vector indicates a spatial displacement of the sub-region of the first frame of the second-view video with respect to a second frame of the first-view video, wherein the first and second frames are captured at a same time instant, and

reconstruct, based on the multi-view data, a portion of the second frame of the first view video; and

a view synthesizer configured to:

determine a scaled disparity vector using the disparity vector and a scaling factor, and

reconstruct a portion of a third frame of a third-view video using the reconstructed portion of the second frame of the first view video and the scaled disparity vector.

2. The decoder of claim 1 , wherein the scaling factor is a value between 0 and 1, and the view synthesizer is configured to multiply the disparity vector with the scaling factor to determine the scaled disparity vector.

3. The decoder of claim 1 , wherein the scaling factor is a value less than 0 or greater than 1, and the view synthesizer is configured to multiply the disparity vector with the scaling factor to determine the scaled disparity vector.

4. The decoder of claim 1 , wherein the scaling factor depends on a spatial location of a third view corresponding to the third-view video relative to a first view corresponding to the first-view video and a second view corresponding to the second-view video.

5. The decoder of claim 1 , further comprising a predictive reconstructor configured to reconstruct the sub-region of the first frame of the second-view video based on the reconstructed portion of the second frame of the first-view video, the disparity vector, and a prediction residual associated with the sub-region obtained from the multi-view data.

6. The decoder of claim 1 , wherein the view synthesizer is further configured to extrapolate or interpolate another portion of the third frame of the third-view video based on the reconstructed portion of the third frame of the third-view video.

7. The decoder of claim 1 , wherein the view synthesizer is further configured to:

determine an interpolated disparity vector based on a plurality of disparity vectors associated with different sub-regions of the second-view video, and

reconstruct another portion of the third frame of the third-view video using the reconstructed portion of the second frame of the first view video and the interpolated disparity vector.

8. The decoder of claim 7 , wherein the different sub-regions of the second-view video belong to a same frame of the second-view video.

9. The decoder of claim 7 , wherein the different sub-regions of the second-view video belong to different frames of the second-view video.

10. The decoder of claim 1 , wherein the extractor is further configured to extract, from the multi-view data, reliability data indicating a reliability factor for the disparity vector, wherein the view synthesizer is configured to reconstruct the portion of the third frame of the third-view video based on the reliability data satisfying a predetermined requirement.

11. An encoder for encoding multi-view data, comprising:

a prediction unit configured to encode multi-view data comprising data related to a first-view video and a second-view video,

wherein to encode the multi-view data, the prediction unit is at least configured to determine a disparity vector associated with a sub-region of a first frame of the second-view video, and a prediction residual associated with the sub-region based on the disparity vector, the disparity vector indicating a spatial displacement of the sub-region of the first frame of the second-view video with respect to a second frame of the first-view video, and the first and second frames are captured at a same time instant; and

a data signal generator configured to insert the multi-view data including the disparity vector and the prediction residual into a data stream,

wherein data associated with a portion of the second frame of the first-view video and a scaled disparity vector are used to synthesize a portion of a third frame of a third-view video, and the scaled disparity vector is determined using the disparity vector and a scaling factor.

12. The encoder of claim 11 , wherein the scaling factor is a value between 0 and 1, and the scaled disparity vector is determined by multiplying the disparity vector with the scaling factor.

13. The encoder of claim 11 , wherein the scaling factor is a value less than 0 or greater than 1, and the scaled disparity vector is determined by multiplying the disparity vector with the scaling factor.

14. The encoder of claim 11 , wherein the scaling factor depends on a spatial location of a third view corresponding to the third-view video relative to a first view corresponding to the first-view video and a second view corresponding to the second-view video.

15. The encoder of claim 11 , wherein the synthesized portion of the frame of the third view video is used to extrapolate or interpolate another portion of the third frame of the third-view video.

16. The encoder of claim 11 , wherein an interpolated disparity vector is determined based on a plurality of disparity vectors associated with different sub-regions of the second-view video, and another portion of the third frame of the third-view video is synthesized using the synthesized portion of the second frame of the first view video and the interpolated disparity vector.

17. The encoder of claim 16 , wherein the different sub-regions of the second-view video belong to a same frame of the second-view video.

18. The encoder of claim 16 , wherein the different sub-regions of the second-view video belong to different frames of the second-view video.

19. The encoder of claim 11 , wherein the prediction unit is further configured to encode, as part of the multi-view data, reliability data indicating a reliability factor for the disparity vector, wherein the portion of the third frame of the third-view video is synthesized based on the reliability data satisfying a predetermined requirement.

20. A method for decoding a video comprising:

receiving multi-view data comprising data related to a first-view video and a second-view video;

obtaining, from the multi-view data, a disparity vector associated with a sub-region of a first frame of the second-view video, wherein the disparity vector indicates a spatial displacement of the sub-region of the first frame of the second-view video with respect to a second frame of the first-view video, wherein the first and second frames are captured at a same time instant;

reconstructing, based on the multi-view data, a portion of the second frame of the first view video;

determining a scaled disparity vector using the disparity vector and a scaling factor; and

reconstructing a portion of a third frame of a third-view video using the reconstructed portion of the second frame of the first view video and the scaled disparity vector.

21. A method for encoding a video comprising:

encoding multi-view data comprising data related to a first-view video and a second-view video, the encoding multi-view data comprising:

determining a disparity vector associated with a sub-region of a first frame of the second-view video, the disparity vector indicating a spatial displacement of the sub-region of the first frame of the second-view video with respect to a second frame of the first-view video, the first and second frames are captured at a same time instant, and

determining a prediction residual associated with the sub-region based on the disparity vector; and

inserting the multi-view data including the disparity vector and the prediction residual into a data stream,

wherein data associated with a portion of the second frame of the first-view video and a scaled disparity vector are used to synthesize a portion of a third frame of a third-view video, and the scaled disparity vector is determined using the disparity vector and a scaling factor.

22. A computer program stored on a non-transitory computer-readable medium comprising a program code for performing, when executed on a computer, a method according to claim 20 .

23. A computer program stored on a non-transitory computer-readable medium comprising a program code for performing, when executed on a computer, a method according to claim 21 .

24. A non-transitory computer-readable medium for storing data associated with a video, comprising:

a data stream stored in the non-transitory computer-readable medium, the data stream comprising multi-view data including at least a first-view video and a second-view video, the multi-view data comprising a disparity vector associated with a sub-region of a first frame of the second-view video, and a prediction residual associated with the sub-region determined based on the disparity vector, wherein

the disparity vector indicates a spatial displacement of the sub-region of the first frame of the second-view video with respect to a second frame of the first-view video,

the first and second frames are captured at a same time instant,

data associated with a portion of the second frame of the first-view video and a scaled disparity vector are used to synthesize a portion of a third frame of a third-view video, and

the scaled disparity vector is determined using the disparity vector and a scaling factor.

Assignments (3)
CHANGE OF NAME Recorded Nov 26, 2024
From: GE VIDEO COMPRESSION, LLC
To: DOLBY VIDEO COMPRESSION, LLC
Reel/Frame 069450/0344 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 6, 2016
From: WIEGAND, THOMAS; MUELLER, KARSTEN; MERKLE, PHILIPP
To: FRAUNHOFER-GESELLSCHAFT ZUR FOERDERUNG DER ANGEWANDTEN FORSCHUNG E.V.
Reel/Frame 039641/0374 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 6, 2016
From: FRAUNHOFER-GESELLSCHAFT ZUR FOERDERUNG DER ANGEWANDTEN FORSCHUNG E.V.
To: GE VIDEO COMPRESSION, LLC
Reel/Frame 039641/0416 →
Continuity (4)
Continuation 14743094 · Jun 18, 2015
Continuation 13739365 · Jan 11, 2013
Continuation PCTEP2010060202 · Jul 15, 2010
Related Publication 20160381392A1 · Dec 29, 2016