IP Library Granted Patent US 12,323,622
Granted Patent B2
US 12,323,622 · App. 18/424,332 · Granted Jun 3, 2025

Hybrid video coding supporting intermediate view synthesis

Inventors: Thomas Wiegand (Berlin, DE); Karsten Mueller (Berlin, DE); Philipp Merkle (Berlin, DE)
Assignee: Dolby Video Compression, LLC
H04N19/597H04N13/111H04N19/105H04N19/109H04N19/11H04N19/124H04N19/17H04N19/172H04N19/177H04N19/30H04N19/46H04N19/513H04N19/521H04N19/567H04N19/61H04N19/80H04N2013/0081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,323,622
App. No.
18/424,332
Granted
Jun 3, 2025
Kind
B2
Abstract

Hybrid video decoder supporting intermediate view synthesis of an intermediate view video from a first- and a second-view video which are predictively coded into a multi-view data signal with frames of the second-view video being spatially subdivided into sub-regions and the multi-view data signal having a prediction mode is provided, having: an extractor configured to respectively extract, from the multi-view data signal, for sub-regions of the frames of the second-view video, a disparity vector and a prediction residual; a predictive reconstructor configured to reconstruct the sub-regions of the frames of the second-view video, by generating a prediction from a reconstructed version of a portion of frames of the first-view video using the disparity vectors and a prediction residual for the respective sub-regions; and an intermediate view synthesizer configured to reconstruct first portions of the intermediate view video.

Claims (29)

1. A decoder for decoding encoded information representing a multi-view video including a first-view video and a second-view video, comprising:

a predictive reconstructor configured to reconstruct a sub-region of a frame of the second-view video based on a reconstructed portion of the first-view video and a disparity vector, wherein the disparity vector indicates a displacement of the sub-region of the frame of the second-view video with respect to a corresponding frame of the first-view video; and

a view synthesizer configured to:

determine another disparity vector based on the disparity vector, and

synthesize a portion of a frame of a synthesized third-view video using the reconstructed portion of the frame of the first view video and the other disparity vector.

2. The decoder of claim 1 , wherein the other disparity vector is further based on a scaling factor, and the view synthesizer is configured to multiply the disparity vector with the scaling factor to determine the other disparity vector.

3. The decoder of claim 2 , wherein the scaling factor is a value less than 0 or greater than 1.

4. The decoder of claim 2 , wherein the scaling factor depends on a spatial location of a third view corresponding to the third-view video relative to a first view corresponding to the first-view video and a second view corresponding to the second-view video.

5. The decoder of claim 1 , wherein the other disparity vector and the disparity vector are equal.

6. The decoder of claim 1 , wherein the view synthesizer is further configured to extrapolate or interpolate another portion of the frame of the third-view video based on the reconstructed portion of the frame of the third-view video.

7. The decoder of claim 1 , wherein the view synthesizer is further configured to: determine an interpolated disparity vector based on a plurality of disparity vectors associated with different sub-regions of the second-view video, and reconstruct another portion of the frame of the third-view video using the reconstructed portion of the frame of the first view video and the interpolated disparity vector.

8. The decoder of claim 7 , wherein the different sub-regions of the second-view video belong to a same frame of the second-view video.

9. The decoder of claim 7 , wherein the different sub-regions of the second-view video belong to different frames of the second-view video.

10. The decoder of claim 1 , further comprising an extractor configured to extract, from the encoded information, reliability data indicating a reliability factor for the disparity vector, wherein the view synthesizer is configured to reconstruct the portion of the frame of the third-view video based on the reliability data satisfying a predetermined requirement.

11. An encoder for encoding information representing a multi-view video, comprising:

a prediction unit configured to encode information representing prediction coding of the multi-view video including data related to a first-view video and a second-view video,

wherein to encode the information, the prediction unit is configured to determine a disparity vector associated with a sub-region of a frame of the second-view video, the disparity vector indicating a displacement of the sub-region of the frame of the second-view video with respect to the first-view video; and

a data signal generator configured to insert the information including the disparity vector into a data stream, wherein data associated with a portion of the frame of the first-view video and a scaled disparity vector are used to synthesize a portion of a frame of a synthesized third-view video, and the scaled disparity vector is determined using the disparity vector.

12. The encoder of claim 11 , wherein the scaled disparity vector is determined further based on a scaling factor, and the scaled disparity vector is determined by multiplying the disparity vector with the scaling factor.

13. The encoder of claim 12 , wherein the scaling factor is a value less than 0 or greater than 1.

14. The encoder of claim 12 , wherein the scaling factor depends on a spatial location of a third view corresponding to the third-view video relative to a first view corresponding to the first-view video and a second view corresponding to the second-view video.

15. The encoder of claim 11 , wherein the synthesized portion of the frame of the third view video is used to extrapolate or interpolate another portion of the frame of the third-view video.

16. The encoder of claim 11 , wherein an interpolated disparity vector is determined based on a plurality of disparity vectors associated with different sub-regions of the second-view video, and another portion of the frame of the third-view video is synthesized using the synthesized portion of the frame of the first view video and the interpolated disparity vector.

17. The encoder of claim 16 , wherein the different sub-regions of the second-view video belong to a same frame of the second-view video.

18. The encoder of claim 16 , wherein the different sub-regions of the second-view video belong to different frames of the second-view video.

19. The encoder of claim 11 , wherein the prediction unit is further configured to encode reliability data indicating a reliability factor for the disparity vector, wherein the portion of the frame of the third-view video is synthesized based on the reliability data satisfying a predetermined requirement.

20. A non-transitory computer-readable medium for storing data associated with a video, comprising:

a data stream stored in the non-transitory computer-readable medium, the data stream comprising information representing a multi-view video including a first-view video and a second-view video, the information including a disparity vector associated with a sub-region of a frame of the second-view video, wherein the disparity vector indicates a displacement of the sub-region of the frame of the second-view video with respect to the first-view video,

wherein data associated with a portion of the frame of the first-view video and a scaled disparity vector are used to synthesize a portion of a frame of a synthesized third-view video, and the scaled disparity vector is determined using the disparity vector.

Assignments (2)
CHANGE OF NAME Recorded Jan 30, 2026
From: GE VIDEO COMPRESSION, LLC
To: DOLBY VIDEO COMPRESSION, LLC
Reel/Frame 074536/0781 →
CHANGE OF NAME Recorded Nov 26, 2024
From: GE VIDEO COMPRESSION, LLC
To: DOLBY VIDEO COMPRESSION, LLC
Reel/Frame 069450/0772 →
Continuity (9)
Continuation 17382862 · Jul 22, 2021
Continuation 16855058 · Apr 22, 2020
Continuation 16043887 · Jul 24, 2018
Continuation 15820687 · Nov 22, 2017
Continuation 15257447 · Sep 6, 2016
Continuation 14743094 · Jun 18, 2015
Continuation 13739365 · Jan 11, 2013
Continuation PCTEP2010060202 · Jul 15, 2010
Related Publication 20240292025A1 · Aug 29, 2024
References Cited (38)
US 5530774A · Fogel · 1996 [cited by applicant]
US 8682087B2 · Tian · 2014 [cited by applicant]
US 20050031035A1 · Vedula et al. · 2005 [cited by applicant]
US 20070014477A1 · Macinnis et al. · 2007 [cited by applicant]
US 20080043095A1 · Vetro et al. · 2008 [cited by applicant]
US 20080247462A1 · Demos · 2008 [cited by applicant]
US 20090129465A1 · Lai · 2009 [cited by applicant]
US 20100002948A1 · Gangwal · 2010 [cited by examiner]
US 20100080287A1 · Ali · 2010 [cited by applicant]
US 20100111183A1 · Jeon · 2010 [cited by applicant]
US 20100259596A1 · Park et al. · 2010 [cited by applicant]
US 20100309294A1 · Ihara · 2010 [cited by examiner]
US 20110001792A1 · Pandit · 2011 [cited by examiner]
US 20110064262A1 · Chen et al. · 2011 [cited by applicant]
US 20110134213A1 · Tsukagoshi · 2011 [cited by applicant]
US 20110222602A1 · Sung · 2011 [cited by applicant]
US 20110254921A1 · Pahalawatta et al. · 2011 [cited by applicant]
US 20110280316A1 · Chen et al. · 2011 [cited by applicant]
US 20120133736A1 · Nishi et al. · 2012 [cited by applicant]
US 20120250982A1 · Ito et al. · 2012 [cited by applicant]
US 20130022113A1 · Chen et al. · 2013 [cited by applicant]
US 20130148722A1 · Zhang et al. · 2013 [cited by applicant]
US 20130170552A1 · Kim et al. · 2013 [cited by applicant]
US 20130229485A1 · Rusanovskyy et al. · 2013 [cited by applicant]
US 20130242051A1 · Balogh · 2013 [cited by applicant]
US 20130279576A1 · Chen et al. · 2013 [cited by applicant]
US 20140028793A1 · Wiegand · 2014 [cited by applicant]
US 20140218473A1 · Hannuksela · 2014 [cited by applicant]
US 20150358598A1 · Lin · 2015 [cited by examiner]
US 20160309156A1 · Park · 2016 [cited by examiner]
US 20170085917A1 · Hannuksela · 2017 [cited by examiner]
US 20190089979A1 · Zhang · 2019 [cited by examiner]
Chang et ai, “Multi-view image compression and intermediate view synthesis for stereoscopic applications”, Circuits and Systems, 2000, Proceedings, ISCAS 2000 Geneva. The 2000 IEEE international symposium on May 28-31, … [cited by applicant]
Lie et al: “Intermediate view synthesis from binocular images for stereoscopic applications”, The 2001 IEEE Intemational symposium on circuits and systems, 2001, ISCAS 2001, vol. 5, May 6-9, 2001, pp. 287-290. [cited by applicant]
Ho et al., “Overview of Multi-view Video Coding,” 14th International Workshop on Systems, Signals and Images Processing and 6th EURASIP Conference Focused on Speech and Image Processing, Multimedia Communications and Se… [cited by applicant]
ISO/IEC JTC1/SC29f/WG11, Text of ISO/IEC 14496-10:2008/FDAM 1 Multiview Video Coding, Doc. N9978, Hannover, Germany, Jul. 2008, ITU-T and ISO/IEC JTCI, 69 pages. [cited by applicant]
Intemational Telecommunication Union, “Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video: Advanced video coding for generic audiovisual services,” ITU-T Recommen… [cited by applicant]
Huang, Yu., et al., “A Layered Method of Visibility Resolving in Depth Image-based Rendering”, ICPR 2008, pp. 1-4. [cited by applicant]