IP Library Granted Patent US 12,231,619
Granted Patent B2
US 12,231,619 · App. 17/762,922 · Granted Feb 18, 2025

Prediction for video encoding and decoding using external reference

Inventors: Philippe Bordes (Laille, FR); Didier Doyen (Cesson-Sévigné, FR); Franck Galpin (Thorigne-Fouillard, FR); Michel Kerdranvat (Chantepie, FR)
Assignee: InterDigital CE Patent Holdings, SAS
H04N19/105H04N19/139H04N19/172H04N19/597
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,231,619
App. No.
17/762,922
Granted
Feb 18, 2025
Kind
B2
Abstract

Various embodiments relate to a video coding system in which some elements required for decoding are generated according to a process that not specified within the video coding system. This process is hereafter referenced to as being the “external” process. This external process may generate “external” reference pictures to be used by a decoder that is adapted to use these external pictures. Encoding method, decoding method, encoding apparatus, decoding apparatus based on this post-processing method are proposed.

Claims (37)

1. An apparatus comprising:

at least one processor configured to:

perform a first process for decoding an encoded stream comprising data representing a video, the first process comprising, for a current picture of the video: obtaining, from the encoded stream, information representative of a use of an external reference picture for reconstructing the current picture;

obtaining, from the encoded stream, information representative of an index in a set of reference pictures of an external reference picture to be used for reconstructing the current picture; and

reconstructing the current picture based on the external reference picture obtained from the set of reference pictures based on the obtained index, wherein the external reference picture is generated locally and not comprised in the encoded stream; and

perform a second process to generate at least one external reference picture.

2. The apparatus of claim 1 , wherein the video is a multi-view video, wherein the external reference picture comprises a texture of a first view and a motion vector map representing disparity information between the first view and a second view and wherein the at least one processor is further configured to reconstruct the second view based on using motion compensation based on the texture of the first view and the disparity information.

3. The apparatus of claim 2 , wherein the at least one processor is further configured to:

copy the first view into a decoded picture buffer of the second view;

associate the first view with a picture order count;

copy the disparity information into the motion information map of a reference picture being co-located; and

predict the second view based on the copied information.

4. The apparatus of claim 3 , wherein the at least one processor is further configured to predict the second view using a temporal motion vector prediction mode.

5. A method for decoding an encoded stream comprising data representing a video, the method comprising, for a current picture of the video:

obtaining, from the encoded stream, information representative of a use of an external reference picture for reconstructing the current picture;

obtaining, from the encoded stream, information representative of an index in a set of reference pictures of an external reference picture to be used for reconstructing the current picture; and

reconstructing the current picture based on the external reference picture obtained from the set of reference pictures based on the obtained index, wherein the external reference picture is generated locally and not comprised in the encoded stream.

6. The method of claim 5 , wherein the video is a multi-view video, wherein the external reference picture comprises a texture of a first view and a motion vector map representing disparity information between the first view and a second view and wherein reconstructing the second view is based on using motion compensation based on the texture of the first view and the disparity information.

7. The method of claim 6 , further comprising:

copying the first view into a decoded picture buffer of the second view;

associating the first view with a picture order count;

copying the disparity information into the motion information map of a reference picture being co-located; and

predicting the second view based on the copied information.

8. The method of claim 7 , wherein the predicting the second view uses a temporal motion vector prediction mode.

9. A non-transitory computer readable medium comprising program code instructions that, when executed by at least one processor, cause the least one processor to:

perform a first process for decoding an encoded stream comprising data representing a video, the instructions causing the at least one processor to, for a current picture of the video:

obtain, from the encoded stream, information representative of a use of an external reference picture for reconstructing the current picture;

obtain, from the encoded stream, information representative of an index in a set of reference pictures of an external reference picture to be used for reconstructing the current picture; and

reconstruct the current picture based on the external reference picture obtained from the set of reference pictures based on the obtained index, wherein the external reference picture is generated locally and not comprised in the encoded stream; and

perform a second process to generate at least one external reference picture.

10. The non-transitory computer readable medium of claim 9 , wherein the video is a multi-view video, wherein the external reference picture comprises a texture of a first view and a motion vector map representing disparity information between the first view and a second view, and wherein reconstructing the second view is based on using motion compensation based on the texture of the first view and the disparity information.

11. The non-transitory computer readable medium of claim 10 , wherein the program code instructions further cause the at least one processor to decode data representing the video by:

copying the first view into a decoded picture buffer of the second view;

associating the first view with a picture order count;

copying the disparity information into the motion information map of a reference picture being co-located; and

predicting the second view based on the copied information.

12. The non-transitory computer readable medium of claim 11 , wherein the predicting the second view uses a temporal motion vector prediction mode.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 1, 2023
From: INTERDIGITAL VC HOLDINGS FRANCE, SAS
To: INTERDIGITAL CE PATENT HOLDINGS, SAS
Reel/Frame 064460/0921 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 23, 2022
From: BORDES, PHILIPPE; DOYEN, DIDIER; GALPIN, FRANCK; KERDRANVAT, MICHEL
To: INTERDIGITAL VC HOLDINGS FRANCE, SAS
Reel/Frame 059377/0235 →
Priority Claims (1)
EP 19306164 · Sep 23, 2019 · regional
Continuity (1)
Related Publication 20220360771A1 · Nov 10, 2022
References Cited (19)
US 20110274174A1 · Francois · 2011 [cited by examiner]
US 20130223525A1 · Zhou · 2013 [cited by examiner]
US 20140003799A1 · Soroushian et al. · 2014 [cited by applicant]
US 20140036033A1 · Takahashi · 2014 [cited by examiner]
US 20180124408A1 · Choi · 2018 [cited by examiner]
US 20190215522A1 · Zhang · 2019 [cited by examiner]
US 20220038733A1 · Hannuksela · 2022 [cited by examiner]
EP 3264767A1 · 2018 [cited by applicant]
Tourapis et al., “Weighted prediction methods for improved motion compensation”, Institute of Electrical and Electronics Engineers (IEEE), 2009 16th IEEE International Conference on Image Processing (ICIP), Cairo, Egypt… [cited by applicant]
Boyce et al, “Draft High Efficiency Video Coding (HEVC) Version 2, Combined Format Range Extensions (RExt), Scalability (SHVC), and Multi-View (MV-HEVC) Extensions”, Joint Collaborative Team on Video Coding (JCT-VC) of … [cited by applicant]
Zuo et al., “Library Based Coding for Videos with Repeated Scenes”, Institute of Electrical and Electronics Engineers (IEEE), 2015 Picture Coding Symposium (PCS), Caims, QLD, Australia, May 31, 2015, 5 pages. [cited by applicant]
Bross et al., “Versatile Video Coding (Draft 6)”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Document: JVET-O2001-vE, 15th Meeting, Gothenburg, Sweden, Jul. 3, 2019, 455 pages. [cited by applicant]
Chen et al., “Algorithm description for Versatile Video Coding and Test Model 6 (VTM 6)”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Document: JVET-O2002-v2, 15th Meeting: Gothenb… [cited by applicant]
Seregin et al., “AHG17: On zero delta POC in reference picture structure”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Document: JVET-O0244-v1, 15th Meeting: Gothenburg, Sweden, Ju… [cited by applicant]
Vasconcelos et al., “Library-based coding: a representation for efficient video compression and retrieval”, Institute of Electrical and Electronics Engineers (IEEE), Proceedings DCC '97, Data Compression Conference, Sno… [cited by applicant]
Anonymous, “High Efficiency Video Coding”, ITU-T Telecommunication Standardization Section of ITU, Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Recommendat… [cited by applicant]
Anonymous, “Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video—Information Technology—Generic coding of moving pictures and associated audio information: Video”, … [cited by applicant]
Zhao et al., “Enhanced Motion-Compensated Video Coding with Deep Virtual Reference Frame Generation”, Institute of Electrical and Electronics Engineers (IEEE), IEEE Transactions on Image Processing, vol. 28, No. 10, Oct… [cited by applicant]
Anonymous, “Information technology—Generic coding of moving pictures and associated audio information: Systems”, International Telecommunication Union, ITU-T Telecommunication Standardization Sector of ITU, Series H: Au… [cited by applicant]