IP Library › Granted Patent US 12,389,027
Granted Patent B2
US 12,389,027 · App. 18/492,501 · Granted Aug 12, 2025

Apparatus, a method and a computer program for video coding and decoding

Inventors: Miska Matias Hannuksela (Tampere, FI); Kemal Ugur (Istanbul, TR)
Assignee: NOKIA TECHNOLOGIES OY
H04N19/463H04N19/30H04N19/61H04N19/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,389,027
App. No.
18/492,501
Granted
Aug 12, 2025
Kind
B2
Abstract

A method comprising encoding a bitstream comprising a base layer, a first enhancement layer and a second enhancement layer; encoding an indication of both the base layer and the first enhancement layer used for prediction for the second enhancement layer in the bitstream; encoding, in the bitstream, an indication of a first set of prediction types that is applicable from the base layer to the second enhancement layer, wherein the first set of prediction types is a subset of all prediction types available for prediction between layers, and encoding, in the bitstream, an indication of a second set of prediction types that is applicable from the base layer or the first enhancement layer to the second enhancement layer, wherein the second set of prediction types is a subset of all prediction types available for prediction between layers.

Claims (80)

1. A method comprising:

encoding a bitstream comprising a base layer, a first enhancement layer and a second enhancement layer;

encoding, in the bitstream, an indication of a number of bits in a prediction type mask syntax element;

encoding, in the bitstream, a first prediction type mask syntax element for a first set of prediction types that is applicable from the base layer to the second enhancement layer, wherein the first set of prediction types is a subset of all prediction types available for prediction between layers, and wherein a prediction type of the first set is represented by a distinctive bit number in a bit mask; and

encoding, in the bitstream, a second prediction type mask syntax element for a second set of prediction types that is applicable from the base layer or the first enhancement layer to the second enhancement layer, wherein the second set of prediction types is a subset of all prediction types available for prediction between layers, and wherein a prediction type of the second set is represented by a distinctive bit number in a bit mask, and

wherein said prediction types of the first set that are available for prediction between layers are adaptively selectable as at least one of the following: sample prediction, motion information prediction or filtering parameter prediction.

2. The method according to claim 1 , wherein the first prediction type mask syntax element is included in a sequence-level syntax structure.

3. The method according to claim 2 , wherein the sequence-level syntax structure comprises at least one of a sequence parameter set or a video parameter set.

4. The method according to claim 1 , further comprising:

encoding a picture of the base layer and a picture of the first enhancement layer; and

encoding a picture of said second enhancement layer using said first set of prediction types from the picture of the base layer and said second set of prediction types from the picture of the first enhancement layer.

5. The method according to claim 1 ,

wherein each of said prediction types available for prediction between layers is represented by a bit number in the first prediction type mask syntax and the second prediction type mask syntax.

6. The method according to claim 1 , wherein said indication of the first set of prediction types and said indication of the second set of prediction types are included in at least one of a sequence parameter set or a video parameter set.

7. The method according to claim 1 , further comprising encoding, in the bitstream, an indication of at least one set of prediction types that is not applicable from the base layer or the first enhancement layer to the second enhancement layer.

8. The method according to claim 1 , wherein the second enhancement layer enhances a first scalability type relative to the base layer and a second scalability type relative to the first enhancement layer, and wherein the first scalability type and second scalability type are selected from at least one of: temporal scalability, quality scalability, spatial scalability, view scalability, depth enhancements, bit-depth scalability, chroma format scalability or color gamut scalability.

9. The method according to claim 8 , wherein the prediction types available for prediction between the second enhancement layer and the base layer are dependent on the first scalability type, and wherein the prediction types available for prediction between the second enhancement layer and the first enhancement layer are dependent on the second scalability type.

10. The method according to claim 1 , wherein the first set of prediction types has a first prediction direction and the second set of prediction types has a second prediction direction, and wherein said first prediction direction and second prediction direction are one of the following: temporal prediction, inter-view prediction, inter-layer prediction or intercomponent prediction.

11. An apparatus comprising:

at least one processor and at least one memory, said at least one memory stored with code thereon, which when executed by said at least one processor, causes the apparatus to perform:

encoding a bitstream comprising a base layer, a first enhancement layer and a second enhancement layer;

encoding, in the bitstream, an indication of a number of bits in a prediction type mask syntax element;

encoding, in the bitstream, a first prediction type mask syntax element for a first set of prediction types that is applicable from the base layer to the second enhancement layer, wherein the first set of prediction types is a subset of all prediction types available for prediction between layers, and wherein a prediction type of the first set is represented by a distinctive bit number in a bit mask; and

encoding, in the bitstream, a second prediction type mask syntax element for a second set of prediction types that is applicable from the base layer or the first enhancement layer to the second enhancement layer, wherein the second set of prediction types is a subset of all prediction types available for prediction between layers, and wherein a prediction type of the second set is represented by a distinctive bit number in a bit mask, and

wherein said prediction types of the first set that are available for prediction between layers are adaptively selectable as at least one of the following: sample prediction, motion information prediction or filtering parameter prediction.

12. The apparatus according to claim 11 , wherein the first prediction type mask syntax element is included in a sequence-level syntax structure.

13. The apparatus according to claim 12 , wherein the sequence-level syntax structure comprises at least one of a sequence parameter set or a video parameter set.

14. The apparatus according to claim 11 , wherein the apparatus is also caused to:

encode a picture of the base layer and a picture of the first enhancement layer; and

encode a picture of said second enhancement layer using said first set of prediction types from the picture of the base layer and said second set of prediction types from the picture of the first enhancement layer.

15. The apparatus according to claim 11 ,

wherein each of said prediction types available for prediction between layers is represented by a bit number in the first prediction type mask syntax and the second prediction type mask syntax.

16. The apparatus according to claim 11 , wherein said indication of the first set of prediction types and said indication of the second set of prediction types are included in at least one of a sequence parameter set or a video parameter set.

17. The apparatus according to claim 11 , wherein the apparatus is further configured to encode, in the bitstream, an indication of at least one set of prediction types that is not applicable from the base layer or the first enhancement layer to the second enhancement layer.

18. The apparatus according to claim 11 , wherein the second enhancement layer enhances a first scalability type relative to the base layer and a second scalability type relative to the first enhancement layer, and wherein the first scalability type and second scalability type are selected from at least one of: temporal scalability, quality scalability, spatial scalability, view scalability, depth enhancements, bit-depth scalability, chroma format scalability or color gamut scalability.

19. The apparatus according to claim 18 wherein the prediction types available for prediction between the second enhancement layer and the base layer are dependent on the first scalability type, and wherein the prediction types available for prediction between the second enhancement layer and the first enhancement layer are dependent on the second scalability type.

20. The apparatus according to claim 11 , wherein the first set of prediction types has a first prediction direction and the second set of prediction types has a second prediction direction, and wherein said first prediction direction and second prediction direction are one of the following: temporal prediction, inter-view prediction, inter-layer prediction or inter-component prediction.

21. A non-transitory computer readable storage medium stored with code thereon for use by an apparatus, which when executed by a processor, causes the apparatus to perform:

encoding a bitstream comprising a base layer, a first enhancement layer and a second enhancement layer;

encoding, in the bitstream, an indication of a number of bits in a prediction type mask syntax element;

encoding, in the bitstream, a first prediction type mask syntax element for a first set of prediction types that is applicable from the base layer to the second enhancement layer, wherein the first set of prediction types is a subset of all prediction types available for prediction between layers, and wherein a prediction type of the first set is represented by a distinctive bit number in a bit mask; and

encoding, in the bitstream, a second prediction type mask syntax for a second set of prediction types that is applicable from the base layer or the first enhancement layer to the second enhancement layer, wherein the second set of prediction types is a subset of all prediction types available for prediction between layers, and wherein a prediction type of the second set is represented by a distinctive bit number in a bit mask, and

wherein said prediction types of the first set that are available for prediction between layers are adaptively selectable as at least one of the following: sample prediction, motion information prediction or filtering parameter prediction.

22. A method comprising:

decoding, from a bitstream, an indication of a number of bits in prediction type mask syntax elements;

decoding, from the bitstream, a first prediction type mask syntax element for a first set of prediction types that is applicable from a base layer to a second enhancement layer, wherein the first set of prediction types is a subset of all prediction types available for prediction between layers, and wherein a prediction type of the first set is represented by a distinctive bit number in a bit mask;

decoding, from the bitstream, a second prediction type mask syntax element for a second set of prediction types that is applicable from a first enhancement layer to the second enhancement layer, wherein the second set of prediction types is a subset of all prediction types available for prediction between layers, and wherein a prediction type of the second set is represented by a distinctive bit number in a bit mask; and

decoding the second enhancement layer using said first set of prediction types from the base layer and said second set of prediction types from the first enhancement layer,

wherein said prediction types of the first set that are available for prediction between layers are at least one of the following: sample prediction, motion information prediction, sample adaptive offset parameter prediction or intra mode information prediction.

23. The method according to claim 22 , wherein the first prediction type mask syntax element is included in a sequence-level syntax structure.

24. The method according to claim 23 , wherein the sequence-level syntax structure comprises at least one of a sequence parameter set or a video parameter set.

25. The method according to claim 22 ,

wherein each of said prediction types available for prediction between layers is represented by a bit number in the first prediction type mask syntax and the second prediction type mask syntax.

26. The method according to claim 22 , wherein said indication of the first set of prediction types and said indication of the second set of prediction types are decoded from at least one of a sequence parameter set or a video parameter set.

27. The method according to claim 22 , further comprising decoding, from the bitstream, an indication of at least one set of prediction types that is not applicable from the base layer or the first enhancement layer to the second enhancement layer.

28. The method according to claim 22 , wherein the second enhancement layer enhances a first scalability type relative to the base layer and a second scalability type relative to the first enhancement layer, and wherein the first scalability type and second scalability type are selected from at least one of: temporal scalability, quality scalability, spatial scalability, view scalability, depth enhancements, bit-depth scalability, chroma format scalability or color gamut scalability.

29. The method according to claim 28 , wherein the prediction types available for prediction between the second enhancement layer and the base layer are dependent on the first scalability type, and wherein the prediction types available for prediction between the second enhancement layer and the first enhancement layer are dependent on the second scalability type.

30. The method according to claim 22 , wherein the first set of prediction types has a first prediction direction and the second set of prediction types has a second prediction direction, and wherein said first prediction direction and second prediction direction are one of the following: temporal prediction, inter-view prediction, inter-layer prediction or inter-component prediction.

31. An apparatus comprising:

at least one processor and at least one memory, said at least one memory stored with code thereon, which when executed by said at least one processor, causes the apparatus to perform:

decoding, from a bitstream, an indication of a number of bits in prediction type mask syntax elements;

decoding, from the bitstream, a first prediction type mask syntax element for a first set of prediction types that is applicable from a base layer to a second enhancement layer, wherein the first set of prediction types is a subset of all prediction types available for prediction between layers, and wherein a prediction type of the first set is represented by a distinctive bit number in a bit mask;

decoding, from the bitstream, a second prediction type mask syntax element for a second set of prediction types that is applicable from a first enhancement layer to the second enhancement layer, wherein the second set of prediction types is a subset of all prediction types available for prediction between layers, and wherein a prediction type of the second set is represented by a distinctive bit number in a bit mask; and

decoding the second enhancement layer using said first set of prediction types from the base layer and said second set of prediction types from the first enhancement layer,

wherein said prediction types of the first set that are available for prediction between layers are at least one of the following: sample prediction, motion information prediction, sample adaptive offset parameter prediction or intra mode information prediction.

32. The apparatus according to claim 31 , wherein the first prediction type mask syntax element is included in a sequence-level syntax structure.

33. The apparatus according to claim 32 , wherein the sequence-level syntax structure comprises at least one of a sequence parameter set or a video parameter set.

34. The apparatus according to claim 31 ,

wherein each of said prediction types available for prediction between layers is represented by a bit number in the first prediction type mask syntax and the second prediction type mask syntax.

35. The apparatus according to claim 31 , wherein said indication of the first set of prediction types and said indication of the second set of prediction types are decoded from at least one of a sequence parameter set or a video parameter set.

36. The apparatus according to claim 31 , wherein the apparatus is further caused to decode, from the bitstream, an indication of at least one set of prediction types that is not applicable from the base layer or the first enhancement layer to the second enhancement layer.

37. The apparatus according to claim 31 , wherein the second enhancement layer enhances a first scalability type relative to the base layer and a second scalability type relative to the first enhancement layer, and wherein the first scalability type and second scalability type are selected from at least one of: temporal scalability, quality scalability, spatial scalability, view scalability, depth enhancements, bit-depth scalability, chroma format scalability or color gamut scalability.

38. The apparatus according to claim 37 , wherein the prediction types available for prediction between the second enhancement layer and the base layer are dependent on the first scalability type, and wherein the prediction types available for prediction between the second enhancement layer and the first enhancement layer are dependent on the second scalability type.

39. The apparatus according to claim 31 , wherein the first set of prediction types has a first prediction direction and the second set of prediction types has a second prediction direction, and wherein said first prediction direction and second prediction direction are one of the following: temporal prediction, inter-view prediction, inter-layer prediction or inter-component prediction.

40. A non-transitory computer readable storage medium stored with code thereon for use by an apparatus, which when executed by a processor, causes the apparatus to perform:

decoding, from a bitstream, an indication of a number of bits in prediction type mask syntax elements;

decoding, from the bitstream, a first prediction type mask syntax element for a first set of prediction types that is applicable from a base layer to a second enhancement layer, wherein the first set of prediction types is a subset of all prediction types available for prediction between layers, and wherein a prediction type of the first set is represented by a distinctive bit number in a bit mask;

decoding, from the bitstream, a second prediction type mask syntax element for a second set of prediction types that is applicable from a first enhancement layer to the second enhancement layer, wherein the second set of prediction types is a subset of all prediction types available for prediction between layers, and wherein a prediction type of the second set is represented by a distinctive bit number in a bit mask; and

decoding the second enhancement layer using said first set of prediction types from the base layer and said second set of prediction types from the first enhancement layer,

wherein said prediction types of the first set that are available for prediction between layers are at least one of the following: sample prediction, motion information prediction, sample adaptive offset parameter prediction or intra mode information prediction.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2023
From: HANNUKSELA, MISKA MATIAS; UGUR, KEMAL
To: NOKIA CORPORATION
Reel/Frame 065312/0644 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2023
From: NOKIA CORPORATION
To: NOKIA TECHNOLOGIES OY
Reel/Frame 065321/0057 →
Continuity (6)
Continuation 17504092 · Oct 18, 2021
Continuation 16689582 · Nov 20, 2019
Continuation 15899129 · Feb 19, 2018
Continuation 14143986 · Dec 30, 2013
Provisional Application 61748938 · Jan 4, 2013
Related Publication 20240056595A1 · Feb 15, 2024
References Cited (55)
US 9900609B2 · Hannuksela et al. · 2018 [cited by applicant]
US 10506247B2 · Hannuksela et al. · 2019 [cited by applicant]
US 20070014346A1 · Wang · 2007 [cited by examiner]
US 20070086521A1 · Wang · 2007 [cited by examiner]
US 20080089411A1 · Wenger · 2008 [cited by examiner]
US 20090003389A1 · Joung et al. · 2009 [cited by applicant]
US 20100202540A1 · Fang · 2010 [cited by applicant]
US 20120044322A1 · Tian et al. · 2012 [cited by applicant]
US 20120056981A1 · Tian et al. · 2012 [cited by applicant]
US 20120230431A1 · Boyce et al. · 2012 [cited by applicant]
US 20120243606A1 · Lainema et al. · 2012 [cited by applicant]
US 20130279576A1 · Chen · 2013 [cited by examiner]
US 20130287093A1 · Hannuksela et al. · 2013 [cited by applicant]
CN 101420609 · 2009 [cited by applicant]
KR 1020080027338A · 2008 [cited by applicant]
KR 1020120024578A · 2012 [cited by applicant]
WO WO20120167712 · 2012 [cited by applicant]
WO WO20121167711A1 · 2012 [cited by applicant]
WO WO20131160559A1 · 2013 [cited by applicant]
“Parameter Values for the HDTV Standards for Production and International Programme Exchange”, Recommendation ITU-R BT.709-5, BT Series Broadcasting service (television), Apr. 2002, 32 pages. [cited by applicant]
“Parameter Values for Ultra-High Definition Television Systems for Production and International Programme Exchange”, Recommendation ITU-R BT.220, BT Series, Broadcasting service (television), Aug. 2012, 7 pages. [cited by applicant]
Advisory Action for U.S. Appl. No. 14/143,986 dated Feb. 13, 2017. [cited by applicant]
Boyce et al., “High Level Syntax Hooks for Future Extensions”, Joint Collaborative Team on Video Coding (JCT-VC) ITU-T SG16 WP3 and ISO/IEC JTC1/SC29NVG11, 8th Meeting, Feb. 1-10, 2012, pp. 1-6. [cited by applicant]
Boyce et al., “NAL Unit Header and Parameter Set Designs for HEVC Extensions”, Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29NVG11, 11th Meeting, Oct. 10-19, 2012, pp. 1-8. [cited by applicant]
Choi, B. et al. “MV-HEVC/SHVC HLS: On interlayer prediction type.” Samsung Electronics Co., Ltd., Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11; 5th Meeting: Vienna, AT, … [cited by applicant]
Decision to Grant for Chinese Application No. 201380074258X dated Sep. 19, 2019, 3 pages. [cited by applicant]
Extended European Search Report for corresponding European Patent Application No. 13870207.1 dated Jun. 15, 2016; 8 pages. [cited by applicant]
Final Office Action for U.S. Appl. No. 16/689,582 dated Jan. 14, 2021, 18 pages. [cited by applicant]
H. Schwarz, et al.; “Overview of the Scalable Video coding extension of the H.264/AVC Standard”; IEEE Transactions on Circuits and Systems for Video Technology; vol. 17, No. 9; Sep. 2007; pp. 1103-1120. [cited by applicant]
Intention to Grant for European Application No. 13 870 207.1 dated Dec. 17, 2019, 10 pages. [cited by applicant]
International Search Report and Written Opinion received for corresponding Patent Cooperation Treaty Application No. PCT/FI2013/051216, dated Apr. 9, 2014, 13 pages. [cited by applicant]
Minutes of the Oral Proceedings for European Application No. 13 870 207.1 dated Dec. 10, 2019, 19 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 17/504,092 dated Sep. 27, 2022. [cited by applicant]
Notice of Allowance for Korean Application No. 10-2015-7020987 dated Mar. 30, 2018, 3 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 14/143,986 dated Oct. 5, 2017. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/899,129 dated Apr. 9, 2019. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/899,129 dated Jul. 29, 2019. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/689,582 dated Jun. 14, 2021. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/504,092 dated Jun. 14, 2023. [cited by applicant]
Office Action for Chinese Application No. 201380074258X dated Jan. 2, 2019, 8 pages. [cited by applicant]
Office Action for Chinese Application No. 201380074258X dated May 8, 2019, 6 pages. [cited by applicant]
Office Action for U.S. Appl. No. 14/143,986 dated Jan. 29, 2016, 13 pages. [cited by applicant]
Office Action for U.S. Appl. No. 14/143,986 dated Oct. 6, 2016, 13 pages. [cited by applicant]
Office Action for U.S. Appl. No. 15/899,129 dated Jul. 27, 2018. [cited by applicant]
Office Action for U.S. Appl. No. 16/689,582 dated Jun. 19, 2020. [cited by applicant]
Office Action from corresponding Chinese Application No. 201380074258.X dated Oct. 20, 2017, with English Translation, 13 pages. [cited by applicant]
Office Action from corresponding European Application No. 13870207.1 dated Sep. 4, 2017, 5 pages. [cited by applicant]
Office Action from corresponding Korean Patent Application No. 2015-7020987 dated Nov. 28, 2016. [cited by applicant]
Office Action from corresponding Korean Patent Application No. 2015-7020987 dated Nov. 28, 2017 with English Summary, 6 pages. [cited by applicant]
Result of Consultation for European Application No. 13 870 207.1 dated Nov. 13, 2019, 2 pages. [cited by applicant]
Search Report and Written Opinion for corresponding Singapore Application No. 11201505278T, dated Jul. 12, 2016, 10 pages. [cited by applicant]
Summons to Attend Oral Proceedings for European Application No. 13870207.1 dated Apr. 25, 2019, 7 pages. [cited by applicant]
U.S. Appl. No. 61/449,079, “Depth Map Coding”, filed Mar. 3, 2011, 63 pages. [cited by applicant]
U.S. Appl. No. 61/706,727, “Method and Technical Equipment for Scalable Video Coding”, filed Sep. 27, 2012, 63 pages. [cited by applicant]
Yamamoto, T. et al. “Non-Square Partition Mode Grouping for CAVLC.” SHARP Corporation, Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11; 7th Meeting: Geneva, CH, Nov. 21-30,… [cited by applicant]