IP Library Granted Patent US 12,284,325
Granted Patent B2
US 12,284,325 · App. 17/892,911 · Granted Apr 22, 2025

Adaptive frame packing for 360-degree video coding

Inventors: Philippe Hanhart (La Conversion, CH); Yuwen He (San Diego, CA); Yan Ye (San Diego, CA)
Assignee: InterDigital VC Holdings, Inc.
H04N13/161H04N19/172H04N19/186H04N19/593H04N19/597
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,284,325
App. No.
17/892,911
Granted
Apr 22, 2025
Kind
B2
Abstract

A video coding device may be configured to periodically select the frame packing configuration (e.g., face layout and/or face rotations parameters) associated with a RAS. The device may receive a plurality of pictures, which may each comprise a plurality of faces. The pictures may be grouped into a plurality of RASs. The device may select a frame packing configuration with the lowest cost for a first RAS. For example, the cost of a frame packing configuration may be determined based on the first picture of the first RAS. The device may select a frame packing configuration for a second RAS. The frame packing configuration for the first RAS may be different than the frame packing configuration for the second RAS. The frame packing configuration for the first RAS and the frame packing configuration for the second RAS may be signaled in the video bitstream.

Claims (62)

1. A device for video decoding, the device comprising:

a processor configured to:

determine a chroma sample location type associated with a picture, wherein the picture comprises a plurality of faces, and luma samples are in a face of the plurality of faces;

based on a frame packing configuration, determine a face rotation associated with the face;

based on the chroma sample location type and the face rotation associated with the face, determine a filter for down-sampling the luma samples associated with the face;

down-sample the luma samples based on the filter; and

predict chroma samples based on the down-sampled luma samples.

2. The device of claim 1 , wherein the face is a first face, the filter is a first filter, and the processor is further configured to:

based on the frame packing configuration, determine a face rotation associated with a second face;

determine a second filter for down-sampling luma samples associated with the second face based on the chroma sample location type and the face rotation associated with the second face;

down-sample the luma samples in the second face based on the second filter; and

predict chroma samples in the second face based on the down-sampled luma samples in the second face.

3. The device of claim 1 , wherein a cross-component linear model (CCLM) is used to predict the chroma samples based on the down-sampled luma samples.

4. The device of claim 1 , wherein the processor is further configured to:

obtain, from video data, a chroma sample location type presence indication configured to indicate whether a chroma sample location type indication is present in the video data; and

based on a condition that the chroma sample location type indication is present, obtain a value of the chroma sample location type indication, wherein the chroma sample location type is determined based on the value of the chroma sample location type indication.

5. The device of claim 4 , wherein the processor is further configured to:

obtain, from the video data, an indication of a chroma format; and

based on the indication of the chroma format, determine that the chroma format is equal to 4:2:0, wherein the value of the chroma sample location type indication is obtained based on the chroma format equaling 4:2:0.

6. The device of claim 1 , wherein the processor is further configured to:

based on the chroma sample location type, determine a first filter candidate and a second filter candidate for down-sampling the luma samples, wherein the filter for down-sampling luma samples is determined by selecting from the first filter candidate and the second filter candidate.

7. A method for video decoding, comprising:

determining a chroma sample location type associated with a picture, wherein the picture comprises a plurality of faces, and luma samples are in a face of the plurality of faces;

based on a frame packing configuration, determining a face rotation associated with the face;

based on the chroma sample location type and the face rotation associated with the face, determining a filter for down-sampling the luma samples associated with the face;

down-sampling the luma samples based on the filter; and

predicting chroma samples based on the down-sampled luma samples.

8. The method of claim 7 , wherein the face is a first face, the filter is a first filter, further comprising:

based on the frame packing configuration, determining a face rotation associated with a second face;

determining a second filter for down-sampling luma samples associated with the second face based on the chroma sample location type and the face rotation associated with the second face;

down-sampling the luma samples in the second face based on the second filter; and

predicting chroma samples in the second face based on the down-sampled luma samples in the second face.

9. The method of claim 7 , wherein a cross-component linear model (CCLM) is used to predict the chroma samples based on the down-sampled luma samples.

10. The method of claim 7 , further comprising:

obtaining, from video data, a chroma sample location type presence indication configured to indicate whether a chroma sample location type indication is present in the video data; and

based on a condition that the chroma sample location type indication is present, obtaining a value of the chroma sample location type indication, wherein the chroma sample location type is determined based on the value of the chroma sample location type indication.

11. The method of claim 10 , further comprising:

obtaining, from the video data, an indication of a chroma format; and

based on the indication of the chroma format, determining that the chroma format is equal to 4:2:0, wherein the value of the chroma sample location type indication is obtained based on the chroma format equaling 4:2:0.

12. The method of claim 7 , further comprising:

based on the chroma sample location type, determining a first filter candidate and a second filter candidate for down-sampling the luma samples;

wherein the filter for down-sampling luma samples is determined by selecting from the first filter candidate and the second filter candidate.

13. A device for video encoding, the device comprising:

a processor configured to:

determine a chroma sample location type associated with a picture, wherein the picture comprises a plurality of faces, and luma samples are in a face of the plurality of faces;

based on a frame packing configuration, determine a face rotation associated with the face;

based on the chroma sample location type and the face rotation associated with the face, determine a filter for down-sampling the luma samples associated with the face;

down-sample the luma samples based on the filter; and

encode chroma samples based on the down-sampled luma samples.

14. The device of claim 13 , wherein the face is a first face, the filter is a first filter, and the processor is further configured to:

based on the frame packing configuration, determine a face rotation associated with a second face;

determine a second filter for down-sampling luma samples associated with the second face based on the chroma sample location type and the face rotation associated with the second face;

down-sample the luma samples in the second face based on the second filter; and

encode chroma samples in the second face based on the down-sampled luma samples in the second face.

15. The device of claim 13 , wherein a cross-component linear model (CCLM) is used to predict the chroma samples based on the down-sampled luma samples.

16. The device of claim 13 , wherein the processor is further configured to:

determine whether to include a chroma sample location type indication in video data;

include a chroma sample location type presence indication in the video data configured to indicate whether the chroma sample location type indication is present in the video data; and

based on a determination to include the chroma sample location type indication, include a value of the chroma sample location type indication in the video data.

17. The device of claim 16 , wherein the processor is further configured to:

determine a chroma format is equal to 4:2:0; and

based on the chroma format being equal to 4:2:0, determine to include the chroma sample location type indication in the video data.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 5, 2024
From: HANHART, PHILIPPE; HE, YUWEN; YE, YAN
To: VID SCALE, INC.
Reel/Frame 069135/0452 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2024
From: VID SCALE, INC.
To: INTERDIGITAL VC HOLDINGS, INC.
Reel/Frame 068284/0031 →
Continuity (4)
Continuation 16960948
Provisional Application 62733371 · Sep 19, 2018
Provisional Application 62617939 · Jan 16, 2018
Related Publication 20230075126A1 · Mar 9, 2023
References Cited (70)
US 11025952B2 · Lv et al. · 2021 [cited by applicant]
US 20110280311A1 · Chen et al. · 2011 [cited by applicant]
US 20120195378A1 · Zheng et al. · 2012 [cited by applicant]
US 20140092998A1 · Zhu · 2014 [cited by examiner]
US 20140112394A1 · Sullivan · 2014 [cited by examiner]
US 20140211842A1 · Zhao et al. · 2014 [cited by applicant]
US 20150195573A1 · Aflaki Beni · 2015 [cited by examiner]
US 20150222928A1 · Tian et al. · 2015 [cited by applicant]
US 20160212433A1 · Zhu · 2016 [cited by examiner]
US 20170150186A1 · Zhang · 2017 [cited by examiner]
US 20170372494A1 · Zhu et al. · 2017 [cited by applicant]
US 20170374385A1 · Huang et al. · 2017 [cited by applicant]
US 20180077426A1 · Zhang · 2018 [cited by examiner]
US 20180164593A1 · Van Der Auwera et al. · 2018 [cited by applicant]
US 20180205934A1 · Abbas et al. · 2018 [cited by applicant]
US 20190007669A1 · Kim et al. · 2019 [cited by applicant]
US 20190089981A1 · Lv et al. · 2019 [cited by applicant]
US 20200236370A1 · Galpin et al. · 2020 [cited by applicant]
US 20210250619A1 · Ye et al. · 2021 [cited by applicant]
CN 102365869A · 2012 [cited by applicant]
CN 103348677A · 2013 [cited by applicant]
CN 103841391A · 2014 [cited by applicant]
CN 104429071A · 2015 [cited by applicant]
CN 105379268A · 2016 [cited by applicant]
CN 107396138A · 2017 [cited by applicant]
WO 2017020021A1 · 2017 [cited by applicant]
WO 2017140946A1 · 2017 [cited by applicant]
WO 2018009746A1 · 2018 [cited by applicant]
WO 2019143551A1 · 2019 [cited by applicant]
Hanhart , InterDigital's Response to the 360° Video Category in Joint Call for Evidence on Video Compression with Capability beyond HEVC, Jul. 13-21, 2017 (Year: 2017). [cited by examiner]
360LIB, Available at <https://jvet.hhi.fraunhofer.de/svn/svn_360Lib/>, 1 page. [cited by applicant]
Abbas et al., “AHG8: New GoPro Test Sequences for Virtual Reality Video Coding”, JVET-D0026, GoPro, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 4th Meeting: Chengdu, CN, Oct. 1… [cited by applicant]
Abbas et al., “AHG8: New Test Sequences for Spherical Video Coding from GoPro”, JVET-G0147, GoPro, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 7th Meeting: Torino, IT, Jul. 13-… [cited by applicant]
Asbun et al., “AHG8: InterDigital Test Sequences for Virtual Reality Video Coding”, JVET-D0039, InterDigital Communications, Inc., Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 4… [cited by applicant]
Asbun et al., “InterDigital Test Sequences for Virtual Reality Video Coding”, JVET-G0055, InterDigital Communications, Inc., Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 7th Mee… [cited by applicant]
Baroncini et al., “Results of the Joint Call for Evidence on Video Compression with Capability beyond Hevc”, JVET-G1004-V2, Jvet, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 7t… [cited by applicant]
Boyce et al., “AHG8: Spherical Rotation Orientation SEI for Coding of 360 Video”, JVET-E0075_V3, Intel, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 5th Meeting: Geneva, CH, Jan… [cited by applicant]
Boyce et al., “EE4: Padded ERP (PERP) Projection Format”, JVET-G0098, Intel Corp., Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 7th Meeting: Torino, IT, Jul. 13-21, 2017, pp. 1-… [cited by applicant]
Boyce et al., “JVET Common Test Conditions and Evaluation Procedures for 360° Video”, JVET-F1030-V4, Editors, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 6th Meeting: Hobart, A… [cited by applicant]
Boyce et al., “Supplemental Enhancement Information for Coded Video Bitstreams (Draft 3)”, JVET-Q2007-V6, Editors, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 17th Meeting: Brussel… [cited by applicant]
Chen et al., “Algorithm Description of Joint Exploration Test Model 7 (JEM 7)”, JVET-G1001-V1, Editors, Joint Video Exploration Team (JVET) of ITU- T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 7th Meeting: Torino, IT, Ju… [cited by applicant]
Chen et al., “Algorithm Description for Versatile Video Coding and Test Model 2 (VTM 2)”, JVET-K1002-V2, Editors, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG11, 11th Meeting: Ljubljana… [cited by applicant]
Choi, Byeongdoo, “Technologies under Consideration for Omnidirectional Media Application Format”, Systems Subgroup, ISO/IEC JTC1/SC29/WG11 N15946, San Diego, CA, US, Feb. 2016, 16 pages. [cited by applicant]
Coban et al., “AHG8: Adjusted Cubemap Projection for 360-Degree Video”, JVET-F0025, Qualcomm Inc., Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 6th Meeting: Hobart, AU, Mar. 31-… [cited by applicant]
Facebook360, “Facebook 360 Video”, Available at <https://facebook360.fb.com/>, pp. 1-5. [cited by applicant]
Github, “Facebook's Equirectangular to Cube Map Tool on GitHub”, Transform 360, Available at <https://github.com/facebook/transform?files=1>, pp. 1-3. [cited by applicant]
Google, “Bringing Pixels Front and Center in VR Video”, Available at <https://www.blog.google/products/google-vr/bringing-pixels-front-and-center-vr-video/>, Mar. 14, 2017, pp. 1-8. [cited by applicant]
Google VR, “Google Cardboard”, Available at <https://www.google.com/get/cardboard/>, pp. 1-4. [cited by applicant]
Hanhart et al., “InterDigital's Response to the 360° Video Category in Joint Call for Evidence on Video Compression with Capability beyond HEVC”, JVET-G0024, InterDigital Communications Inc., Joint Video Exploration Tea… [cited by applicant]
He et al., “AHG8: Algorithm Description of Projection Format Conversion in 360Lib”, JVET-E0084, InterDigital Communications Inc., Samsung Electronics Co. Ltd., MediaTek Inc., Zhejiang University, Qualcomm Inc., OwlReali… [cited by applicant]
HTC, “HTC Vive”, Available at <https://www.htcvive.com/us/>, pp. 1-3. [cited by applicant]
ISO/IEC, “Requirements for OMAF”, Requirements, ISO/IEC JTC1/SC29/WG11 N16143, San Diego, CA, US, Feb. 2016, 2 pages. [cited by applicant]
Kuzyakov et al., “Next-Generation Video Encoding Techniques for 360 Video and VR”, Facebook Code, Available at <https://code.facebook.com/posts/1126354007399553/next-generation-video-encoding-techniques-for-360-video-an… [cited by applicant]
Ma et al., “Simplification of the Common Test Condition for Fast Simulation”, JVET-B0036, Huawei Co., Ltd., Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 2nd Meeting: San Diego, … [cited by applicant]
Norkin et al., “Call for Test Materials for Future Video Coding Standardization”, JVET-B1002, ITU-T Q6/16 Visual Coding (VCEG) and ISO/IEC JTC1/SC29/WG11 Coding of Moving Pictures and Audio (MPEG), Joint Video Explorati… [cited by applicant]
Oculus, “Oculus Rift”, Available at <https://www.oculus.com/en-us/rift/>, pp. 1-19. [cited by applicant]
Schwarz et al., “Tampere Pole Vaulting Sequence for Virtual Reality Video Coding”, JVET-D0143, Nokia, Tampere University of Technology, Rakka Creative, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC… [cited by applicant]
Segall et al., “Draft Joint Call for Proposals on Video Compression with Capability beyond HEVC”, JVET-G1002, Editors, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 7th Meeting: … [cited by applicant]
Sullivan et al., “Meeting Notes of the 3rd Meeting of the Joint Video Exploration Team (JVET)”, JVET-C1000, Responsible Coordinators, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11… [cited by applicant]
Sullivan et al., “Meeting Report of the 7th meeting of the Joint Video Exploration Team (JVET), Torino, IT”, JVET-G_Notes_d2, Responsible Coordinators, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC… [cited by applicant]
Sun et al., “Test Sequences for Virtual Reality Video Coding from LetinVR”, JVET-G0053, Letin VR Digital Technology Co., Ltd., Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 7th M… [cited by applicant]
Sun et al., “Test Sequences for Virtual Reality Video Coding from LetinVR”, JVET-D0179, Letin VR Digital Technology Co., Ltd., Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 4th M… [cited by applicant]
Thomas et al., “5G and Future Media Consumption”, TNO, ISO/IEC JTC1/SC29/WG11 MPEG2016/m37604, San Diego, CA, US, Feb. 2016, 10 pages. [cited by applicant]
Wien et al., “Joint Call for Evidence on Video Compression with Capability Beyond HEVC”, JVET-F1002, JVET, Joint Video Exploration Team (JVET), of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Hobart, AU, Mar. 31-Apr.… [cited by applicant]
Xiu et al., “Description of SDR, HDR and 360° Video Coding Technology Proposal by InterDigital Communications and Dolby Laboratories”, JVET-J0015-V1, InterDigital Communications, Inc., Dolby Laboratories, Inc., Joint Vi… [cited by applicant]
Ye et al., “Algorithm Descriptions of Projection Format Conversion and Video Quality Metrics in 360Lib”, JVET-F1003-V1, Editors, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 6th… [cited by applicant]
Youtube, “360 Video”, Virtual Reality, Available at <https://www.youtube.com/channel/UCzuqhhs6NWbgTzMuM09WKDQ>, pp. 1-3. [cited by applicant]
CN 103841391 (A), Cited in Notice of Allowance in related Chinese Application No. 201980008336.3 on Oct. 10, 2023. [cited by applicant]
Hanhart et al., “CE13: Adaptive frame packing (Tests 4.3 and 4.4)”, JVET-K0331, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 11th Meeting: Ljubljana, SI, Jul. 10-18, 2018, pp. 1… [cited by applicant]
Hanhart et al., “CE13-related: Adaptive frame packing on top of CMP, MCP, and PAU”, JVET-K0332-r1, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 11th Meeting: Ljubljana, SI, Jul.… [cited by applicant]