IP Library Granted Patent US 12,513,329
Granted Patent B2
US 12,513,329 · App. 17/860,629 · Granted Dec 30, 2025

360-degree video coding using geometry projection

Inventors: Yuwen He (San Diego, CA); Yan Ye (San Diego, CA); Philippe Hanhart (La Conversion, CH); Xiaoyu Xiu (San Diego, CA)
Assignee: InterDigital VC Holdings, Inc.
H04N19/597G06T17/10G06T17/30H04N13/117H04N13/161H04N13/194H04N13/344H04N13/383H04N19/105H04N19/132H04N19/172H04N19/563H04N19/593H04N23/698
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,513,329
App. No.
17/860,629
Granted
Dec 30, 2025
Kind
B2
Abstract

Processing video data may include capturing the video data with multiple cameras and stitching the video data together to obtain a 360-degree video. A frame-packed picture may be provided based on the captured and stitched video data. A current sample location may be identified in the frame-packed picture. Whether a neighboring sample location is located outside of a content boundary of the frame-packed picture may be determined. When the neighboring sample location is located outside of the content boundary, a padding sample location may be derived based on at least one circular characteristic of the 360-degree video content and the projection geometry. The 360-degree video content may be processed based on the padding sample location.

Claims (61)

1 . A method of video decoding comprising:

obtaining a current block location in a video content, wherein the current block location is associated with a current block of the video content, the video content comprising a plurality of faces;

obtaining a spatial neighboring block location, wherein the current block is configured to be decoded based on the spatial neighboring block location;

determining that the spatial neighboring block location is located outside of a content boundary of the video content;

based on the determination that the spatial neighboring block location is located outside of the content boundary of the video content, obtaining a mapped block location associated with the spatial neighboring block location including:

projecting the spatial neighboring block location along one direction being a perpendicular direction or a diagonal direction to a diagonal line extending out of a corner of a face of the plurality of faces to obtain an intermediate location; and

projecting the intermediate location along the one direction to a second content boundary to obtain the mapped block location associated with the spatial neighboring block location;

obtaining a prediction mode attribute associated with the mapped block location, wherein the prediction mode attribute includes an intra mode of a block at the mapped block location; and

based on the prediction mode attribute, decoding the current block.

2 . The method of claim 1 , wherein the method further comprises:

reconstructing the current block based on the intra mode of the block at the mapped block location.

3 . The method of claim 1 , wherein the video content comprises a plurality of faces, and obtaining the mapped block location comprises:

calculating a 3D position associated with the spatial neighboring block location, wherein the spatial neighboring block location is associated with a first face;

based on the calculated 3D position associated with the spatial neighboring block location, identifying a second face that is associated with the mapped block location; and

applying a geometry projection with the 3D position of the spatial neighboring block location to derive a 2D planar position of the spatial neighboring block location in the second face as the mapped block location.

4 . The method of claim 3 , wherein the video content comprises a plurality of faces associated with a first geometry projection, and wherein the method further comprises:

converting a coordinate associated with the first geometry projection into an intermediate coordinate, wherein the intermediate coordinate is associated with a second geometry projection, wherein the 3D position of the spatial neighboring block location is obtained in the intermediate coordinate, and the 2D planar position of the spatial neighboring block location in the second face is obtained in the intermediate coordinate; and

converting the derived 2D planar position of spatial neighboring block location associated with the second geometry projection back to the coordinate associated with the first geometry projection.

5 . The method of claim 1 , wherein the video content comprises a plurality of faces, wherein the spatial neighboring block location is located in a first face, wherein the mapped block location is located in a second face, and wherein the second face differs from the first face.

6 . A non-transitory computer readable medium storing instructions, which, when executed by one or more processors, cause the one or more processors to perform the method of claim 1 .

7 . An apparatus for video decoding comprising:

a processor configured to:

obtain a current block location in a video content, wherein the current block location is associated with a current block of the video content, the video content comprising a plurality of faces;

obtain a spatial neighboring block location, wherein the current block is configured to be decoded based on the spatial neighboring block location;

determine that the spatial neighboring block location is located outside of a content boundary of the video content;

based on the determination that the spatial neighboring block location is located outside of the content boundary of the video content, obtain a mapped block location associated with the spatial neighboring block location including:

project the spatial neighboring block location along one direction being a perpendicular direction or a diagonal direction to a diagonal line extending out of a corner of a face of the plurality of faces to obtain an intermediate location; and

project the intermediate location along the one direction to a second content boundary to obtain the mapped block location associated with the spatial neighboring block location;

obtain a prediction mode attribute associated with the mapped block location, wherein the prediction mode attribute includes an intra mode of a block at the mapped block location; and

based on the prediction mode attribute, decode the current block.

8 . The apparatus of claim 7 , wherein the processor is further configured to:

reconstruct the current block based on the intra mode of the block at the mapped block location.

9 . The apparatus of claim 7 , wherein the video content comprises a plurality of faces and to obtain the mapped block location comprises the processor being configured to:

calculate a 3D position associated with the spatial neighboring block location, wherein the spatial neighboring block location is associated with a first face;

based on the calculated 3D position associated with the spatial neighboring block location, identify a second face that is associated with the mapped block location; and

apply a geometry projection with the 3D position of the spatial neighboring block location to derive a 2D planar position of the spatial neighboring block location in the second face as the mapped block location.

10 . The apparatus of claim 9 , wherein the video content comprises a plurality of faces associated with a first geometry projection, and wherein the processor is further configured to:

convert a coordinate associated with the first geometry projection into an intermediate coordinate, wherein the intermediate coordinate is associated with a second geometry projection, wherein the 3D position of the spatial neighboring block location is obtained in the intermediate coordinate, and the 2D planar position of the spatial neighboring block location in the second face is obtained in the intermediate coordinate; and

convert the derived 2D planar position of the spatial neighboring block location associated with the second geometry projection back to the coordinate associated with the first geometry projection.

11 . A method of video encoding comprising:

obtaining a current block location in a video content, wherein the current block location is associated with a current block of the video content, the video content comprising a plurality of faces;

obtaining a spatial neighboring block location, wherein the current block is configured to be encoded based on the spatial neighboring block location;

determining that the spatial neighboring block location is located outside of a content boundary of the video content;

based on the determination that the spatial neighboring block location is located outside of the content boundary of the video content, obtaining a mapped block location associated with the spatial neighboring block location including:

projecting the spatial neighboring block location along one direction being a perpendicular direction or a diagonal direction to a diagonal line extending out of a corner of a face of the plurality of faces to obtain an intermediate location; and

projecting the intermediate location along the one direction to a second content boundary to obtain the mapped block location associated with the spatial neighboring block location;

obtaining a prediction mode attribute associated with the mapped block location, wherein the prediction mode attribute includes an intra mode of a block at the mapped block location; and

encoding the current block based on the obtained prediction mode attribute associated with the mapped block location.

12 . A non-transitory computer readable medium storing instructions, which, when executed by one or more processors, cause the one or more processors to perform the method of claim 11 .

13 . The apparatus of claim 7 , wherein the video content comprises a plurality of faces, wherein the spatial neighboring block location is located in a first face, wherein the mapped block location is located in a second face, and wherein the second face differs from the first face, wherein the content boundary comprises at least one of a frame packed picture boundary or a face boundary, and wherein the video content comprises a 360-degree video content, and the mapped block location is obtained based on a circular characteristic of the 360-degree video content.

14 . A apparatus for video encoding comprising:

a processor configured to:

obtain a current block location in a video content, wherein the current block location is associated with a current block of the video content, the video content comprising a plurality of faces;

obtain a spatial neighboring block location, wherein the current block is configured to be encoded based on the spatial neighboring block location;

determine that the spatial neighboring block location is located outside of a content boundary of the video content;

based on the determination that the spatial neighboring block location is located outside of the content boundary of the video content, obtain a mapped block location associated with the spatial neighboring block location including:

project the spatial neighboring block location along one direction being a perpendicular direction or a diagonal direction to a diagonal line extending out of a corner of a face of the plurality of faces to obtain an intermediate location; and

project the intermediate location along the one direction to a second content boundary to obtain the mapped block location associated with the spatial neighboring block location;

obtain a prediction mode attribute associated with the mapped block location, wherein the prediction mode attribute includes an intra mode of a block at the mapped block location; and

encode the current block based on the obtained prediction mode attribute associated with the mapped block location.

15 . The apparatus of claim 14 , wherein the video content comprises a plurality of faces.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2024
From: HE, YUWEN; YE, YAN; HANHART, PHILIPPE; XIU, XIAOYU
To: VID SCALE, INC.
Reel/Frame 069086/0106 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2024
From: VID SCALE, INC.
To: INTERDIGITAL VC HOLDINGS, INC.
Reel/Frame 068284/0031 →
Continuity (7)
Continuation 17111683 · Dec 4, 2020
Continuation 16315447
Provisional Application 62500605 · May 3, 2017
Provisional Application 62463242 · Feb 24, 2017
Provisional Application 62404017 · Oct 4, 2016
Provisional Application 62360112 · Jul 8, 2016
Related Publication 20220368947A1 · Nov 17, 2022
References Cited (59)
US 9578329B2 · Yang et al. · 2017 [cited by applicant]
US 10863182B2 · Hannuksela · 2020 [cited by applicant]
US 10887621B2 · He et al. · 2021 [cited by applicant]
US 20040247173A1 · Nielsen et al. · 2004 [cited by applicant]
US 20060034374A1 · Park et al. · 2006 [cited by applicant]
US 20060034529A1 · Park et al. · 2006 [cited by applicant]
US 20060034530A1 · Park · 2006 [cited by applicant]
US 20130335532A1 · Tanaka et al. · 2013 [cited by applicant]
US 20140133758A1 · Kienzle · 2014 [cited by applicant]
US 20140219356A1 · Nishitani et al. · 2014 [cited by applicant]
US 20150195573A1 · Aflaki Beni et al. · 2015 [cited by applicant]
US 20160012855A1 · Krishnan · 2016 [cited by applicant]
US 20160112704A1 · Grange et al. · 2016 [cited by applicant]
US 20170085917A1 · Hannuksela · 2017 [cited by applicant]
US 20170214937A1 · Lin · 2017 [cited by examiner]
US 20170230668A1 · Lin · 2017 [cited by examiner]
US 20170332107A1 · Abbas · 2017 [cited by examiner]
US 20190007669A1 · Kim · 2019 [cited by examiner]
CN 1491403A · 2004 [cited by applicant]
CN 101853552A · 2010 [cited by applicant]
CN 103443582A · 2013 [cited by applicant]
CN 103782595A · 2014 [cited by applicant]
EP 1064817A1 · 2001 [cited by applicant]
EP 1162830A2 · 2001 [cited by applicant]
EP 1441307A1 · 2004 [cited by applicant]
JP 2013016930A · 2013 [cited by applicant]
KR 20060015225A · 2006 [cited by applicant]
KR 1020060050350A · 2006 [cited by applicant]
KR 1020150129548A · 2015 [cited by applicant]
WO 0008889A1 · 2000 [cited by applicant]
WO 2017205648A1 · 2017 [cited by applicant]
“VR Coaster”, Available at <http://www.vrcoaster.com/>, 2014-2018, pp. 1-7. [cited by applicant]
Abbas, Adeel, “GoPro Test Sequences for Virtual Reality Video Coding”, JVET-C0021, GoPro, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 3rd Meeting: Geneva, CH, May 26-31, 2016, … [cited by applicant]
Bang et al., “Description of 360 3D Video Application Exploration Experiments on Divergent Multi-View Video”, Requirements, ISO/IEC JTC1/SC29/WG11 MPEG2015/ M16129, San Diego, US, Feb. 2016, 5 pages. [cited by applicant]
Budagavi et al., “360 Degrees Video Coding using Region Adaptive Smoothing”, 2015 IEEE International Conference on Image Processing (ICIP), Quebec City, QC, Canada, Sep. 27-30, 2015, pp. 750-754. [cited by applicant]
Carbotte, Kevin, “Google Looks to Solve VR Video Quality Issues With Equi-Angular Cubemaps (EAC)”, Tom's Hardware, Available at <https://www.tomshardware.com/news/google-equi-angulra-cubemap-projection-technology,33917.… [cited by applicant]
Choi et al., “Test Sequence Formats for Virtual Reality Video Coding”, JVET-C0050, Samsung Electronics Co., Ltd., Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 3rd Meeting: Genev… [cited by applicant]
Choi, Byeongdoo, “Technologies under Consideration for Omnidirectional Media Application Format”, Systems Subgroup, ISO/IEC JTC1/SC29/WG11 N15946, San Diego, CA, US, Feb. 2016, 16 pages. [cited by applicant]
Coban et al., “AHG8: Adjusted Cubemap Projection for 360-Degree Video”, JVET-F0025, Qualcomm Inc., Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 6th Meeting: Hobart, AU, Mar. 31-… [cited by applicant]
Facebook360, “Facebook 360 Video”, Available at <https://facebook360.fb.com/>, pp. 1-5. [cited by applicant]
Github, “Facebook's Equirectangular to Cube Map Tool on GitHub”, Transform 360, Available at <https://github.com/facebook/transform?files=1>, pp. 1-3. [cited by applicant]
Google VR, “Google Cardboard”, Available at <https://www.google.com/get/cardboard/>, pp. 1-4. [cited by applicant]
Habe et al., “Report of EE1 on Omni-Directional Video”, Jeita 3DMM Committee, ISO/IEC JTC1/SC29/WG11 MPEG2003/M9480, Pattaya, Mar. 2003, pp. 1-13. [cited by applicant]
Hanhart et al., “AHG8: Reference Samples Derivation Using Geometry Padding for Intra Coding”, JVET-D0092, InterDigital Communications Inc., Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29… [cited by applicant]
He et al., “AHG8: Geometry Padding for 360 Video Coding”, JVET-D0075, InterDigital Communications Inc., Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC1/SC 29/WG 11, 4th Meeting: Chengdu, CN, Oct… [cited by applicant]
Ho et al., “Unicube for Dynamic Environment Mapping”, IEEE Transactions on Visualization and Computer Graphics, vol. 17, No. 1, Jan. 2011, pp. 51-63. [cited by applicant]
HTC, “HTC Vive”, Available at <https://www.htcvive.com/us/>, pp. 1-3. [cited by applicant]
ISO/IEC, “Requirements for OMAF”, Requirements, ISO/IEC JTC1/SC29/WG11 N16143, San Diego, CA, US, Feb. 2016, 2 pages. [cited by applicant]
Kuzyakov et al., “Next-Generation Video Encoding Techniques for 360 Video and VR”, Facebook Code, Available at <https://code.facebook.com/posts/1126354007399553/next-generation-video-encoding-techniques-for-360-video-an… [cited by applicant]
Licea-Kane et al., “ARB Seamless Cube Map”, The Khronos Group, Inc., Jul. 3, 2009, 3 pages. [cited by applicant]
Norkin et al., “Call for Test Materials for Future Video Coding Standardization”, JVET-B1002, ITU-T Q6/16 Visual Coding (VCEG) and ISO/IEC JTC1/SC29/WG11 Coding of Moving Pictures and Audio (MPEG), Joint Video Explorati… [cited by applicant]
Oculus, “Oculus Rift”, Available at <https://www.oculus.com/en-us/rift/>, pp. 1-19. [cited by applicant]
Ridge et al., “Nokia Test Sequences for Virtual Reality Video Coding”, JVET-C0064, Nokia Technologies, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 3rd Meeting: Geneva, CH, May … [cited by applicant]
Shih et al., “AHG8: Face-Based Padding Scheme for Cube Projection”, JVET-E0057, MediaTek Inc., Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC1/SC 29/WG 11, 5th Meeting: Geneva, CH, Jan. 12-20, 2… [cited by applicant]
Sullivan et al., “Meeting Notes of the 3rd Meeting of the Joint Video Exploration Team (JVET)”, JVET-C1000, Responsible Coordinators, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11… [cited by applicant]
Thomas et al., “5G and Future Media Consumption”, TNO, ISO/IEC JTC1/SC29/WG11 MPEG2016/m37604, San Diego, CA, US, Feb. 2016, 10 pages. [cited by applicant]
Youtube, “360 Video”, Virtual Reality, Available at <https://www.youtube.com/channel/UCzuqhhs6NWbgTzMuM09WKDQ>, pp. 1-3. [cited by applicant]
Yu et al., “A Framework to Evaluate Omnidirectional Video Coding Schemes”, IEEE International Symposium on Mixed and Augmented Reality, Sep. 29-Oct. 3, 2015, pp. 31-36. [cited by applicant]
Yu et al., “Content Adaptive Representations of Omnidirectional Videos for Cinematic Virtual Reality”, Proceedings of the 3rd International Workshop on Immersive Media Experiences, Brisbane, Australia, Oct. 30, 2015, pp… [cited by applicant]