IP Library › Granted Patent US 12,301,825
Granted Patent B2
US 12,301,825 · App. 18/097,165 · Granted May 13, 2025

Resolution-based video encoding

Inventors: Maneli Noorkami (Menlo Park, CA); Afshin Taghavi Nasrabadi (Santa Clara, CA); Ranjit Desai (Cupertino, CA)
Assignee: Apple Inc.
H04N19/136H04N19/107H04N19/119H04N19/159H04N19/172H04N19/176H04N19/30H04N19/50
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,301,825
App. No.
18/097,165
Granted
May 13, 2025
Kind
B2
Abstract

Aspects of the subject technology relate to encoding of video frames having content with a variable resolution that varies within the video frame. Aspects of the subject technology can provide an efficient encoding by using smaller macroblocks for lower resolution content within the video frame, and larger macroblocks for higher resolution content within the video frame. An encoder may be provided with resolution information for the content of the video frame, which can be used by the encoder to determine macroblock sizes, macroblock divisions, and/or prediction modes for the encoding of the video frame.

Claims (68)

1. A method, comprising:

obtaining a video frame having content with a variable resolution that varies within the video frame;

obtaining resolution information for the content of the video frame;

providing the video frame and the resolution information as inputs to an encoder; and

encoding the video frame with the encoder based on the resolution information, at least in part by determining a prediction mode for at least one macroblock for a first portion of the video frame based on the resolution information.

2. The method of claim 1 , wherein the resolution information comprises a resolution map for the content of the video frame.

3. The method of claim 1 , wherein the resolution information comprises a projection indicator for the content of the video frame.

4. The method of claim 3 , wherein encoding the video frame with the encoder based on the resolution information comprises determining, by the encoder, a plurality of resolution factors for the content of a respective plurality of portions of the video frame based on the projection indicator and based on previously stored projection information for the projection indicator.

5. The method of claim 1 , wherein encoding the video frame with the encoder based on the resolution information comprises:

obtaining a first resolution factor for the content of a first portion of the video frame based on the resolution information;

determining a first macroblock size of at least one macroblock for the first portion of the video frame based on the first resolution factor;

obtaining a second resolution factor for the content of a second portion of the video frame based on the resolution information; and

determining a second macroblock size, different from the first macroblock size, of at least one macroblock for the second portion of the video frame based on the second resolution factor.

6. The method of claim 5 , wherein encoding the video frame with the encoder based on the resolution information further comprises generating a resolution-based encoded video stream that includes the first resolution factor, an encoded version of the at least one macroblock for the first portion of the video frame, the second resolution factor, and an encoded version of the at least one macroblock for the second portion of the video frame.

7. The method of claim 5 , wherein encoding the video frame with the encoder based on the resolution information further comprises:

dividing the at least one macroblock for the first portion of the video frame based on the resolution information.

8. The method of claim 1 , wherein the video frame comprises a gaze-based foveated video frame, and wherein the determining the prediction mode comprises selecting an intra-frame prediction mode or an inter-frame prediction mode for the at least one macroblock for the first portion of the video frame.

9. The method of claim 1 , wherein the video frame comprises one of a foveated video frame, an equirectangular projection video frame, and a fish-eye projection video frame.

10. A method, comprising:

receiving, at an electronic device, a resolution-based encoded video stream comprising a plurality of encoded macroblocks for a video frame and a plurality of respective resolution factors for the plurality of encoded macroblocks, each of the plurality of encoded macroblocks having been encoded based at least in part on a prediction mode selected based on resolution information for a portion of the video frame; and

decoding the resolution-based encoded video stream, with a decoder of the electronic device using the plurality of respective resolution factors received in the resolution-based encoded video stream, to generate the video frame.

11. The method of claim 10 , wherein decoding the resolution-based encoded video stream comprises:

decoding a first one of the plurality of encoded macroblocks using a first respective one of the plurality of respective resolution factors to generate a first portion of the video frame; and

decoding a second one of the plurality of encoded macroblocks using a second respective one of the plurality of respective resolution factors to generate a second portion of the video frame.

12. The method of claim 11 , wherein the first portion of the video frame includes first content having a first resolution corresponding to the first respective one of the plurality of respective resolution factors, the second portion of the video frame has second content having a second resolution corresponding to the second one respective one of the plurality of respective resolution factors, and the first resolution is different from the second resolution.

13. The method of claim 12 , wherein the first resolution is higher than the second resolution, and wherein a size, after decoding, of the first one of the plurality of encoded macroblocks is larger than a size, after decoding, of the second one of the plurality of encoded macroblocks.

14. The method of claim 10 , wherein the video frame comprises one of a foveated video frame, an equirectangular projection video frame, and a fish-eye projection video frame.

15. An electronic device, comprising:

a memory; and

one or more processors configured to:

obtain a video frame having content with a variable resolution that varies within the video frame;

obtain resolution information for the video frame;

provide the video frame and the resolution information as inputs to an encoder; and

encode the video frame with the encoder based on the resolution information, at least in part by:

determining, based on the resolution information, an initial macroblock size, of a macroblock size for the at least one macroblock to obtain a final macroblock size for the at least one macroblock, and

including the at least one macroblock having the final macroblock size in an encoded video stream.

16. The electronic device of claim 15 , wherein the one or more processors are configured to encode the video frame with the encoder based on the resolution information, at least in part, by:

obtaining a first resolution factor for the content of a first portion of the video frame based on the resolution information;

determining a first macroblock size of at least one macroblock for the first portion of the video frame based on the first resolution factor;

obtaining a second resolution factor for the content of a second portion of the video frame based on the resolution information; and

determining a second macroblock size, different from the first macroblock size, of at least one macroblock for the second portion of the video frame based on the second resolution factor.

17. The electronic device of claim 16 , wherein the one or more processors are further configured to encode the video frame with the encoder based on the resolution information, at least in part, by generating a resolution-based encoded video stream that includes the first resolution factor, an encoded version of the at least one macroblock for the first portion of the video frame, the second resolution factor, and an encoded version of the at least one macroblock for the second portion of the video frame.

18. The electronic device of claim 16 , wherein the one or more processors are further configured to encode the video frame with the encoder based on the resolution information, at least in part, by:

dividing the at least one macroblock for the first portion of the video frame based on the resolution information.

19. The electronic device of claim 16 , wherein the one or more processors are further configured to encode the video frame with the encoder based on the resolution information, at least in part, by:

determining a prediction mode for the at least one macroblock for the first portion of the video frame based on the resolution information.

20. The electronic device of claim 15 , wherein the one or more processors are configured to encode the video frame with the encoder based on the resolution information by performing a power efficient optimization of a macroblock size, a macroblock division, and a prediction mode selection using the resolution information.

21. The electronic device of claim 15 , wherein the variable resolution of the content of the video frame is different from a variable resolution of content of a temporally adjacent video frame, and wherein the one or more processors are configured to encode the video frame with the encoder based on the resolution information by scaling at least one macroblock for the video frame using the resolution information, for comparison of the scaled macroblock with a reference block.

22. A method, comprising:

obtaining a video stream and a projection indicator for the video stream; and

generating, based on the projection indicator, an encoded video stream that includes:

a first resolution factor corresponding to an encoded version of at least one macroblock for a first portion of the video stream,

the at least one macroblock for the first portion of the video stream,

a second resolution factor corresponding to an encoded version of at least one macroblock for a second portion of the video stream, and

the at least one macroblock for the second portion of the video stream.

23. The method of claim 22 , wherein the first portion of the video stream and the second portion of the video stream:

are different portions of one video frame; or

correspond to a first and second video frame respectively.

24. The method of claim 23 , further comprising:

obtaining the first resolution factor for first content of the first portion of the video stream;

determining a first macroblock size of the at least one macroblock for the first portion of the video stream based on the first resolution factor;

obtaining the second resolution factor for second content of a second portion of the video stream; and

determining a second macroblock size, different from the first macroblock size, of the at least one macroblock for the second portion of the video stream based on the second resolution factor.

25. The method of claim 24 , wherein the one video frame includes content with a variable resolution that varies within the one video frame.

26. The method of claim 1 , wherein encoding the video frame with the encoder based on the resolution information comprises:

determining, based on the resolution information, an initial macroblock size for at least one macroblock for the video frame,

performing, based on the initial macroblock size, an optimization of a macroblock size for the at least one macroblock to obtain a final macroblock size for the at least one macroblock, and

including the at least one macroblock having the final macroblock size in an encoded video stream.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2023
From: NOORKAMI, MANELI; DESAI, RANJIT; TAGHAVI NASRABADI, AFSHIN
To: APPLE INC.
Reel/Frame 065314/0136 →
Continuity (2)
Provisional Application 63320552 · Mar 16, 2022
Related Publication 20230300338A1 · Sep 21, 2023
References Cited (38)
US 6252989B1 · Geisler · 2001 [cited by examiner]
US 8547435B2 · Mimar · 2013 [cited by examiner]
US 8619857B2 · Zhao · 2013 [cited by examiner]
US 9615063B2 · Tanaka · 2017 [cited by examiner]
US 10425643B2 · Shen · 2019 [cited by examiner]
US 10546364B2 · Bastani · 2020 [cited by examiner]
US 10798306B1 · Johansen · 2020 [cited by examiner]
US 10904508B2 · Zhou · 2021 [cited by applicant]
US 10997693B2 · Newman · 2021 [cited by examiner]
US 11792381B2 · Wang · 2023 [cited by examiner]
US 20050018911A1 · Deever · 2005 [cited by examiner]
US 20100086032A1 · Chen et al. · 2010 [cited by applicant]
US 20100215098A1 · Chung · 2010 [cited by applicant]
US 20110002391A1 · Uslubas et al. · 2011 [cited by applicant]
US 20120269274A1 · Kim et al. · 2012 [cited by applicant]
US 20130039415A1 · Kim et al. · 2013 [cited by applicant]
US 20140241421A1 · Orton-Jay · 2014 [cited by applicant]
US 20150063437A1 · Murakami et al. · 2015 [cited by applicant]
US 20160360209A1 · Gosling et al. · 2016 [cited by applicant]
US 20180174619A1 · Roy · 2018 [cited by examiner]
US 20180220119A1 · Horvitz et al. · 2018 [cited by applicant]
US 20180288363A1 · Amengual Galdon · 2018 [cited by examiner]
US 20180343472A1 · Wang · 2018 [cited by examiner]
US 20200288164A1 · Lin et al. · 2020 [cited by applicant]
US 20200366934A1 · Edpalm et al. · 2020 [cited by applicant]
US 20210136397A1 · Lashmikantha et al. · 2021 [cited by applicant]
US 20210142443A1 · Eble et al. · 2021 [cited by applicant]
US 20230350202A1 · Mayrand · 2023 [cited by examiner]
KR 101007381B1 · 2011 [cited by applicant]
Videoconferencing using spatially varying sensing with multiple moving foveae; Basu—1994; (Year: 1994). [cited by examiner]
Foveated video compression with optimal rate control; Lee—2001; (Year: 2001). [cited by examiner]
Foveated Neural Network Gaze Prediction on Egocentric Videos; Zhang—2017; (Year: 2017). [cited by examiner]
Visual pattern image sequence coding; Silsbee—1993; (Year: 1993). [cited by examiner]
Exploring with a foveated robot eye system; Baron—1994; (Year: 1994). [cited by examiner]
Videoconference using spatially varying sensing with multiple moving foveae; Basu—1994; (Year: 1994). [cited by examiner]
PoLin Lai, et al., “Signalling of VR spatial relationship in MPD,” International Organisation for Standardisation ISO/IEC JTC1/SC29/WG11 Coding of Moving Pictures and Audio, Oct. 2016, 15 pages. [cited by applicant]
International Search Report and Written Opinion from PCT/US2023/014753, dated Jul. 6, 2023, 17 pages. [cited by applicant]
Lee, et al., “Foveated Video Compression with Optimal Rate Control,” IEEE Transactions on Image Processing, Jul. 2001, vol. 10, No. 7, pp. 977-992. [cited by applicant]