IP Library Granted Patent US 12,477,231
Granted Patent B2
US 12,477,231 · App. 18/137,021 · Granted Nov 18, 2025

Apparatus and methods for image encoding using spatially weighted encoding quality parameters

Inventors: Balineedu Chowdary Adsumilli (San Francisco, CA); Adeel Abbas (Carlsbad, CA); Sumit Chawla (San Carlos, CA)
Assignee: GoPro, Inc.
H04N23/698H04N13/178H04N13/243H04N13/296H04N13/344H04N19/115H04N19/117H04N19/122H04N19/126H04N19/167H04N19/57H04N19/597
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,477,231
App. No.
18/137,021
Granted
Nov 18, 2025
Kind
B2
Abstract

Visual content that includes spatial portions is obtained. A determination is made that one of the spatial portions includes a face. Based on the determination, encoding quality parameters for the one of the spatial portions is identified. The encoding quality parameters are obtained by combining a first distortion model related to the obtaining the visual content with a second model that emphasizes the one of the spatial portions The visual content is encoded. The encoding quality parameters are stored, in association with but separate from, the one of the spatial portions. After decoding, the one of the spatial portions are rendered based on the encoding quality parameters. The encoding quality parameters are obtained by combining a first distortion model related to the obtaining the visual content with a second model that emphasizes the one of the spatial portions.

Claims (40)

1 . A method, comprising:

obtaining visual content, the visual content comprising spatial portions;

determining that one of the spatial portions includes a face;

identifying, based on the determination, encoding quality parameters for the one of the spatial portions, wherein the encoding quality parameters are obtained by combining a first distortion model related to the obtaining the visual content with a second model that provides higher encoding quality for the one of the spatial portions;

encoding the visual content;

storing, in association with but separate from the one of the spatial portions, the encoding quality parameters; and

rendering, after decoding, the one of the spatial portions based on the encoding quality parameters.

2 . The method of claim 1 , wherein the visual content is captured using a non- uniform image capture device, and wherein the visual content is one of a binocular image content, a spherical image content, or a panoramic image content.

3 . The method of claim 1 , wherein the second model provides the higher encoding quality for the one of the spatial portions by concentrating processing resources toward high encoding quality for the one of the spatial portions.

4 . The method of claim 1 , wherein the encoding quality parameters are identified based on a distortion model associated with a lens type used to obtain the visual content.

5 . The method of claim 4 , wherein the distortion model is an equirectangular projection that is such that a maximum quality equirectangular projection is at an equator of the visual content.

6 . The method of claim 1 , wherein the determination that the one of the spatial portions includes a face is based on user input.

7 . The method of claim 6 , wherein the user input is obtained after encoding the visual content.

8 . The method of claim 1 , wherein the determination that the one of the spatial portions includes a face is made by an encoding process that encodes the visual content.

9 . A device, comprising:

a memory; and

a processor, the processor configured to execute instructions stored in the memory to:

obtain encoded visual content from visual content, the encoded visual content comprising an encoded spatial portion, wherein the encoded visual content is obtained from an encoding process by instructions to:

determine that a spatial portion corresponding to the encoded spatial portion includes a face;

encode the visual content using the encoding process; and

store an indication of the spatial portion that includes the face;

identify, based on the indication, encoding quality parameters for the spatial portion; and

render, after decoding of at least the encoded spatial portion, the spatial portion based on the encoding quality parameters.

10 . The device of claim 9 , wherein the encoding quality parameters are further identified in response to a request to render the spatial portion.

11 . The device of claim 9 ,

wherein the encoded spatial portion is a first encoded spatial portion, the spatial portion is a first spatial portion, and the encoding quality parameters are first encoding quality parameters,

wherein the encoded visual content includes a second encoded spatial portion of a second spatial portion, and

wherein the processor is further configured to identify, for the second spatial portion, second encoding quality parameters that are based on a non-uniform capture of the visual content.

12 . The device of claim 11 , wherein the second encoding quality parameters are dependent on a location of the second spatial portion in the visual content.

13 . The device of claim 11 , wherein the second encoding quality parameters are identified based on a mathematical relationship.

14 . The device of claim 11 , wherein the second encoding quality parameters are identified based on a relationship of spatial portions to respective encoding quality parameters.

15 . A non-transitory memory, comprising executable instructions that, when executed by a processor, facilitate performance of operations, the operations comprising operations to:

obtain visual content, the visual content comprising spatial portions;

identify a face in a subset of the spatial portions; and

encode at least some of the spatial portions of the visual content based on spatially weighted encoding quality parameters, wherein the spatially weighted encoding quality parameters are obtained by combining a first distortion model related to the obtained visual content with a second model that provides higher encoding quality for the subset of the spatial portions.

16 . The non-transitory memory of claim 15 , wherein the visual content is captured using a non-uniform image capture, and wherein the visual content is one of a binocular image content, a spherical image content, or a panoramic image content.

17 . The non-transitory memory of claim 15 , wherein the second model provides the higher encoding quality for the subset of the spatial portions by concentrating processing resources toward high encoding quality for the subset of the spatial portions.

18 . The non-transitory memory of claim 15 , wherein the first distortion model is based on a lens type used to obtain the visual content.

19 . The non-transitory memory of claim 15 , wherein the first distortion model is an equirectangular projection that is such that a maximum quality equirectangular projection is at an equator of the visual content.

20 . The non-transitory memory of claim 15 , wherein the face is identified by an encoding process.

Assignments (3)
SECURITY INTEREST Recorded Aug 4, 2025
From: GOPRO, INC.
To: FARALLON CAPITAL MANAGEMENT, L.L.C., AS AGENT
Reel/Frame 072340/0676 →
SECURITY INTEREST Recorded Aug 4, 2025
From: GOPRO, INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 072358/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 20, 2023
From: ADSUMILLI, BALINEEDU CHOWDARY; ABBAS, ADEEL; CHAWLA, SUMIT
To: GOPRO, INC.
Reel/Frame 063387/0406 →
Continuity (5)
Continuation 17215362 · Mar 29, 2021
Continuation 16363668 · Mar 25, 2019
Continuation 15432700 · Feb 14, 2017
Provisional Application 62351818 · Jun 17, 2016
Related Publication 20240129636A1 · Apr 18, 2024
References Cited (49)
US 8218080B2 · Xu · 2012 [cited by examiner]
US 8467627B2 · Gwak · 2013 [cited by examiner]
US 8606073B2 · Woodman · 2013 [cited by applicant]
US 9171577B1 · Newman · 2015 [cited by applicant]
US 10244167B2 · Adsumilli · 2019 [cited by examiner]
US 10965868B2 · Adsumilli · 2021 [cited by applicant]
US 20030007567A1 · Newman · 2003 [cited by applicant]
US 20050249296A1 · Axnas · 2005 [cited by applicant]
US 20070126884A1 · Xu · 2007 [cited by examiner]
US 20080225129A1 · Viinikanoja · 2008 [cited by applicant]
US 20090135269A1 · Nozaki · 2009 [cited by examiner]
US 20100033552A1 · Ogawa · 2010 [cited by applicant]
US 20130202039A1 · Song · 2013 [cited by applicant]
US 20130265227A1 · Julian · 2013 [cited by examiner]
US 20130343665A1 · Maeda · 2013 [cited by applicant]
US 20140143823A1 · Manchester · 2014 [cited by examiner]
US 20150163498A1 · Shimada · 2015 [cited by applicant]
US 20160239340A1 · Chauvet · 2016 [cited by applicant]
US 20160274338A1 · Davies · 2016 [cited by applicant]
US 20170366814A1 · Adsumilli · 2017 [cited by applicant]
US 20190289208A1 · Adsumilli · 2019 [cited by applicant]
US 20210218891A1 · Adsumilli · 2021 [cited by applicant]
Achanta R., et al., ‘Slic Superpixeis Compared to State-of-The-Art Superpixei Methods,’ IEEE Transactions on Pattern Analysis and Machine intelligence, 2012, vol. 34 (11), pp. 2274-2282. [cited by applicant]
Allene C, et al,, ‘Seamless Image-based Texture Atlases Using Multi-band Blending,’ Pattern Recognition, 2008. ICPR 2008. 19th International Conference on, 2008. 4 pages. [cited by applicant]
Badrinarayanan V., et al., ‘Segnet: a Deep Convoiutional Encoder-Decoder Architecture for Image Segmentation,’ arXiv preprint arXiv: 1511.00561, 2015. 14 pages. [cited by applicant]
Barghout L. and Sheynin J., ‘Real-world scene perception and perceptual organization: Lessons from Computer Vision’. Journal of Vision, 2013, vol. 13 (9). (Abstract). 1 page. [cited by applicant]
Barghout L., ‘Visual Taxometric approach Image Segmentation using Fuzzy-Spatial Taxon Cut Yields Contextually Relevant Regions,’ Communications in Computer and Information Science (CCIS), Springer-Verlag, 2014, pp. 163-… [cited by applicant]
Bay H., et al., ‘Surf: Speeded up Robust Features,’ European Conference on Computer Vision, Springer Berlin Heidelberg, 2006, pp. 404-417. [cited by applicant]
Beier et al., ‘Feature-Based Image Metamorphosis,’ in Computer Graphics Journal, Jul. 1992, vol. 28 (2), pp. 35-42. [cited by applicant]
Brainard R.C., et al., “Low-Resolution TV: Subjective Effects of Frame Repetition and Picture Replenishment,” Bell Labs Technical Journal, Jan. 1967, vol. 46 (1), pp. 261-271. [cited by applicant]
Burt et al., ‘A Multiresolution Spline with Application to Image Mosaics,’ in ACM Transactions on Graphics (TOG), 1983, vol. 2, No. 4, pp. 217-236. [cited by applicant]
Chan et al., ‘Active contours without edges’. IEEE Transactions on Image Processing, 2001, 10 (2), pp. 266-277 (hereinafter ‘Chan’). [cited by applicant]
Chang H., etal., ‘Super-resolution Through Neighbor Embedding,’ Computer Vision and Pattern Recognition, 2004. CVPR2004. Proceedings of the 2004 IEEE Computer Society Conference on, vol. 1, 2004. 8 pages. [cited by applicant]
Elen, ‘Whatever happened to Ambisonics’ AudioMedia Magazine, Nov. 1991. 18 pages. [cited by applicant]
Gracias, et al., ‘Fast Image Blending Using Watersheds and Graph Cuts,’ Image and Vision Computing, 2009, vol. 27 (5), pp. 597-607. [cited by applicant]
Herbst E., et al., ‘Occlusion Reasoning for Temporal Interpolation Using Optical Flow,’ Department of Computer Science and Engineering, University of Washington, Tech. Rep. UW-CSE-09-08-01, 2009. 41 pages. [cited by applicant]
Jakubowski M., et al., ‘Block-based motion estimation algorithmsa survey,’ Opto-Eiectronics Review 21, No. 1 (2013), pp. 88-102. [cited by applicant]
Kendall A., et al., ‘Bayesian Segnet: Model Uncertainty in Deep Convolutional Encoder-Decoder Architectures for Scene Understanding,’ arXiv: 1511.02680, 2015. (11 pages). [cited by applicant]
Lowe, David G. “Object recognition from local scale-invariant features.” In Computer vision, 1999. The proceedings of the seventh IEEE international conference on, vol. 2, pp. 1150-1157. IEEE 1999. (Year: 1999). [cited by applicant]
Mitzel D., et al., ‘Video Super Resolution Using Duality Based TV-11 Optical Flow,’ Joint Pattern Recognition Symposium, 2009, pp. 432-441. [cited by applicant]
Perez et al., ‘Poisson Image Editing,’ in ACM Transactions on Graphics (TOG), 2003, vol. 22, No. 3, pp. 313-318. [cited by applicant]
Schick A., et al., “Improving Foreground Segmentations with Probabilistic Superpixel Markov Random Fields,” 2012 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, 2012, pp. 27-31. [cited by applicant]
Suzuki et al., ‘Inter Frame Coding with Template Matching Averaging,’ in IEEE international Conference on Image Processing Proceedings (2007), vol. (ill), pp. 409-412. [cited by applicant]
Szeliski R., “Computer Vision: Algorithms and Applications,” Springer Science & Business Media, 2010, 979 pages. [cited by applicant]
Thaipanich T., et al., “Low Complexity Algorithms for Robust Video frame rate up-conversion (FRUC) technique,” IEEE Transactions on Consumer Electronics, Feb. 2009, vol. 55 (1), pp. 220-228. [cited by applicant]
Xiao, et al., ‘Multiple View Semantic Segmentation for Street View Images,’ 2009 IEEE 12th International Conference on Computer Vision, 2009, pp. 686-693. [cited by applicant]
Xiong Y et al., ‘Gradient Domain Image Blending and Implementation on Mobile Devices,’ International Conference on Mobile Computing, Applications, and Services, Springer Berlin Heidelberg, 2009, pp. 293-306. [cited by applicant]
Zhai et al., “A Low Complexity Motion Compensated Frame Interpolation Method,” in IEEE International Symposium on Circuits and Systems (2005), pp. 4927-4930. [cited by applicant]
Zhang., “A Flexible New Technique for Camera Calibration” IEEE Transactions, dated Nov. 2000, vol. 22, No. 11, pp. 1330-1334. [cited by applicant]