IP Library Granted Patent US 12,266,113
Granted Patent B2
US 12,266,113 · App. 17/617,560 · Granted Apr 1, 2025

Automatically segmenting and adjusting images

Inventors: Orly Liba (Palo Alto, CA); Florian Kainz (San Rafael, CA); Longqi Cai (Mountain View, CA); Yael Pritch Knaan (Mountain View, CA)
Assignee: Google LLC
G06T7/11G06T5/60G06T5/70G06T5/94G06V10/764
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,266,113
App. No.
17/617,560
Granted
Apr 1, 2025
Kind
B2
Abstract

A device automatically segments an image into different regions and automatically adjusts perceived exposure-levels or other characteristics associated with each of the different regions, to produce pictures that exceed expectations for the type of optics and camera equipment being used and in some cases, the pictures even resemble other high-quality photography created using professional equipment and photo editing software. A machine-learned model is trained to automatically segment an image into distinct regions. The model outputs one or more masks that define the distinct regions. The mask(s) are refined using a guided filter or other technique to ensure that edges of the mask(s) conform to edges of objects depicted in the image. By applying the mask(s) to the image, the device can individually adjust respective characteristics of each of the different regions to produce a higher-quality picture of a scene.

Claims (47)

1. A computer-implemented method comprising:

receiving, by a processor of a computing device, an original image captured by a camera;

automatically segmenting, by the processor, the original image into multiple regions of pixels by outputting a mask indicating which pixels of the original image or a down-sampled version of the original image are contained within each of the multiple regions;

using a guided filter to add matting to each of the multiple regions to refine the mask to generate a refined mask;

independently adjusting, by the processor, a respective characteristic of each of the multiple regions as indicated by the refined mask;

combining, by the processor, the multiple regions to form a new image after independently adjusting the respective characteristic of each of the multiple regions; and

outputting, by the processor and for display, the new image.

2. The computer-implemented method of claim 1 , wherein automatically segmenting the original image into the multiple regions comprises:

inputting the original image or the downsampled version of the original image into a machine-learned model, wherein the machine-learned model is configured to output the mask indicating which pixels of the original image or the downsampled version of the original image are contained within each of the multiple regions.

3. The computer-implemented method of claim 2 , wherein the mask further indicates a respective degree of confidence associated with each of the pixels of the original image or the downsampled version of the original image that the pixel is contained within each of the multiple regions.

4. The computer-implemented method of any of claim 1 , wherein independently adjusting the respective characteristic of each of the multiple regions comprises independently applying, by the processor, a respective auto-white-balancing to each of the multiple regions.

5. The computer-implemented method of claim 4 , wherein independently applying the respective auto-white-balancing to each of the multiple regions comprises:

determining a respective type classifying each of the multiple regions;

selecting the respective auto-white-balancing for each of the multiple regions based on the respective type; and

applying the respective auto-while-balancing selected for each of the multiple regions.

6. The computer-implemented method of claim 5 , wherein the respective type classifying each of the multiple regions includes a sky region including a pixel representation of a sky or a non-sky region including a pixel representation of one or more objects other than the sky.

7. The computer-implemented method of any of claim 4 , wherein selecting the respective auto-white-balancing for each of the multiple regions based on the respective type comprises selecting a first auto-white-balancing for each of the multiple regions that is determined to be a sky region and selecting a second, different auto-white-balancing for each of the multiple regions that is determined to be a non-sky region.

8. The computer-implemented method of any of claim 1 , wherein independently adjusting the respective characteristic of each of the multiple regions comprises independently adjusting, by the processor and according to a respective type classifying each of the multiple regions, a respective brightness associated with each of the multiple regions.

9. The computer-implemented method of any of claim 1 , wherein independently adjusting the respective characteristic of each of the multiple regions comprises independently removing, by the processor and according to a respective type classifying each of the multiple regions, noise from each of the multiple regions.

10. The computer-implemented method of claim 9 , further comprising:

determining the respective type classifying each of the multiple regions; and

in response to determining the respective type classifying a particular region from the multiple regions is a sky region type, averaging groups of noisy pixels in the particular region from the multiple regions that are a size and frequency that satisfies a threshold.

11. The computer-implemented method of claim 10 , further comprising:

further in response to determining the respective type classifying the particular region from the multiple regions is the sky region type, retaining groups of noisy pixels in the particular region from the multiple regions that are a size and frequency that is less than or greater than the threshold.

12. The computer-implemented method of any of claim 1 , wherein outputting the new image comprises outputting, for display, the new image automatically and in response to receiving an image capture command from an input component of the computing device, the image capture command directing the camera to capture the original image.

13. A computing device comprising at least one processor configured to:

receive an original image captured by a camera;

automatically segment the original image into multiple regions of pixels by outputting a mask indicating which pixels of the original image or a downsampled version of the original image are contained within each of the multiple regions;

use a guided filter to add matting to each of the multiple regions to refine the mask to generate a refined mask;

independently adjust a respective characteristic of each of the multiple regions as indicated by the refined mask;

combine the multiple regions to form a new image after independently adjusting the respective characteristic of each of the multiple regions; and

output, for display, the new image.

14. The computing device of claim 13 , wherein the at least one processor is configured to automatically segment the original image into the multiple regions by:

inputting the original image or the downsampled version of the original image into a machine-learned model, wherein the machine-learned model is configured to output the mask indicating which pixels of the original image or the downsampled version of the original image are contained within each of the multiple regions.

15. The computing device of claim 14 , wherein the mask further indicates a respective degree of confidence associated with each of the pixels of the original image or the downsampled version of the original image that the pixel is contained within each of the multiple regions.

16. The computing device of claim 13 , wherein the at least one processor is configured to independently adjust the respective characteristic of each of the multiple regions by independently applying a respective auto-white-balancing to each of the multiple regions.

17. The computing device of claim 13 , wherein the at least one processor is configured to independently apply the respective auto-white-balancing to each of the multiple regions by:

determining a respective type classifying each of the multiple regions;

selecting the respective auto-white-balancing for each of the multiple regions based on the respective type; and

applying the respective auto-while-balancing selected for each of the multiple regions.

18. A non-transitory computer readable medium comprising program instructions executable by at least one processor to:

receive an original image captured by a camera;

automatically segment the original image into multiple regions of pixels by outputting a mask indicating which pixels of the original image or a downsampled version of the original image are contained within each of the multiple regions;

use a guided filter to add matting to each of the multiple regions to refine the mask to generate a refined mask;

independently adjust a respective characteristic of each of the multiple regions as indicated by the refined mask;

combine the multiple regions to form a new image after independently adjusting the respective characteristic of each of the multiple regions; and

output, for display, the new image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 8, 2021
From: LIBA, ORLY; KAINZ, FLORIAN; CAI, LONGQI; KNAAN, YAEL PRITCH
To: GOOGLE LLC
Reel/Frame 058340/0114 →
Continuity (1)
Related Publication 20220230323A1 · Jul 21, 2022
References Cited (51)
US 6804409B2 · Sobol et al. · 2004 [cited by applicant]
US 7099056B1 · Kindt · 2006 [cited by applicant]
US 8483479B2 · Kunkel et al. · 2013 [cited by applicant]
US 8890975B2 · Baba et al. · 2014 [cited by applicant]
US 9002068B2 · Wang · 2015 [cited by applicant]
US 10521952B2 · Ackerson et al. · 2019 [cited by applicant]
US 10742899B1 · Ma · 2020 [cited by examiner]
US 20090207266A1 · Yoda · 2009 [cited by applicant]
US 20100149420A1 · Zhang · 2010 [cited by examiner]
US 20100322513A1 · Xu · 2010 [cited by examiner]
US 20110090303A1 · Wu et al. · 2011 [cited by applicant]
US 20110117959A1 · Rolston · 2011 [cited by examiner]
US 20120002074A1 · Baba et al. · 2012 [cited by applicant]
US 20120051635A1 · Kunkel et al. · 2012 [cited by applicant]
US 20130314558A1 · Ju et al. · 2013 [cited by applicant]
US 20140118578A1 · Sasaki · 2014 [cited by examiner]
US 20140126780A1 · Wang · 2014 [cited by applicant]
US 20190130630A1 · Ackerson et al. · 2019 [cited by applicant]
US 20190171908A1 · Salavon · 2019 [cited by applicant]
US 20210329150A1 · Wang et al. · 2021 [cited by applicant]
KR 20040077240 · 2004 [cited by applicant]
WO WO2017071644A1 · 2017 [cited by examiner]
WO 2021010974 · 2021 [cited by applicant]
Z. Yan, H. Zhang, B. Wang, S. Paris, and Y. Yu, “Automatic photo adjustment using Deep Neural Networks,” ACM Transactions on Graphics, vol. 35, No. 2, pp. 1-15, Feb. 2016. doi: 10.1145/2790296 (Year: 2016). [cited by examiner]
K. He, J. Sun and X. Tang, “Guided Image Filtering,” in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 35, No. 6, pp. 1397-1409, Jun. 2013, doi: 10.1109/TPAMI.2012.213 (Year: 2013). [cited by examiner]
Kaufman, D. Lischinski, and M. Werman, “Content-Aware automatic photo enhancement,” Computer Graphics Forum, vol. 31, No. 8, pp. 2528-2540, Sep. 2012. doi: 10.1111/j.1467-8659.2012.03225.x (Year: 2012). [cited by examiner]
B. Jiang et al., “Single image fog and haze removal based on self-adaptive guided image filter and color channel information of Sky Region,” Multimedia Tools and Applications, vol. 77, No. 11, pp. 13513-13530, Jul. 2017… [cited by examiner]
S. Nam, and S.J. Kim, “Deep Semantics-Aware Photo Adjustment”, ArXiv Computer Vision and Pattern Recognition, Jun. 2017. doi: 1706.08260 (Year: 2017). [cited by examiner]
G. Hu and J. Clark, “Instance Segmentation Based Semantic Matting for Compositing Applications,” 2019 16th Conference on Computer and Robot Vision (CRV), Kingston, QC, Canada, 2019, pp. 135-142, doi: 10.1109/CRV.2019.00… [cited by examiner]
Xu, N., Price, B., Cohen, S., & Huang, T. (2017). Deep image matting. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 2970-2979). (Year: 2017). [cited by examiner]
“Improving Autofocus Speed in Macrophotography Applications”, Technical Disclosure Commons—https://www.tdcommons.org/dpubs_series/5133, May 11, 2022, 9 pages. [cited by applicant]
“Reinforcement Learning—Wikipedia”, Retrieved from: https://en.wikipedia.org/wiki/Reinforcement_learning—on May 11, 2022, 16 pages. [cited by applicant]
Gordon, “How to Use HDR Mode When Taking Photos”, Dec. 17, 2020, 13 pages. [cited by applicant]
Hasinoff, et al., “Burst photography for high dynamic range and low-light imaging on mobile cameras”, Nov. 2016, 12 pages. [cited by applicant]
Hollister, “How to guarantee the iPhone 13 Pro's macro mode is on”, Retrieved at: https://www.theverge.com/22745578/iphone-13-pro-macro-mode-how-to, Oct. 26, 2021, 11 pages. [cited by applicant]
Kains, et al., “Astrophotography with Night Sight on Pixel Phones”, https://ai.googleblog.com/2019/11/astrophotography-with-night-sight-on.html, Nov. 26, 2019, 8 pages. [cited by applicant]
Konstantinova, et al., “Fingertip Proximity Sensor with Realtime Visual-based Calibration”, Oct. 2016, 6 pages. [cited by applicant]
“International Preliminary Report on Patentability”, Application No. PCT/US2019/041863, Jan. 18, 2022, 10 pages. [cited by applicant]
“International Search Report and Written Opinion”, PCT Application No. PCT/US2019/041863, Mar. 16, 2020, 17 pages. [cited by applicant]
Chatterjee, et al., “Clustering-Based Denoising with Locally Learned Dictionaries”, IEEE Transactions on Image Processing, vol. 18, No. 7, Jul. 2009, 14 pages. [cited by applicant]
Fried, et al., “Perspective-Aware Manipulation of Portrait Photos”, SIGGRAPH '16 Technical Paper, Jul. 24-28, 2016, Anaheim, CA, 10 pages. [cited by applicant]
Gao, et al., “Scene Metering and Exposure Control for Enhancing High Dynamic Range Imaging”, Technical Disclosure Commons; Retrieved from https://www.tdcommons.org/dpubs_series/3092, Apr. 1, 2020, 12 pages. [cited by applicant]
Gupta, “Techniques for Automatically Determining a Time-Lapse Frame Rate”, Technical Disclosure Commons, Retrieved from https://www.tdcommons.org/dpubs_series/3237, May 15, 2020, 11 pages. [cited by applicant]
Huth, et al., “Significantly Improved Precision of Cell Migration Analysis in Time-Lapse Video Microscopy Through Use of a Fully Automated Tracking System”, Apr. 2010, 12 pages. [cited by applicant]
Kaufman, et al., “Content-Aware Automatic Photo Enhancement”, Computer Graphics Forum, vol. 31, No. 8, pp. 2528-2540, 2012, 13 pages. [cited by applicant]
La Place, et al., “Segmenting Sky Pixels in Images”, 2017, 11 pages. [cited by applicant]
Mohtasham, et al., “Real-time Image Signal Processor Stats Management to Save Power and CPU Cycles”, Technical Disclosure Commons; Retrieved from https://www.tdcommons.org/dpubs_series/3308, Jun. 10, 2020, 7 pages. [cited by applicant]
Moraldo, “Virtual Camera Image Processing”, Technical Disclosure Commons; Retrieved from https://www.tdcommons.org/dpubs_series/3072, Mar. 30, 2020, 10 pages. [cited by applicant]
Yang, et al., “Improved Object Detection in an Image by Correcting Regions with Distortion”, Technical Disclosure Commons; Retrieved from https://www.tdcommons.org/dpubs_series/3090, Apr. 1, 2020, 8 pages. [cited by applicant]
Yang, et al., “Using Image-Processing Settings to Determine an Optimal Operating Point for Object Detection on Imaging Devices”, Technical Disclosure Commons; Retrieved from https://www.tdcommons.org/dpubs_series/2985, … [cited by applicant]
Zhu, et al., “What's the Role of Image Matting in Image Segmentation?”, Proceeding of the IEEE International Conference on Robotics and Biomimetrics, Shenzhen, China, Dec. 2013, 4 pages. [cited by applicant]