IP Library › Granted Patent US 12,236,640
Granted Patent B2
US 12,236,640 · App. 17/656,796 · Granted Feb 25, 2025

Machine learning based image calibration using dense fields

Inventors: Jianming Zhang (Campbell, CA); Linyi Jin (Ann Arbor, MI); Kevin Matzen (San Jose, CA); Oliver Wang (Seattle, WA); Yannick Hold-Geoffroy (San Jose, CA)
Assignee: Adobe Inc.
G06T7/80G06N3/045G06T9/002G06T11/00G06T2207/20081G06T2207/20084G06V10/764
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,236,640
App. No.
17/656,796
Granted
Feb 25, 2025
Kind
B2
Abstract

Systems and methods for image dense field based view calibration are provided. In one embodiment, an input image is applied to a dense field machine learning model that generates a vertical vector dense field (VVF) and a latitude dense field (LDF) from the input image. The VVF comprises a vertical vector of a projected vanishing point direction for each of the pixels of the input image. The latitude dense field (LDF) comprises a projected latitude value for the pixels of the input image. A dense field map for the input image comprising the VVF and the LDF can be directly or indirectly used for a variety of image processing manipulations. The VVF and LDF can be optionally used to derive traditional camera calibration parameters from uncontrolled images that have undergone undocumented or unknown manipulations.

Claims (42)

1. A method comprising:

generating, with a dense field machine learning model, a vertical vector dense field from an input image, the vertical vector dense field comprising a vertical vector of a projected vanishing point direction for a plurality of pixels of the input image;

generating, with the dense field machine learning model, a latitude dense field from the input image, the latitude dense field comprising a projected latitude value for the plurality of pixels of the input image; and

generating, with an image processing application, at least one modification to a first image based on the vertical vector dense field and the latitude dense field, to produce an output comprising a second image.

2. The method of claim 1 , further comprising:

computing, with the image processing application, one or more camera parameters from the vertical vector dense field and the latitude dense field.

3. The method of claim 1 , further comprising:

computing, with the image processing application, one or more camera parameters from the vertical vector dense field and the latitude dense field by optimizing for a set of camera parameters that predict a dense field view represented by the vertical vector dense field and the latitude dense field.

4. The method of claim 1 , further comprising:

for each of the plurality of pixels of the input image, generating, with the dense field machine learning model, a first confidence vector corresponding to a set of vertical vector bins; and

for each of the plurality of pixels of the input image, generating, with the dense field machine learning model, a second confidence vector corresponding to a set of latitude value bins.

5. The method of claim 1 , further comprising:

generating, with the dense field machine learning model, at least one dense field map for the input image comprising the vertical vector dense field and the latitude dense field.

6. The method of claim 1 , wherein the projected latitude value comprises a world spherical coordinate system latitude value.

7. The method of claim 1 , further comprising:

converting, with an image converter, the input image from a non-raster image to a raster image.

8. The method of claim 1 , wherein the dense field machine learning model comprises an encoder-decoder machine learning neural network that includes a feature encoder, a vertical vector dense field decoder, and a latitude dense field decoder.

9. The method of claim 1 , wherein the vertical vector dense field and the latitude dense field indicate a confidence for each vertical vector and each projected latitude value computed by the dense field machine learning model.

10. An image processing environment comprising:

a dense field machine learning model configured to receive an input image, wherein the dense field machine learning model is trained to:

generate a vertical vector dense field from the input image, the vertical vector dense field comprising a vertical vector of a projected vanishing point direction for a plurality of pixels of the input image; and

generate a latitude dense field from the input image, the latitude dense field comprising a projected latitude value for the plurality of pixels of the input image.

11. The image processing environment of claim 10 , further comprising:

an image processing application executed by at least one controller; and

wherein the image processing application is configured to receive a first image from an image source and provide a raster version of the first image to the dense field machine learning model as the input image.

12. The image processing environment of claim 10 , further comprising:

an image processing application executed by at least one controller, wherein the image processing application comprises an optimizer; and

wherein the optimizer is programmed to receive the vertical vector dense field and the latitude dense field from the dense field machine learning model and compute one or more camera parameters from the vertical vector dense field and the latitude dense field.

13. The image processing environment of claim 10 , wherein the dense field machine learning model produces at least one dense field map for the input image comprising the vertical vector dense field and the latitude dense field.

14. The image processing environment of claim 13 , further comprising:

an image processing application executed by at least one controller, wherein the image processing application generates at least one modification to a first image based on the vertical vector dense field and the latitude dense field, to produce an output comprising a second image.

15. The image processing environment of claim 10 , wherein the dense field machine learning model comprises an encoder-decoder machine learning neural network that includes a feature encoder, a vertical vector dense field decoder, and a latitude dense field decoder.

16. The image processing environment of claim 10 , wherein for each of the plurality of pixels of the input image, the dense field machine learning model generates a first confidence vector for each vertical vector of the vertical vector dense field, and a second confidence vector for each projected latitude value of the latitude dense field.

17. A non-transitory computer-readable medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising:

causing a dense field machine learning model to produce at least one dense field map for an input image, the at least one dense field map comprising a vertical vector dense field and a latitude dense field; and

applying at least one imaging function to a first input image based on the at least one dense field map to produce a second image comprising a processed image.

18. The non-transitory computer-readable medium of claim 17 , the operations further comprising:

computing one or more camera parameters from the vertical vector dense field and the latitude dense field by optimizing for a set of camera parameters that predict the vertical vector dense field and the latitude dense field.

19. The non-transitory computer-readable medium of claim 17 , wherein:

the vertical vector dense field comprises a vertical vector of a projected up direction for a plurality of pixels of the input image; and

the latitude dense field comprises a projected latitude value for the plurality of pixels of the input image.

20. The non-transitory computer-readable medium of claim 17 , wherein the at least one dense field map indicates a confidence computed by the dense field machine learning model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 27, 2022
From: ZHANG, JIANMING; HOLD-GEOFFROY, YANNICK; WANG, OLIVER; JIN, LINYI; MATZEN, KEVIN
To: ADOBE INC.
Reel/Frame 059751/0050 →
Continuity (1)
Related Publication 20230306637A1 · Sep 28, 2023
References Cited (50)
Hold-Geoffroy et al., A perceptual measure for deep single image camera calibration, CVPR (Year: 2018). [cited by examiner]
Lee et al., Ctrl-c: Camera calibration transformer with line-classification, ICCV (Year: 2021). [cited by examiner]
Lee et al., Automatic upright adjustment of photographs with robust camera calibration, IEEE TPAMI (Year: 2014). [cited by examiner]
Bogdan et al, Deepcalib: A deep learning approach for automatic intrinsic calibration of wide field-of-view cameras, CVMP (Year: 2018). [cited by examiner]
Ahmadyan, A., Zhang, L., Ablavatski, A., Wei ,J., Grundmann, M., “Objectron: A large scale dataset of object-centric videos in the wild with pose annotations,” CVPR, 2021, <https://arxiv.org/abs/2012.09988>, 10 pages. [cited by applicant]
Antunes, M., Barreto, J.P., Aquada, D., Ottersten, B., “Unsupervised vanishing point detection and camera calibration from a single manhattan image with radial distortion,” CVPR, 2017, <https://www.researchgate.net/publ… [cited by applicant]
Armeni, I., Sax, S., Zamir, A.R., Savarese, S., “Joint 2d-3d-semantic data for indoor scene understanding,” 2017, <https://arxiv.org/abs/1702.01105>, 9 pages. [cited by applicant]
Azadi, S., Pathak, D., Ebrahimi, S., Darrell, T., “Compositional gan: Learning image-conditional binary composition,” International Journal of Computer Vision, 128(10), 2020, <https://arxiv.org/abs/1807.07560>, 15 pages. [cited by applicant]
Bansal, A., Russell, B., Gupta, A., “Marr Revisited: 2D-3D model alignment via surface normal prediction,” CVPR, 2016, <https://arxiv.org/abs/1604.01347>, 11 pages. [cited by applicant]
Beardsley, P., Murray, D., “Camera calibration using vanishing points,” BMVC, 1992, <https://www.researchgate.net/publication/239405791_A_new_camera_calibration_method_from_vanishing_points_in_a_vision_system>, 10 pages. [cited by applicant]
Chen, Q., Wu, H., Wada, T. “Camera calibration with two arbitrary coplanar circles,” ECCV, 2004, <https://www.researchgate.net/publication/221304888_Camera_Calibration_with_Two_Arbitrary_Coplanar_Circles>, 13 pages. [cited by applicant]
Esser, P., Rombach, R., Ommer, B., “Taming transformers for high-resolution imagesynthesis,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, <https://openaccess.thecvf.com/conten… [cited by applicant]
Furukawa, Y. et al., “Accurate camera calibration from multi-view stereo and bundle adjustment,” IJCV, 2009, <https://www.di.ens.fr/willow/pdfs/cvpr08a.pdf>, 12 pages. [cited by applicant]
Heikkila, J. et al., “A four-step camera calibration procedure with implicit image correction,” CVPR, 1997, <http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.94.7536&rep=rep1&type=pdf>, 7 pages. [cited by applicant]
Hold-Geoffroy, Y. et al., “A perceptual measure for deep single image camera calibration,” CVPR, 2018, <https://arxiv.org/abs/1712.01259?context=cs>, 10 pages. [cited by applicant]
Huang, J. et al., “Framenet: Learning local canonical frames of 3d surfaces from a single rgb image,” Proceedings of the IEEE/CVF, International Conference on Computer Vision, 2019, <https://arxiv.org/abs/1903.12305>, 1… [cited by applicant]
Karsch, K. et al., “Automatic scene inference for 3d object compositing,” ACM Transactions on Graphics, 2014, <https://arxiv.org/abs/1912.12297>, 14 pages. [cited by applicant]
Kirillov, A. et al., “PointRend: Image segmentation as rendering,” 2019, <https://arxiv.org/abs/1912.08193>, 10 pages. [cited by applicant]
Lalonde, J.F. et al., “Photo clip art,” ACM transactions on graphics, 2007. [cited by applicant]
Lee, D. et al., “Context-aware synthesis and placement of object instances,” Advances in neural information processing systems 31, 2018, <https://papers.nips.cc/paper/2018/hash/c6969ae30d99f73951cb976b88a457af-Abstract.… [cited by applicant]
Lee, H., et al., “Automatic upright adjustment of photographs with robust camera calibration,” TPAMI, 2014, <https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.722.5040&rep=rep1&type=pdf>, 13 pages. [cited by applicant]
Lee, J. et al., “Ctrl-c: Camera calibration transformer with line-classification,” ICCV, 2021, <https://arxiv.org/abs/2109.02259>, 14 pages. [cited by applicant]
Lee, J. et al., “Neural geometric parser for single image camera calibration,” ECCV, 2020, <https://arxiv.org/abs/2007.11855>, 25 pages. [cited by applicant]
Li, X. et al., “Blind geometric distortion correction on images through deep learning,” CVPR, 2019, <https://arxiv.org/abs/1909.03459>, 10 pages. [cited by applicant]
Lin, C.H. et al., “St-gan: Spatial transformer generative adversarial networks for image compositing,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, <https://arxiv.org/abs/1803.01837>, 13 page… [cited by applicant]
Lin, T.Y. et al., “Microsoft coco: Common objects in context,” ECCV, 2014, <https://arxiv.org/abs/1405.0312>, 15 pages. [cited by applicant]
Lopez, M. et al., “Deep single image camera calibration with radial distortion,” CVPR, 2019, <https://openaccess.thecvf.com/content_CVPR_2019/papers/Lopez_Deep_Single_Image_Camera_Calibration_With_Radial_Distortion_CVPR… [cited by applicant]
Mei, C. et al., “Single view point omnidirectional camera calibration from planar grids,” ICRA, 2007, <https://www.robots.ox.ac.uk/˜cmei/articles/single_viewpoint_calib_mei_07.pdf>, 6 pages. [cited by applicant]
Melo, R. et al., “Unsupervised intrinsic calibration from a single frame using a ‘plumb-line’ approach,” ICCV, 2013, <https://home.deec.uc.pt/˜jpbar/Publication_Source/meloICCV2013.pdf>, 8 pages. [cited by applicant]
Park, T. et al., “Semantic image synthesis with spatially-adaptive normalization,” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, <https://arxiv.org/abs/1903.07291?source=post_p… [cited by applicant]
Rother, Carsten, “A new approach to vanishing point detection in architectural environments,” BMVC, 2002, <http://www.bmva.org/bmvc/2000/papers/p39.pdf>, 10 pages. [cited by applicant]
Strecha, C. et al., “On benchmarking camera calibration and multi-view stereo for high resolution imagery,” In: CVPR, 2008, <http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.329.6115&rep=rep1&type=pdf>. [cited by applicant]
Sturm, P.F. et al., “On plane-based camera calibration: A general algorithm, singularities, applications,” In: CVPR, 1999, <https://hal.inria.fr/inria-00525681/document>, 8 pages. [cited by applicant]
Wang, R., et al.,“VPLNet: Deep single view normal estimation with vanishing points and lines,” In: CVPR, 2020, <http://szeliski.org/papers/Wang_VPLNet_CVPR20.pdf>, 10 pages. [cited by applicant]
Wang, W. et al., “TartanAir: A dataset to push the limits of visual slam,” 2020, <https://arxiv.org/abs/2003.14338>, 8 pages. [cited by applicant]
Wang, Y. et al., “Repopulating street scenes,” In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5110-5119, 2021, <https://arxiv.org/abs/2103.16183>, 10 pages. [cited by applicant]
Workman, S. et al., “Deepfocal: A method for direct focal length estimation,” In: ICIP, 2015, <http://cs.uky.edu/˜scott/resources/deepfocal.pdf>, 5 pages. [cited by applicant]
Workman, S. et al., “Horizon lines in the wild,” BMVC, 2016, <https://a?rxiv.org/abs/1604.02129>, 12 pages. [cited by applicant]
Xian, W. et al., “Uprightnet: geometry-aware camera orientation estimation from single images,” In: ICCV, 2019, <https://arxiv.org/abs/1908.07070>, 10 pages. [cited by applicant]
Xie, E. et al., “SegFormer: simple and efficient design for semantic segmentation with transformers,” In: NeurIPS, 2021, <https://arxiv.org/abs/2105.15203>, 18 pages. [cited by applicant]
Zhai. M. et al., “Detecting vanishing points using global image context in a non-manhattan world,” In: CVPR, 2016, <https://arxiv.org/abs/1608.05684>, 9 pages. [cited by applicant]
Zhang, L. et al., “Learning object placement by inpainting for compositional data augmentation,” In: European Conference on Computer Vision, 2020, <https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123580562.pdf>… [cited by applicant]
Zhang, Z. et al., “A flexible new technique for camera calibration,” TPAMI, 2000, https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/tr98-71.pdf>, 22 pages. [cited by applicant]
Zhu, R. et al., “Single view metrology in the wild,” In: European Conference on Computer Vision, 2020, <https://arxiv.org/abs/2007.09529>, 17 pages. [cited by applicant]
Roberts, M. et al., “Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding,” In ICCV 2021, 2021, https://arxiv.org/abs/2011.02523>, 11 pages. [cited by applicant]
Kullback, S. and Leibler, R.A., “On Information and Sufficiency,” 1951, <https://projecteuclid.org/journals/annals-of-mathematical-statistics/volume-22/issue-1/On-Information-and-Sufficiency/10.1214/aoms/1177729694.full… [cited by applicant]
Luo, X. et al., “Consistent Video Depth Estimation,” ACM Transactions on Graphics, vol. 39, No. 4, Article 71, Jul. 2020, <https://arxiv.org/abs/2004.15021>, 13 pages. [cited by applicant]
“Tilt-shift photography,” Wikipedia, accessed Jun. 2022, <https://en.wikipedia.org/wiki/Tilt%E2%80%93shift_photography>, 13 pages. [cited by applicant]
“The Great Panini Projection,” Wikipedia, accessed Jun. 2022, <https://wiki.panotools.org/The_General_Panini_Projection>, 4 pages. [cited by applicant]
“Box plot,” Wikipedia, accessed Jun. 2022, <https://en.wikipedia.org/wiki/Box_plot>, 8 pages. [cited by applicant]