IP Library Granted Patent US 12,677,069
Granted Patent B1
US 12,677,069 · App. 16/717,858 · Granted Jul 7, 2026

Panorama generation using one or more neural networks

Inventors: Siddhant Pardeshi (Pune, IN); Pranit P. Kothari (Pune, IN); Vinayak Vilas Gaikwad (Pune, IN)
Assignee: NVIDIA Corporation
H04N23/698G06N3/045G06N3/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,677,069
App. No.
16/717,858
Granted
Jul 7, 2026
Kind
B1
Abstract

Apparatuses, systems, and techniques are presented to generate panoramas from a set of images. In at least one embodiment, one or more generative neural networks are used to generate a spherical panoramic image using features extracted from input images.

Claims (55)

1 . One or more processors, comprising: circuitry to:

use one or more neural networks to:

generate encodings of one or more first features based, at least in part, on one or more second features extracted from a plurality of images of an environment captured from a same position in the environment facing different directions;

generate a spherical panoramic image based, at least in part, on the encodings; and

cause the spherical panoramic image to be presented using a display of a virtual reality headset.

2 . The one or more processors of claim 1 , wherein the circuitry is further to extract representative features from at least six images of the plurality of images, the representative features being provided as input to the one or more neural networks.

3 . The one or more processors of claim 2 , wherein the one or more neural networks include one or more generative adversarial networks (GANs) or variational autoencoders (VAEs) for generating the spherical panoramic image based at least in part on the representative features.

4 . The one or more processors of claim 2 , wherein the representative features are used to generate a cube map that is transformed into the spherical panoramic image.

5 . The one or more processors of claim 1 , wherein gaps in view directions of the plurality of images are filled with content generated by the one or more neural networks.

6 . The one or more processors of claim 1 , wherein the circuitry is further to perform post-processing of the spherical panoramic image to cause the spherical panoramic image to be in a format for a specified use.

7 . A system comprising:

one or more processors to:

use one or more neural networks to:

generate encodings of one or more first features based, at least in part, on one or more second features extracted from a plurality of images of an environment captured from a same position in the environment facing different directions;

generate a spherical panoramic image based, at least in part, on the encodings; and

cause the spherical panoramic images to be presented using a display of a virtual reality headset.

8 . The system of claim 7 , wherein the one or more processors are further to extract representative features from at least six images of the plurality of images, the representative features being provided as input to the one or more neural networks.

9 . The system of claim 8 , wherein the one or more neural networks include one or more generative adversarial networks (GANs) or variational autoencoders (VAEs) for generating the spherical panoramic image using the representative features.

10 . The system of claim 8 , wherein the representative features are used to generate a cube map that is transformed into the spherical panoramic image.

11 . The system of claim 7 , wherein gaps in view directions of the plurality of images are filled with content generated by the one or more neural networks.

12 . The system of claim 7 , wherein the one or more processors are further to perform post-processing of the spherical panoramic image to cause the spherical panoramic image to be in a format for a specified use.

13 . A method comprising:

using one or more neural networks to:

generate encodings of one or more first features based, at least in part, on one or more second features extracted from a plurality of images of an environment captured from a same position in the environment facing different directions;

generate a spherical panoramic image based, at least in part, on the encodings; and

causing the spherical image to be presented using a display of a virtual reality headset.

14 . The method of claim 13 , further comprising:

extracting representative features from at least six images of the plurality of images, the representative features being provided as input to the one or more neural networks.

15 . The method of claim 14 , wherein the one or more neural networks include one or more generative adversarial networks (GANs) or variational autoencoders (VAEs) for generating the spherical panoramic image using the representative features.

16 . The method of claim 14 , wherein the representative features are used to generate a cube map that is transformed into the spherical panoramic image.

17 . The method of claim 13 , wherein gaps in view directions of the plurality of images are filled with content generated by the one or more neural networks.

18 . The method of claim 13 , further comprising: performing post-processing of the spherical panoramic image to cause the spherical panoramic image to be in a format for a specified use.

19 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:

use one or more neural networks to:

generate encodings of one or more first features based, at least in part, on one or more second features extracts from a plurality of images of an environment captured from a same positions in the environment facing different directions;

generate a spherical panoramic image based, at least in part, on the encodings; and

cause the spherical image to be presented using a display of a virtual reality headset.

20 . The non-transitory machine-readable medium of claim 19 , wherein the instructions if performed further cause the one or more processors to:

extract representative features from at least six images of the plurality of images, the representative features being provided as input to the one or more neural networks.

21 . The non-transitory machine-readable medium of claim 20 , wherein the one or more neural networks include one or more generative adversarial networks (GANs) or variational autoencoders (VAEs) for generating the spherical panoramic image using the representative features.

22 . The non-transitory machine-readable medium of claim 20 , wherein the representative features are used to generate a cube map that is transformed into the spherical panoramic image.

23 . The non-transitory machine-readable medium of claim 19 , wherein gaps in view directions of the plurality of images are filled with content generated by the one or more neural networks.

24 . The non-transitory machine-readable medium of claim 19 , wherein the instructions if performed further cause the one or more processors to:

perform post-processing of the spherical panoramic image to cause the spherical panoramic image to be in a format for a specified use.

25 . A panorama generation system, comprising:

one or more processors to:

use one or more neural networks to:

generate encodings of one or more first features based, at least in part, on one or more second features extracted from a plurality of images of an environment captured from a same position in the environment facing different directions;

generate a spherical panoramic image based, at least in part, on the encodings; and

cause the spherical image to be presented using a display of a virtual reality headset.

26 . The panorama generation system of claim 25 , wherein the one or more processors are further to extract representative features from at least six images of the plurality of images, the representative features being provided as input to the one or more neural networks.

27 . The panorama generation system of claim 26 , wherein the one or more neural networks include one or more generative adversarial networks (GANs) or variational autoencoders (VAEs) for generating the spherical panoramic image using the representative features.

28 . The panorama generation system of claim 26 , wherein the representative features are used to generate a cube map that is transformed into the spherical panoramic image.

29 . The panorama generation system of claim 25 , wherein gaps in view directions of the plurality of images are filled with content generated by the one or more neural networks.

30 . The panorama generation system of claim 25 , wherein the one or more processors are further to perform post-processing of the spherical panoramic image to cause the spherical panoramic image to be in a format for a specified use.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2019
From: PARDESHI, SIDDHANT; KOTHARI, PRANIT P; GAIKWAD, VINAYAK VILAS
To: NVIDIA CORPORATION
Reel/Frame 051336/0563 →
References Cited (128)
US 10672188B2 · Bleyer et al. · 2020 [cited by applicant]
US 10706890B2 · Somanath et al. · 2020 [cited by applicant]
US 10974147B1 · Sarria, Jr. · 2021 [cited by applicant]
US 11107228B1 · Shrivastava · 2021 [cited by applicant]
US 11631156B2 · Singh et al. · 2023 [cited by applicant]
US 20090209348A1 · Roberts et al. · 2009 [cited by applicant]
US 20160034809A1 · Trenholm et al. · 2016 [cited by applicant]
US 20180035165A1 · Ogle et al. · 2018 [cited by applicant]
US 20180075581A1 · Shi et al. · 2018 [cited by applicant]
US 20180137611A1 · Kwon et al. · 2018 [cited by applicant]
US 20180302614A1 · Toksvig et al. · 2018 [cited by applicant]
US 20180349771A1 · Kamilov et al. · 2018 [cited by applicant]
US 20190007669A1 · Kim · 2019 [cited by examiner]
US 20190026956A1 · Gausebeck et al. · 2019 [cited by applicant]
US 20190066733A1 · Somanath et al. · 2019 [cited by applicant]
US 20190094542A1 · Langner et al. · 2019 [cited by applicant]
US 20190108396A1 · Dal Mutto et al. · 2019 [cited by applicant]
US 20190114807A1 · Saa-Garriga · 2019 [cited by examiner]
US 20190279075A1 · Liu et al. · 2019 [cited by applicant]
US 20190289327A1 · Lin et al. · 2019 [cited by applicant]
US 20190325644A1 · Bleyer et al. · 2019 [cited by applicant]
US 20190355126A1 · Sun · 2019 [cited by examiner]
US 20190362539A1 · Kurz et al. · 2019 [cited by applicant]
US 20190385358A1 · Chui et al. · 2019 [cited by applicant]
US 20200118253A1 · Eisenmann · 2020 [cited by examiner]
US 20200118255A1 · Bazin et al. · 2020 [cited by applicant]
US 20200162715A1 · Chaudhuri et al. · 2020 [cited by applicant]
US 20200175329A1 · Malaya · 2020 [cited by applicant]
US 20200260062A1 · Sharma et al. · 2020 [cited by applicant]
US 20200302579A1 · Eisenmann et al. · 2020 [cited by applicant]
US 20200342652A1 · Rowell et al. · 2020 [cited by applicant]
US 20200410767A1 · Williams et al. · 2020 [cited by applicant]
US 20210012561A1 · Sunkavalli et al. · 2021 [cited by applicant]
US 20210048881A1 · Huang et al. · 2021 [cited by applicant]
US 20210178274A1 · St-Pierre et al. · 2021 [cited by applicant]
US 20210227192A1 · Li · 2021 [cited by examiner]
US 20210248811A1 · Shan et al. · 2021 [cited by applicant]
US 20210295606A1 · Kim et al. · 2021 [cited by applicant]
US 20210312698A1 · He et al. · 2021 [cited by applicant]
US 20210357655A1 · Park et al. · 2021 [cited by applicant]
US 20220317055A1 · Higa · 2022 [cited by applicant]
US 20230168382A1 · Steinberg et al. · 2023 [cited by applicant]
CN 108230338A · 2018 [cited by applicant]
CN 108268845A · 2018 [cited by applicant]
CN 108305229A · 2018 [cited by applicant]
CN 108537864A · 2018 [cited by applicant]
CN 109218619A · 2019 [cited by applicant]
CN 109788270A · 2019 [cited by examiner]
CN 110223370A · 2019 [cited by applicant]
CN 110555892A · 2019 [cited by applicant]
CN 110798673A · 2020 [cited by applicant]
EP 3321881A1 · 2018 [cited by applicant]
EP 3579198A1 · 2019 [cited by applicant]
EP 3975065A1 · 2022 [cited by applicant]
KR 102173942B1 · 2020 [cited by applicant]
WO 2020242508A1 · 2020 [cited by applicant]
WO 2022042831A1 · 2022 [cited by applicant]
Flores, M., Valiente, D., Peidró, A. et al. Generating a full spherical view by modeling the relation between two fisheye images. [cited by examiner]
18th century Royal warrant commanding the preparation of letters patent granting Robert Barker use for fourteen years of his invention ‘La nature a coup d'oeil . . . for displaying views of nature . . . by oil painting’… [cited by examiner]
Goldman, A. (2021) [cited by examiner]
Livesey, J. (2019) [cited by examiner]
[cited by examiner]
Adobe, “Extensible Metadata Platform (XMP),” retrieved from https://www.adobe.com/products/xmp.html, 2012, 4 pages. [cited by applicant]
Brown, “AutoStitch: A New Dimension in Automatic Image Stitching,” retrieved from http://matthewalunbrown.com/autostitch/autostitch.html, Jul. 5, 2018, 3 pages. [cited by applicant]
Facebook, “Graph API,” retrieved from https://developers.facebook.com/docs/graph-api/reference/photo/, 3 pages. [cited by applicant]
Goodfellow et al., Generative Adversarial Networks, 2014, 9 pages. [cited by applicant]
Google, “Photo Sphere XMP Metadata,” retrieved from https://developers.google.com/streetview/spherical-metadata, Jun. 14, 2018, 10 pages. [cited by applicant]
Mordvintsev et al., “Introduction to SIFT (Scale-Invariant Feature Transform),” retrieved from https://web.archive.org/web/20161122072557/https://opencv-python tutroals.readthedocs.io/en/latest/py_tutorials/py_feature2d… [cited by applicant]
Razavi et al., “Generating Diverse High-Fidelity Images with VQ-VAE-2,” Jun. 2, 2019, 15 pages. [cited by applicant]
Wu et al., “Deep Portrait Image Completion and Extrapolation,” IEEE Transactions on Image Processing, 2019, 12 pages. [cited by applicant]
Wu et al., “Light Field Super-Resolution Using Global and Local Multi-Views,” Computer Applied Research, 36(5):, May 2019, 7 pages. [cited by applicant]
Yamazaki, “On Depth and Complexity of Generative Adversarial Networks,” 2017, 80 pages. [cited by applicant]
Yang et al., “Weakly-Supervised Disentangling with Recurrent Transformation for 3D View Synthesis,” Advances in Neural Information Processing Systems, 2015, 9 pages. [cited by applicant]
Yu et al., “Free-Form Image Inpainting with Gated Convolution,” ICCV, 2019, 10 pages. [cited by applicant]
Zhang et al., “Framebreak: Dramatic Image Extrapolation by Guided Shift-Maps,” CVPR, 2013, 8 pages. [cited by applicant]
Zhao et al., “Multi-View Image Generation from a Single-View,” Feb. 27, 2018, 9 pages. [cited by applicant]
Akimoto et al., “360-Degree Image Completion by Two-Stage Conditional GANS,” IEEE International Conference on Image Processing, 2019, 5 pages. [cited by applicant]
Chen et al. “Photographic Image Synthesis with Cascaded Refinement Networks”, arXiv:707.09405v1, dated Jul. 28, 2017, 10 pages. [cited by applicant]
Extended European Search Report for Application No. 20214107.3, mailed Apr. 21, 2021, 11 pages. [cited by applicant]
Goodfellow et al., “Generative Adversarial Nets,” In Advances in Neural Information Processing Systems, Jun. 10, 2014, 9 pages. [cited by applicant]
Habibie et al., “A Recurrent Variational Autoencoder for Human Motion Synthesis,” retrieved from https://pure.mpg.de/rest/items/item_3190428/component/file_3190429/content, 2017, 12 pages. [cited by applicant]
Han et al., “Human Body Action Identifying Method Based on the 3D Convolutional Neural Network,” 2015, 7 pages. [cited by applicant]
IEEE “IEEE Standard for Floating-Point Arithmetic”, Microprocessor Standards Committee of the IEEE Computer Society, IEEE Std 754-2008, dated Jun. 12, 2008, 70 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
Iizuka et al., “Globally and Locally Consistent Image Completion,” ACM Transactions on Graphics (ToG) 36.4, 2017, 14 pages. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/US2022/11099, mailed Mar. 30, 2022, filed Jan. 4, 2022, 9 pages. [cited by applicant]
Lee et al., “Context-Aware Synthesis and Placement of Object Instances,” Advances in Neural Information Processing Systems, 2018, 11 pages. [cited by applicant]
Li et al., “Convolutional Neural Network Based Inter-Frame Enhancement for 360-Degree Video Streaming,” Pacific Rim Conference on Multimedia, 2018, 10 pages. [cited by applicant]
Li et al., “StoryGAN: A Sequential Conditional GAN for Story Visualization,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, 10 pages. [cited by applicant]
Liang et al., “Generative Semantic Manipulation with Contrasting GAN,” 2017, 12 pages. [cited by applicant]
Lin et al. “COCO-GAN: Generation by Parts via Conditional Coordinating,” Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, 10 pages. [cited by applicant]
Lu et al., “2D-to-Stereo Panorama Conversion Using GAN and Concentric Mosaics,” IEEE Access 7, 2019, 10 pages. [cited by applicant]
McFadden et al., “Automatic 1-19 Environment Map Construction for Mixed Reality Robotics Applications,” Dec. 10, 2016, 10 pages. [cited by applicant]
Naseer et al., “Indoor Scene Understanding in 2.5/3d: Survey,” 2018, 29 pages. [cited by applicant]
Notice of Decsion to Grant for Chinese Application No. 202110962578.5, mailed Jul. 29, 2025, 5 pages. [cited by applicant]
Notice of Intention to Grant for United Kingdom Application No. GB2112182.7, mailed Nov. 20, 2024, 2 pages. [cited by applicant]
Office Action for Chinese Application No. 202011538415.6, mailed Jun. 1, 2023, 9 pages. [cited by applicant]
Office Action for Chinese Application No. 202011538415.6, mailed Mar. 12, 2024, 9 pages. [cited by applicant]
Office Action for Chinese Application No. 202011538415.6, mailed Nov. 16, 2023, 7 pages. [cited by applicant]
Office Action for Chinese Application No. 202011538415.6, mailed Sep. 26, 2024, 10 pages. [cited by applicant]
Office Action for Chinese Application No. 202110962578.5, mailed Mar. 17, 2025, 10 pages. [cited by applicant]
Office Action for Chinese Application No. 202110962578.5, mailed Sep. 14, 2024, 25 pages. [cited by applicant]
Office Action for Chinese Application No. 202280002267.7, mailed Nov. 19, 2025, 17 pages. [cited by applicant]
Office Action for European Application No. 20214107.3, mailed Apr. 28, 2023, 4 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2112182.7, mailed Jan. 17, 2024, 2 pages. [cited by applicant]
Okubo et al., “Omni-Directional Image Generation from Single Snapshot Image,” IEEE International Conference on Systems, 2020, 8 pages. [cited by applicant]
Oord et al., “Pixel Recurrent Neural Networks,” ICML, 2016, 10 pages. [cited by applicant]
Ozcinar et al., “Super-Resolution of Omnidirectional Images Using Adversarial Learning,” International Workshop on Multimedia Signal Processing, 2019, 6 pages. [cited by applicant]
Park et al., “Mc-gan: Multi-Conditional Generative Adversarial Network for Image Synthesis,” 2018, 13 pages. [cited by applicant]
Park et al., “Megan: Mixture of Experts of Generative Adversarial Networks for Multimodal Image Generation,” 2018, 7 pages. [cited by applicant]
Reed et al., “Learning What and Where to Draw,” Advances in Neural Information Processing Systems, 2019, 9 pages. [cited by applicant]
Sabini et al., “Painting Outside the Box: Image Outpainting with GANs,” 2018,. [cited by applicant]
Song et al., “Im2pano3d: Extrapolating 360 Structure and Semantics Beyond the Field of View,” CVPR, 2018, 10 pages. [cited by applicant]
Streetview, “Photo Sphere XMP Metadata,” retrieved from https://developers.google.com/streetview/spherical-metadata, Jun. 14, 2018, 10 pages. [cited by applicant]
Sultana et al., “Evolution of Image Segmentation using Deep Convolutional Neural Network: A Survey,” May 29, 2020, 38 pages. [cited by applicant]
Sumantri et al., “360 Panorama Synthesis from a Sparse Set of Images with Unknown Field of View,” Dec. 22, 2019, 10 pages. [cited by applicant]
Teterwak et al., “Boundless: Generative Adversarial Networks for Image Extension,” ICCV, 2019, 10 pages. [cited by applicant]
United Kingdom Combined Search and Examination Report for Patent Application No. 2112182.7 dated Jun. 8, 2022, 9 pages. [cited by applicant]
United Kingdom Search Report for Application No. GB2112182.7, mailed Nov. 27, 2024, 2 pages. [cited by applicant]
Verma et al., “Generalized Zero-Shot Learning via Synthesized Examples,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, 9 pages. [cited by applicant]
Wang et al., “Wide-Context Semantic Image Extrapolation,” CVPR, 2019, 10 pages. [cited by applicant]
Wessman et al., “Generation of Artificial Training Data for Deep Learning,” Master's Theses in Mathematical Sciences, 2018, 84 pages. [cited by applicant]
White, “Sampling Generative Networks,” 2016, 11 pages. [cited by applicant]
Wiles et al., “SynSin: End-to-end View Synthesis from a Single Image,” CVPR, 2020, 22 pages. [cited by applicant]
Wiles et al., “SynSin: End-to-End View Synthesis from a Single Image,” IEEE Conference on Computer Vision and Pattern Recognition, 2020, 11 pages. [cited by applicant]
Wu et al., “A Survey of Image Synthesis and Editing with Generative Adversarial Networks,” Tsinghua Science and Technology, 22(6): Dec. 2017, 15 pages. [cited by applicant]
Office Action for Chinese Application No. 202280002267.7, mailed Apr. 10, 2026, 20 pages. [cited by applicant]
Examination Report for United Kingdom Application No. GB2206726.8, mailed Apr. 28, 2026, 1 page. [cited by applicant]