IP Library › Granted Patent US 12,444,140
Granted Patent B2
US 12,444,140 · App. 18/517,455 · Granted Oct 14, 2025

Virtual walkthrough experience generation based on neural radiance field model renderings

Inventors: Carlos Montero, Jr. (Brier, WA); Marcos Seefelder de Assis Araujo (London, GB); Cardin Everett Moffett (Denver, CO); Charles Goran (Lafayette, CO)
Assignee: GOOGLE LLC
G06T19/003G06F16/787G06T7/10H04N13/354
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,444,140
App. No.
18/517,455
Granted
Oct 14, 2025
Kind
B2
Abstract

Systems and methods for generating and providing a virtual walkthrough interface can include generating a virtual walkthrough video based on view synthesis renderings generated by neural radiance field model. The neural radiance field model can be trained based on a plurality of images of an environment and may generate the view synthesis renderings based on processing positions along a determined walkthrough path. The generated virtual walkthrough video can then be scrubbed through to provide the virtual walkthrough interface.

Claims (86)

1. A computer-implemented method for generating a virtual walkthrough video, the method comprising:

obtaining, by a computing system comprising one or more processors, one or more neural radiance field models associated with an environment, wherein the one or more neural radiance field models were trained to generate view renderings of the environment, wherein the environment is associated with a geographic location;

processing, by the computing system, a plurality of positions with the one or more neural radiance field models to generate a plurality of view synthesis renderings of the environment, wherein the plurality of positions are associated with a plurality of locations within the environment, and wherein the plurality of view synthesis renderings are descriptive of the environment from the plurality of positions;

generating, by the computing system, a virtual walkthrough video based on the plurality of view synthesis renderings of the environment, wherein the virtual walkthrough video is descriptive of a sequence of views of the environment, wherein generating, by the computing system, the virtual walkthrough video based on the plurality of view synthesis renderings of the environment comprises:

generating, by the computing system, a first rendering video based on rendering the sequence of views of the environment in a first direction;

generating, by the computing system, a second rendering video based on rendering the sequence of views of the environment in a second direction; and

generating, by the computing system, the virtual walkthrough video by combining the first rendering video and the second rendering video; and

storing, by the computing system, the virtual walkthrough video in a database, wherein storing the virtual walkthrough video comprises indexing the virtual walkthrough video with the geographic location associated with the environment.

2. The method of claim 1 , further comprising:

obtaining, by the computing system, a search query;

determining, by the computing system, the search query is associated with the geographic location;

in response to determining the search query is associated with the geographic location, obtaining, by the computing system, the virtual walkthrough video from the database based on the geographic location; and

providing, by the computing system, the virtual walkthrough video for display.

3. The method of claim 2 , wherein the search query comprises a text string, wherein the text string is associated with one or more entities; and

wherein determining the geographic location associated with the virtual walkthrough video is associated with the search query comprises: determining the one or more entities is associated with the geographic location.

4. The method of claim 1 , further comprising:

providing, by the computing system, a map interface for display, wherein the map interface comprises map information associated with the geographic location;

obtaining, by the computing system, a selection of a virtual walkthrough user interface element;

determining, by the computing system, the virtual walkthrough video is associated with the geographic location; and

providing, by the computing system, the virtual walkthrough video for display.

5. The method of claim 1 , wherein the first rendering video is associated with a first portion of the virtual walkthrough video, and wherein the second rendering video is associated with a second portion of the virtual walkthrough video.

6. The method of claim 1 , wherein processing, by the computing system, the plurality of positions with the one or more neural radiance field models to generate the plurality of view synthesis renderings of the environment comprises:

for each position of the plurality of positions:

processing, by the computing system, the position with the one or more neural radiance field models to generate a plurality of directional view synthesis renderings, wherein the plurality of directional view synthesis renderings are associated with a plurality of view directions for the position; and

generating, by the computing system, a respective view synthesis rendering for the position by stitching the plurality of directional view synthesis renderings to generate a panoramic image rendering for the position.

7. The method of claim 1 , further comprising:

obtaining, by the computing system, a plurality of images of the environment;

training, by the computing system, one or more neural radiance field models based on the plurality of images, wherein training, by the computing system, the one or more neural radiance field models based on the plurality of images comprises:

determining, by the computing system, a plurality of respective scene positions and a plurality of respective scene view directions for the plurality of images based on comparing feature location and feature sizes between images;

processing, by the computing system, one or more respective scene positions of the plurality of respective scene positions and one or more respective scene view directions of the plurality of respective scene view directions with the one or more neural radiance field models to generate one or more predicted view synthesis renderings, wherein the one or more predicted view synthesis renderings comprise one or more predicted color values and one or more predicted opacity values;

evaluating, by the computing system, a loss function that evaluates a difference between the one or more predicted view synthesis renderings and one or more respective images of the plurality of images; and

adjusting, by the computing system, one or more parameters of the one or more neural radiance field model based at least in part on the loss function.

8. The method of claim 1 , further comprising:

obtaining, by the computing system, a plurality of images of the environment;

processing, by the computing system, the plurality of images with a segmentation model to generate a plurality of segmented images, wherein the segmentation model generates segmentation masks for segmenting occlusions from an image;

generating, by the computing system, replacement data for the plurality of segmented images, wherein the replacement data is descriptive of predicted pixels for replacing masked regions of the plurality of segmented images; and

generating, by the computing system, a plurality of augmented images based on the plurality of segmented images and the replacement data, wherein the one or more neural radiance field models are trained on the plurality of augmented images.

9. The method of claim 1 , wherein the one or more neural radiance field models were trained on a plurality of images of the environment and lidar data for the environment, wherein training comprises evaluating a depth loss based on the lidar data.

10. A computing system for providing a virtual walkthrough interface, the system comprising:

one or more processors; and

one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations comprising:

obtaining one or more neural radiance field models, wherein the one or more neural radiance field models were trained to generate view renderings of one or more rooms;

processing a plurality of positions with the one or more neural radiance field models to generate a plurality of view synthesis renderings of the one or more rooms, wherein the plurality of positions are associated with a plurality of locations within the one or more rooms, and wherein the plurality of view synthesis renderings are descriptive of the one or more from the plurality of positions;

generating a forward-directional video based on the plurality of view synthesis renderings of the one or more rooms, wherein the forward-directional video is descriptive of a first sequence of views associated with traveling in a first direction through the one or more rooms;

generating a backward-directional video based on the plurality of view synthesis renderings of the one or more rooms, wherein the backward-directional video is descriptive of a second sequence of views associated with traveling in a second direction through the one or more rooms, wherein the second direction is opposite of the first direction;

generating a multi-directional video based on the forward-directional video and the backward-directional video, wherein the multi-directional video is descriptive of the first sequence of views and the second sequence of views; and

providing a virtual walkthrough interface for the one or more rooms by providing an interface for navigating through the multi-directional video.

11. The system of claim 10 , wherein the operations further comprise:

determining a walkthrough path based on processing a plurality of images of the one or more rooms; and

determining the plurality of positions based on the walkthrough path.

12. The system of claim 11 , wherein determining the walkthrough path based on processing the plurality of images of the one or more rooms comprises:

processing the plurality of images to determine a plurality of room landmarks associated with features of interest in the one or more rooms; and

generating the walkthrough path based on the plurality of room landmarks.

13. The system of claim 11 , wherein determining the plurality of positions based on the walkthrough path comprises:

determining a plurality of points on the walkthrough path, wherein the plurality of points comprise varying spacing based on regions of interest within the one or more rooms.

14. The system of claim 10 , wherein navigating through the multi-directional video comprises scrubbing through the multi-directional video.

15. The system of claim 10 , wherein the one or more rooms are associated with a restaurant, and wherein the virtual walkthrough interface is provided in a knowledge panel for the restaurant in a search results interface.

16. The system of claim 10 , wherein generating the multi-directional video based on the forward-directional video and the backward-directional video comprises:

determining a plurality of frame associations between the forward-directional video and the backward-directional video, wherein each of the plurality of frame associations is descriptive of corresponding frames associated with a same position in the one or more rooms;

generating multi-directional metadata based on the plurality of frame associations, wherein the multi-directional metadata is descriptive of the corresponding frames, and wherein the multi-directional metadata is configured to provide instructions for the virtual walkthrough interface to navigate to different portions of the multi-directional video based on a walkthrough direction and the plurality of frame associations; and

generating the multi-directional video by stitching the forward-directional video and the backward-directional video and embedding the multi-directional metadata.

17. One or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations, the operations comprising:

obtaining a plurality of images of an environment;

determining a plurality of image-capture positions and a plurality of image-capture view directions associated with the plurality of images;

training one or more neural radiance field models based on the plurality of images, the plurality of image-capture positions, and the plurality of image-capture view directions, wherein the one or more neural radiance field models are trained to generate view renderings of the environment;

processing a plurality of positions with the one or more neural radiance field models to generate a plurality of view synthesis renderings of the environment, wherein the plurality of positions are associated with a plurality of locations within the environment, wherein the plurality of view synthesis renderings are descriptive of the environment from the plurality of positions, and wherein the plurality of view synthesis renderings comprise a plurality of three-hundred and sixty degree view renderings, wherein the plurality of three-hundred and sixty degree view renderings are generated by generating a plurality of direction-based view renderings for a position and stitching the plurality of direction-based view renderings together to generate a panoramic image;

generating a virtual walkthrough video based on the plurality of view synthesis renderings of the environment, wherein the virtual walkthrough video is descriptive of a sequence of views of the environment; and

storing the virtual walkthrough video in a database, wherein storing the virtual walkthrough video comprises indexing the virtual walkthrough video with a geographic location associated with the environment.

18. The one or more non-transitory computer-readable media of claim 17 , wherein the virtual walkthrough video is descriptive of a sequence of views of the environment rendered in a forward progressing sequence and a backwards progressing sequence.

19. The one or more non-transitory computer-readable media of claim 17 , wherein the virtual walkthrough video comprises a three-hundred and sixty degree view video, wherein the virtual walkthrough video is formatted to be selectively cropped by a video player to provide one or more video directions for display during playback.

20. A computer-implemented method for providing a virtual walkthrough interface, the method comprising:

obtaining, by a computing system comprising one or more processors, a location-based query, wherein the location-based query is associated with obtaining information associated with a particular location;

obtaining, by the computing system, a virtual walkthrough video associated with the particular location based on the location-based query, wherein the virtual walkthrough video was generated by generating a plurality of view renderings with a neural radiance field model, and wherein the virtual walkthrough video is descriptive a sequence of views of an environment associated with the location rendered in a forward progressing sequence and a backwards progressing sequence, wherein the forward progressing sequence is associated with a first portion of the virtual walkthrough video, and wherein the backwards progressing sequence is associated with a second portion of the virtual walkthrough video;

providing, by the computing system, playback of a first set of frames of the virtual walkthrough video;

obtaining, by the computing system and during display of a particular frame in the first portion of the virtual walkthrough video, a navigation input, wherein the navigation input is descriptive of a request to perform a virtual walkthrough in an opposite direction;

determining, by the computing system, a corresponding frame in the second portion of the virtual walkthrough video, wherein the corresponding frame is associated with the particular frame; and

providing, by the computing system, playback of a second set of frames of the virtual walkthrough video starting with the corresponding frame.

21. The method of claim 20 , wherein the particular frame and the corresponding frame are associated with a same position in the environment.

22. A computing system for providing a virtual walkthrough interface, the system comprising:

one or more processors; and

one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations comprising:

obtaining a location-based query, wherein the location-based query is associated with obtaining information associated with a particular location;

obtaining a virtual walkthrough video associated with the particular location based on the location-based query, wherein the virtual walkthrough video was generated by generating a plurality of view synthesis renderings with a neural radiance field model, and wherein the plurality of view synthesis renderings comprise a plurality of three-hundred and sixty degree view renderings;

providing a virtual walkthrough associated with a first view direction by cropping the plurality of three-hundred and sixty degree view renderings of the virtual walkthrough video to display a first portion of the plurality of three-hundred and sixty degree view renderings associated with the first view direction;

obtaining a view direction input, wherein the view direction input is descriptive of a request to adjust a focal direction to a second view direction; and

providing the second view direction for display by cropping the plurality of three-hundred and sixty degree view renderings of the virtual walkthrough video to display a second portion of the plurality of three-hundred and sixty degree view renderings associated with the second view direction.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 15, 2023
From: MONTERO, CARLOS, JR.; SEEFELDER DE ASSIS ARAUJO, MARCOS; MOFFETT, CARDIN EVERETT; GORAN, CHARLES
To: GOOGLE LLC
Reel/Frame 065884/0554 →
Continuity (1)
Related Publication 20250166311A1 · May 22, 2025
References Cited (56)
US 11868672B1 · Dehkordi · 2024 [cited by examiner]
US 12328569B2 · Taylor · 2025 [cited by examiner]
US 20070198951A1 · Frank · 2007 [cited by applicant]
US 20070282792A1 · Bailly et al. · 2007 [cited by applicant]
US 20130097197A1 · Rincover et al. · 2013 [cited by applicant]
US 20200302241A1 · White et al. · 2020 [cited by applicant]
US 20240202987A1 · Sadr · 2024 [cited by examiner]
US 20240404191A1 · Liu · 2024 [cited by examiner]
US 20250086896A1 · Litany · 2025 [cited by examiner]
US 20250088617A1 · Iwao · 2025 [cited by examiner]
US 20250106370A1 · Chen · 2025 [cited by examiner]
US 20250131642A1 · Yang · 2025 [cited by examiner]
Yao, Mingyuan, et al. “Neural radiance field-based visual rendering: a comprehensive review.” arXiv preprint arXiv:2404.00714 (2024). (Year: 2024). [cited by examiner]
Xu, Linning, et al. “VR-NeRF: High-fidelity virtualized walkable spaces.” SIGGRAPH Asia 2023 Conference Papers. 2023. (Year: 2023). [cited by examiner]
Alley et al., “Neural Point-Based Graphics”, arXiv:1906.08240v3, Apr. 5, 2020, 16 pages. [cited by applicant]
Bojanowski et al., “Optimizing the Latent Space of Generative Networks.”, Thirty-Fifth International Conference on Machine Learning, Stockholm, Sweden, Jul. 10-15, 10 pages. [cited by applicant]
Brooks et al., “Unprocessing Images for Learned Raw Denoising”, arXiv:1811.11127v1, Nov. 27, 2018, 9 pages. [cited by applicant]
Buehler et al., “Unstructured Lumigraph Rendering.”, Twenty-eighth Conference on Computer Graphics and Interactive Techniques, Los Angeles, California, United States, Aug. 12-17, 2001, pp. 425-432 [cited by applicant]
Chen et al., “Rethinking Atrons Convolution for Semantic Image Segmentation.”, arXiv:1706.05587v3, Jun. 17, 2017, 14 pages. [cited by applicant]
Flynn et al., “DeepStereo: Learning to Predict New Views from the World's Imagery.”, Institute of Electrical and Electronics Engineers Conference on Computer Vision and Pattern Recognition, Las Vegas, Nevada, United Sta… [cited by applicant]
Flyan et al., “DeepView: View Synthesis with Learned Gradient Descent.”, Institute of Electrical and Electronics Engineers/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, California, United States… [cited by applicant]
Frahm et al., “Building Rome on a Cloudless Day.”, Eleventh European Conference on Computer Vision, Crete, Greece, Sep. 5-11, 2010, pp. 368-381. [cited by applicant]
Hartley et al., “Multiple View Geometry in Computer Vision.”, Cambria ty Press, New York, New York, United States, 2003, 673 pages. [cited by applicant]
Isola et al., “Image-to-Image Translation with Conditional Adversarial Networks.”, Institute of Electrical and Electronics Engineers Conference on Computer Vision and Pattern Recognition, Honolulu, Hawaii, United States… [cited by applicant]
Jin et al., “Image Matching Across Wide Baselines: From Paper to Practice.”, arXiv:2003.01587v5, Feb. 11, 2021, 27 pages. [cited by applicant]
Johnson et al., “Perceptual Losses for Real-Time Style Transfer and Super-Resolution.”, Fourteenth European Conference on Computer Vision, Amsterdam, Netherlands, Oct. 8-16, 2016, 17 pages. [cited by applicant]
Kendall et al., “What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?”, arXiv:1703.04977v2, Oct. 5, 2017, 12 pages. [cited by applicant]
Kingma, “Adam: A Method for Stochastic Optimization”, arXiv:1412.6980v9, Jan. 30, 2017, 15 pages. [cited by applicant]
Laffont et al., “Coherent Intrinsic Images from Photo Collections.”, Association for Computing Machinery Transactions on Graphics, vol. 31, Issue 6, No. 2012, 13 pages. [cited by applicant]
Levoy et al., “Light Field Rendering.”, Twenty-third Annual Conference on Computer Graphics and Interactive Techniques, New Orleans, Lousiana, United States, Aug. 4-9, 1996, 12 pages. [cited by applicant]
Li et al., “Crowdsampling the Plenoptic Function.”, Sixteenth European Conference on Computer Vision, Virtual, Aug. 23-28, 2020, 18 pages. [cited by applicant]
Lombardi et al., “Neural Volumes: Learning Dynamic Renderable Columes from Images.”, Association for Computing Machinery Transactions on Graphics, vol. 28, No. 4, Jul. 2019, pp. 1-14. [cited by applicant]
Martin-Brualla et al. “LookinGood: Enhancing Performance Capture with Real-time Nural Re-Rendering”, arXiv:1811.05029v1, Nov. 12, 2018, 14 pages. [cited by applicant]
Max, “Optical Models for Direct Volume Rendering.”, Institute of Electrical and Electronics Engineers Transactions on Visualization and Computer Graphics, vol. 1, Issue 2, Jun. 1995, pp. 99-108. [cited by applicant]
Pappas et al., “Perceptual Criteria for Image Quality Evaluation.”, Hanbook of Image and Video Processing, Academic Press, 2000, 974 pages. [cited by applicant]
Price et al., “Augmentin Crowd-Sourced 3D Reconstructions using Semantic Detections.”, Institute of Electrical and Electronics Engineers/CVP Conference on Computer Vision and Pattern Recognition, Salt Lake City, Utah, U… [cited by applicant]
Schonberger et al., “Structure-from-Motion Revisited.”, Institute of Electrical and Electronics Engineers Conference on Computer Vision and Pattern Recognition, Las Vegas, Nevada, United States, Jun. 27-30, 2016, pp. 41… [cited by applicant]
Shan et al., “The Visual Turing Test for Scene Reconstruction.”, International Conference on 3D Vision, Seattle, Washington, United States, Jun. 29-Jul. 1, 2013, 8 pages. [cited by applicant]
Sinim, “Image Based Rendering.”, Springer Scientific+Business Media LLC, New York, New York, United States, 2007, 425 pages. [cited by applicant]
Sitzmann et al., “DeepVoxels: Learning Persistent 3D Feature Embeddings.”, Institute of Electrical and Electronics Engineers/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, California, United Stat… [cited by applicant]
Sitzmang et al., “Seene Representation Networks: Continuons 3D-Structure-Aware Neural Scene Representations.”, Thirty-third Conference on Neural Information Processing Systems, Vancouver, Canada, Dec. 8-14, 2019, 12 pag… [cited by applicant]
Snavely et al., “Photo Tourism: Exploring Collections in 3D”, Association for Computing Machinery's Transactions on Graphics, vol. 25, Issue 3, Jul. 1, 2006, pp. 835-846. [cited by applicant]
Talebi et al., “NIMA: Neural Image Assessment”, axXiv:1709.05424v2, Apr. 26, 2018, 15 pages. [cited by applicant]
Tancik et al., “Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains.”, Thirty-fourth Conference on Neural Information Processing Systems, Vancouver, Canada, Dec. 6-12, 2020, 11 pages. [cited by applicant]
Tewari et al., “State of the Art on Neural Rendering”, arXiv:2004.03805v1 Apr. 8, 2020, 27 pages. [cited by applicant]
Thies et al., Deferred Neural Rendering: Image Synthesis using Neural Textures, arXiv:1904.12356v1, Apr. 28, 2019, 12 pages. [cited by applicant]
Thung, “A Survey of Image Quality Measures”, 2009 International Conference for Technical Postgraduates, Kuala Lumpur, Malaysia, Dec. 14-15, 2009, 5 pages. [cited by applicant]
Triggs et al., “Bundle Adjustment—A Modern Synthesis.”, International Workshop on Vision Algorithms, Corfu, Greece, Sep. 21-22, 1999, 72 pages. [cited by applicant]
Wang et al., “Multi-Scale Structural Similarity for Image Quality Assessment.”, Thirty-seventh Institute of Electrical and Electronics Engineer's Asilomar Conference on Signals, Systems, and Computer, Pacific Grove, Cal… [cited by applicant]
Wang et al., “A Universal Image Quality Index.”, Institute of Electrical and Electronics Engineers Signal Processing Letters, vol. 9, No. 3, Mar. 2002, 4 pages. [cited by applicant]
Invitation to Pay Additional Fees for Application No. PCT/US2024/052794, mailed Jan. 27, 2025, 5 pages. [cited by applicant]
Rematas et al., “Urban Radiance Fields”, IEEE/ CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, 11 pages. [cited by applicant]
Tancik et al., “Block-NeRF: Scalable Large Scene Neural View Synthesis”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, 11 pages. [cited by applicant]
Xie et al., “S-NERF: Neural Radiance Fields For Street Views”, arXiv:2303.00749v1, 24 pages. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/US2024/052794, mailed Mar. 20, 2025, 23 pages. [cited by applicant]
Artificial Intelligence, “Block NeRF: Scalable Large Scene Neural View Synthesis | CVPR 2022”, YouTube, Jun. 30, 2022, Retrieved from: https://www.youtube.com/watch?v=sibCaWMaDx8. [cited by applicant]