IP Library › Granted Patent US 12,394,135
Granted Patent B2
US 12,394,135 · App. 18/065,600 · Granted Aug 19, 2025

Computing images of dynamic scenes

Inventors: Marek Adam Kowalski (Cambridge, GB); Matthew Alastair Johnson (Cambridge, GB); Jamie Daniel Joseph Shotton (Cambridge, GB)
Assignee: Microsoft Technology Licensing, LLC
G06T15/06A63F13/52G06N3/08G06N20/00G06T7/75G06T7/80G06T15/08G06T17/10G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,394,135
App. No.
18/065,600
Granted
Aug 19, 2025
Kind
B2
Abstract

Computing an output image of a dynamic scene. A value of E is selected which is a parameter describing desired dynamic content of the scene in the output image. Using selected intrinsic camera parameters and a selected viewpoint, for individual pixels of the output image to be generated, the method computes a ray that goes from a virtual camera through the pixel into the dynamic scene. For individual ones of the rays, sample at least one point along the ray. For individual ones of the sampled points, a viewing direction being a direction of the corresponding ray, and E, query a machine learning model to produce colour and opacity values at the sampled point with the dynamic content of the scene as specified by E. For individual ones of the rays, apply a volume rendering method to the colour and opacity values computed along that ray, to produce a pixel value of the output image.

Claims (52)

1. A computer-implemented method of training a machine learning model, the method comprising:

accessing a plurality of training images of a dynamic scene, the training images having been captured from a plurality of different viewpoints and at a plurality of different times;

for an individual image of the training images:

specifying a viewing direction of a pixel of the individual image according to a viewpoint of a capture device which captured the individual image;

specifying a value of E which is a parameter describing desired dynamic content of the dynamic scene in an output image of the dynamic scene; and

training the machine learning model using supervised learning given the training images such that the machine learning model produces a radiance value of an output three-dimensional point in the dynamic scene, the radiance value of the output three-dimensional point comprising a color value and an opacity value, given points in the dynamic scene, the viewing direction, and the value of E; the training including:

training the machine learning model to generate the color value based on both a location of the output three-dimensional point on a ray and a direction of the ray, and

training the machine learning model to generate the opacity value based on the location of the output three-dimensional point on the ray.

2. The computer-implemented method of claim 1 , further comprising:

the value of E having a type and a format;

the type of the value of E depending on the training images; and

the format of the value of E depending on the training images.

3. The computer-implemented method of claim 1 , wherein the dynamic scene is three-dimensional and comprises a moving object.

4. The computer-implemented method of claim 1 , further comprising, for individual ones of the training images, specifying intrinsic parameter values of the capture device associated with the output image, and a viewpoint for the capture device.

5. The computer-implemented method of claim 1 , wherein the value of E is specified using one or more of: a time when the individual image was captured, a value of parameters of a 3D model of an object in the dynamic scene at the time when the individual image was captured.

6. The computer-implemented method of claim 1 , wherein the machine learning model is a neural network with a plurality of layers, each layer comprising a plurality of nodes where each node has a weight.

7. The computer-implemented method of claim 6 , further comprising modifying the weight using the value of E.

8. An apparatus comprising:

a processor; and

a memory storing instructions that, when executed by the processor, perform a method of training a machine learning model comprising:

accessing a plurality of training images of a dynamic scene, the training images having been captured from a plurality of different viewpoints and at a plurality of different times;

for an individual image of the training images:

specifying a viewing direction of a pixel of the individual image according to a viewpoint of a capture device which captured the individual image;

specifying a value of E which is a parameter describing desired dynamic content of the dynamic scene in an output image of the dynamic scene; and

training the machine learning model using supervised learning given the training images such that the machine learning model produces a radiance value of an output three-dimensional point in the dynamic scene, the radiance value of the output three-dimensional point comprising a color value and an opacity value, given points in the dynamic scene, the viewing direction, and the value of E; the training including:

training the machine learning model to generate the color value based on both a location of the output three-dimensional point on a ray and a direction of the ray, and

training the machine learning model to generate the opacity value based on the location of the output three-dimensional point on the ray.

9. The apparatus of claim 8 , further comprising:

the value of E having a type and a format;

the type of the value of E depending on the training images; and

the format of the value of E depending on the training images.

10. The apparatus of claim 8 , wherein the dynamic scene is three-dimensional and comprises a moving object.

11. The apparatus of claim 8 , further comprising, for individual ones of the training images, specifying intrinsic parameter values of the capture device associated with the output image, and a viewpoint for the capture device.

12. The apparatus of claim 8 , wherein the value of E is specified using one or more of: a time when the individual image was captured, a value of parameters of a 3D model of an object in the dynamic scene at the time when the individual image was captured.

13. The apparatus of claim 8 , wherein the machine learning model is a neural network with a plurality of layers, each layer comprising a plurality of nodes where each node has a weight.

14. The apparatus of claim 13 , further comprising modifying the weight using the value of E.

15. A computer storage medium storing computer executable instructions that upon execution by a processor perform a method of training a machine learning model comprising:

accessing a plurality of training images of a dynamic scene, the training images having been captured from a plurality of different viewpoints and at a plurality of different times;

for an individual image of the training images:

specifying a viewing direction of a pixel of the individual image according to a viewpoint of a capture device which captured the individual image;

specifying a value of E which is a parameter describing desired dynamic content of the dynamic scene in an output image of the dynamic scene; and

training the machine learning model using supervised learning given the training images such that the machine learning model produces a radiance value of an output three-dimensional point in the dynamic scene, the radiance value of the output three-dimensional point comprising a color value and an opacity value, given points in the dynamic scene, the viewing direction, and the value of E; the training including:

training the machine learning model to generate the color value based on both a location of the output three-dimensional point on a ray and a direction of the ray, and

training the machine learning model to generate the opacity value based on the location of the output three-dimensional point on the ray.

16. The computer storage medium of claim 15 , further comprising:

the value of E having a type and a format;

the type of the value of E depending on the training images; and

the format of the value of E depending on the training images.

17. The computer storage medium of claim 15 , wherein the dynamic scene is three-dimensional and comprises a moving object.

18. The computer storage medium of claim 15 , further comprising, for individual ones of the training images, specifying intrinsic parameter values of the capture device associated with the output image, and a viewpoint for the capture device.

19. The computer storage medium of claim 15 , wherein the value of E is specified using one or more of: a time when the individual image was captured, a value of parameters of a 3D model of an object in the dynamic scene at the time when the individual image was captured.

20. The computer storage medium of claim 15 , wherein the machine learning model is a neural network with a plurality of layers, each layer comprising a plurality of nodes where each node has a weight, wherein the weight is modified using the value of E.

Priority Claims (1)
GB 2009058 · Jun 15, 2020 · national
Continuity (2)
Division 16927928 · Jul 13, 2020
Related Publication 20230116250A1 · Apr 13, 2023
References Cited (9)
US 10839557B1 · Arora · 2020 [cited by examiner]
US 20080246770A1 · Kiefer · 2008 [cited by applicant]
US 20120213430A1 · Nutter · 2012 [cited by applicant]
US 20180260975A1 · Sunkavalli · 2018 [cited by applicant]
US 20210367702A1 · Fang · 2021 [cited by examiner]
Chen et al. (Deep Video-Based Performance Synthesis from Sparse Multi-View Capture, Pacific Graphics 2019) (Year: 2019). [cited by examiner]
Mildenhall et al. (NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, ECCV, Mar. 2020) (Year: 2020). [cited by examiner]
Communication pursuant to Article 94(3) Received in European Patent Application No. 21731623.1, mailed on Mar. 4, 2025, 06 pages. [cited by applicant]
First Examination Report Received for Indian Application No. 202247070644, mailed on Jul. 7, 2025, 08 pages. [cited by applicant]
Cited By (1)
US 12,664,726