IP Library Granted Patent US 12670664
Granted Patent B2
US 12670664 · App. 18/110,487 · Granted Jun 30, 2026

Temporally coherent volumetric video

Inventors: Vsevolod Kagarlitsky (Ramat Gan, IL); Shirley Keinan (Tel Aviv, IL); Amir Green (Mitzpe Netofa, IL); Michal Heker (Tel Aviv, IL); Michael Birnboim (Holon, IL); Gilad Talmon (Tel Aviv, IL)
Assignee: TAKE-TWO INTERACTIVE SOFTWARE, INC.
G06T17/20G06T13/20G06T15/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12670664
App. No.
18/110,487
Filed
Feb 16, 2023
Granted
Jun 30, 2026
Kind
B2
Art Unit
2615
USPC
345/423
Abstract

A method for finding a deformation field transforming a source frame of a volumetric video into a target frame of a volumetric video comprising steps of building a texture implicit function for said target frame and training a neural network to generate said deformation field between said source frame and said target frame, said texture implicit function for said target frame being a texture matching loss for said neural network.

Claims (46)

1 . A method for finding a deformation field transforming a source frame of a volumetric video into a target frame of a volumetric video comprising steps of:

building a texture implicit function for said target frame;

training a neural network to generate said deformation field between said source frame and said target frame, said texture implicit function for said target frame being a texture matching loss for said neural network; and

no deformation field being findable between said source frame and said target frame, selecting an intermediate frame and generating a source-intermediate deformation field between said source frame and said intermediate frame using a texture implicit function for said intermediate frame, generating an intermediate-target deformation field between said intermediate frame and said target frame, and generating a source-target deformation field from said source-intermediate deformation field and said intermediate-target deformation field.

2 . The method of claim 1 , additionally comprising at least one of the following steps:

a. selecting said 3D model representation to be a signed distance field;

b. generating a lighting implicit function for said target frame, and using said lighting implicit function as a lighting matching loss for said neural network; or

c. generating a semantics implicit function for said target frame, and using said semantics implicit function as a semantics matching loss for said neural network.

3 . The method of claim 1 , additionally comprising steps of generating a 3D model implicit function for said target frame from a 3D model representation of said frame, and using said 3D model implicit function as a geometry matching loss for said neural network.

4 . The method of claim 1 , additionally comprising steps of:

selecting a generation path, said generation path comprising:

a representative frame R; and

a propagation path between pairs of frames (source frame k, target frame k) where target frame k of a frame in said generation path becomes source frame k for a next frame in said propagation path; and

applying the method for finding a deformation field transforming a source frame of a volumetric video into a target frame of a volumetric video to each pair of frames on said generation path.

5 . The method of claim 4 , additionally comprising at least one of the following steps:

a. said propagation path comprising a member selected from a group consisting of R-frame propagation where R frame propagation is propagation from representative frame R to all other frames, or moving propagation where moving propagation is propagating (representative frame R to frame n1), propagating (frame n1 to frame n2) and repeating propagation (frame nx to frame ny) until all frames are linkable, no two of said at least one generation path having both the same representative frame R and the same propagation path;

b. having a plurality of generation paths, and for each of said target frame k in said volumetric video, using the deformation fields from said plurality of generation paths having said target frame k to find an averaged deformation field having said target frame k, thereby generating an averaged volumetric video;

c. for a 3D object representation not comprising a mesh, for said representative frame R; applying an algorithm configured to generate a textured mesh for each object in said 3D object representation of R, and transforming said texture and said mesh of said representative frame R into texture and mesh for all other frames in said volumetric video via said source frame k-target frame k deformation fields;

d. storing topology only for representative frames.

6 . The method of claim 4 , additionally comprising a step of storing texture only for representative frames.

7 . The method of claim 4 , additionally comprising a step of selecting said representative frame R from frames comprising a characteristic selected from a group consisting of: a frame comprising a predetermined object, a frame comprising a predetermined pose, a predetermined frame in said volumetric video, or any combination thereof.

8 . The method of claim 4 , additionally comprising steps of storing or downloading said volumetric video as a frame-and-deformation field file, said frame-and-deformation field file comprising all between-frame deformation fields and, for each of said generation path, storing topology and texture for representative frame R.

9 . The method of claim 4 , additionally comprising a step of storing said representative frame as a textured mesh.

10 . A non-transitory computer-readable medium comprising computer-executable instructions which, when executed by a computing device, cause the computing device to carry out a method for finding a deformation field transforming a source frame of a volumetric video into a target frame of a volumetric video, the method comprising steps of:

building a texture implicit function for said target frame;

training a neural network to generate said deformation field between said source frame and said target frame, said texture implicit function for said target frame being a texture matching loss for said neural network; and

no deformation field being findable between said source frame and said target frame, to select an intermediate frame and generate a source-intermediate deformation field between said source frame and said intermediate frame using a texture implicit function for said intermediate frame, generate an intermediate-target deformation field between said intermediate frame and said target frame, and generate a source-target deformation field from said source-intermediate deformation field and said intermediate-target deformation field.

11 . The non-transitory computer-readable medium of claim 10 , wherein the computer-executable instructions are additionally configured, when executed, to perform at least one of:

a. select said 3D model representation to be a signed distance field;

b. generate a 3D model implicit function for said target frame from a 3D model representation of said frame, and use said 3D model implicit function as a geometry matching loss for said neural network; or

c. generate a semantics implicit function for said target frame, and use said semantics implicit function as a semantics matching loss for said neural network.

12 . The non-transitory computer-readable medium of claim 10 , wherein the computer-executable instructions are additionally configured, when executed, to generate a lighting implicit function for said target frame, and use said lighting implicit function as a lighting matching loss for said neural network.

13 . The non-transitory computer-readable medium of claim 10 , wherein the computer-executable instructions are additionally configured, when executed, to:

select a generation path, said generation path comprising:

a representative frame R; and

a propagation path between pairs of frames (source frame k, target frame k) where target frame k of a frame in said generation path becomes source frame k for a next frame in said propagation path; and

apply the method for finding a deformation field transforming a source frame of a volumetric video into a target frame of a volumetric video to each pair of frames on said generation path.

14 . The non-transitory computer-readable medium of claim 13 , wherein the computer-executable instructions are additionally configured, when executed, to perform at least one of:

a. said propagation path comprising a member selected from a group consisting of R-frame propagation where R frame propagation is propagation from representative frame R to all other frames, or moving propagation where moving propagation is propagating (representative frame R to frame n1), propagate (frame n1 to frame n2) and repeat propagation (frame nx to frame ny) until all frames are linkable, no two of said at least one generation path having both the same representative frame R and the same propagation path;

b. have a plurality of generation paths, and for each of said target frame k in said volumetric video, use the deformation fields from said plurality of generation paths having said target frame k to find an averaged deformation field having said target frame k, thereby generating an averaged volumetric video; or

c. for a 3D object representation not comprising a mesh, for said representative frame R; apply an algorithm configured to generate a textured mesh for each object in said 3D object representation of R, and transform said texture and said mesh of said representative frame R into texture and mesh for all other frames in said volumetric video via said source frame k-target frame k deformation fields;

d. store topology only for representative frames.

15 . The non-transitory computer-readable medium of claim 13 , wherein the computer-executable instructions are additionally configured, when executed, to store texture for representative frames only.

16 . The non-transitory computer-readable medium of claim 13 , wherein the computer-executable instructions are additionally configured, when executed, to select said representative frame R from frames comprising a characteristic selected from a group consisting of: a frame comprising a predetermined object, a frame comprising a predetermined pose, a predetermined frame in said volumetric video, or any combination thereof.

17 . The non-transitory computer-readable medium of claim 13 , wherein the computer-executable instructions are additionally configured, when executed, to store or download said volumetric video as a frame-and-deformation field file, said frame-and-deformation field file comprising all between-frame deformation fields and, for said generation path, store topology and texture for representative frame R.

18 . The non-transitory computer-readable medium of claim 13 , wherein the computer-executable instructions are additionally configured, when executed, to store said representative frame as a textured mesh.